Staff Research Engineer
Job description
About the role
Synthesia is building interactive models that perceive and respond to users' actions and emotions, not just their words. This role focuses on advancing real-time, interactive experiences by integrating text, audio, and video models. You will help define and implement the core components of this vision within the R&D department.
Key facts
What you'll do
- Define the roadmap for new model capabilities and customer functionality, both short-term and long-term.
- Propose new multi-modal system architectures, especially for text and voice.
- Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis.
- Design solutions that enhance emotional expressiveness and natural interaction.
- Implement designs from pretraining through post-training.
- Integrate and test new architectures (neural codecs, diffusion, flow-matching) to improve realism and responsiveness.
- Create new evaluation metrics for conversational systems, including latency-aware and interaction-based measurements.
- Stay updated on the latest research in audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs.
- Curate new datasets to supplement existing data.
- Lead post-training initiatives like DPO, fine-tuning, and distillation to achieve production-ready models.
- Deploy models to production with optimized runtime and address customer feedback.
Requirements
- Ability to introduce new ideas and designs that advance interactive multimodal systems.
- Strong understanding of generative modeling, particularly for sequential or multimodal data.
- Hands-on experience with large language models or similar transformer-based architectures.
- High proficiency in PyTorch, including distributed training and model optimization.
- Solid understanding of time-series modeling and tokenization, preferably in audio, speech, or video contexts.
- Demonstrated ability to prototype quickly, test hypotheses, and iterate efficiently.
- Proven experience training deep learning models end-to-end, from data preparation to evaluation.
- Strong general software engineering skills for contributing to shared research infrastructure.
- Experience shipping a generative model into a live product at meaningful scale.
- Experience working on conversational or interactive systems where latency, responsiveness, and user experience were primary constraints.
- Experience with large-scale LLM trainings leading to models with good reasoning capabilities.
- Ownership of a research problem from architecture proposal through pretraining, post-training, and production deployment.
- Collaboration across modalities or teams (e.g., audio and video, or research and product) to deliver a unified system.
Nice to have
- Experience with real-time or streaming architectures.
- Familiarity with architectures in audio and speech generation, such as diffusion models, neural codecs, flow-matching models, or autoregressive decoders.
- Excellence in one or more of the following modalities: voice, text, video.
- Evidence of original research contributions, such as publications or open-source work at top-tier venues (e.g., NeurIPS, CVPR, ICML, ICLR, Interspeech).
Skills & tools
- PyTorch
- Generative modeling
- Transformer-based architectures
- Time-series modeling
- Tokenization
- Distributed training
- Model optimization
- DPO
- Fine-tuning
- Distillation
- Neural codecs
- Diffusion models
- Flow-matching models
- Autoregressive decoders
Practical notes
- This is a remote position within Europe.
About the company
Synthesia Limited is a British multinational artificial intelligence company based in London, United Kingdom. The firm develops synthetic media generation software that creates AI-generated video content. Its technology produces audio-visual agents and cloned avatars for professional use. The platform allows organizations to generate video material without traditional filming equipment or production crews. Users can select from a library of diverse avatars or create custom digital twins. The system supports multiple languages and accents for global communication needs. The company serves a significant portion of the corporate market. It is used by seventy percent of FTSE 100 companies. Over ninety percent of Fortune 100 companies also rely on the platform. Clients apply the technology for internal training and onboarding programs. Marketing teams use it for personalized outreach and product explainers. Customer support departments deploy AI agents for scalable service delivery. The software reduces the time and cost associated with conventional video production. It enables rapid updates to content when information changes. The firm focuses on enterprise-grade security and compliance standards. It operates as Britain's largest generative AI company by market presence. The leadership team continues to invest in research to improve avatar realism and motion quality. The goal is to make video creation accessible to anyone with a script.