Senior Research Scientist | Model Steering
Job description
About the role
Become a part of an innovative AI company that specializes in developing intelligent solutions to tackle intricate business problems. In this position, you will be at the forefront of enhancing DeepL's next-generation translation models powered by large language models (LLMs). Your responsibilities will include leading efforts in fine-tuning, post-training, model steerability, and reinforcement learning. You will play a crucial role in conducting extensive research, prototyping novel concepts, executing large-scale experiments, and transitioning breakthroughs into practical applications.
Key facts
What you'll do
- Design and refine translation models that adapt to user preferences, contextual rules, and specific scenarios.
- Engage in hands-on research and development focused on post-training techniques for core translation models, including supervised fine-tuning, knowledge distillation, preference optimization, and reinforcement learning to enhance translation quality.
- Develop reward and evaluator models for translation tasks, utilizing rubric-based and reference-based grading, while addressing challenges such as reward hacking and failures in quality estimation.
- Enhance models to effectively integrate multimodal content and context, thereby improving the overall translation quality.
- Oversee the entire model delivery process, which includes prototyping, conducting ablation studies, training, evaluating, optimizing, and deploying models in real-time systems.
- Establish practices for evaluation, ensuring reproducibility, monitoring, and continuous improvement of models in production settings.
- Provide mentorship to fellow researchers and engineers, promoting teamwork and elevating model quality standards across the team.
- Collaborate with cross-functional teams to align research initiatives with product development and engineering objectives.
Requirements
- Proven experience in making large models steerable and capable of following instructions, employing techniques such as instruction tuning, latent space manipulation, steering vectors, or constrained encoding/decoding.
- In-depth practical knowledge of LLM post-training methods (including supervised fine-tuning and direct preference optimization), knowledge distillation (teacher-student frameworks), and reinforcement learning techniques (such as RLHF/RLAIF, PPO/GSPO, and reward modeling).
- Strong data-centric mindset for constructing synthetic and preference data pipelines, generating LLM-as-judge evaluations, and managing data curation, filtering, and reasoning about data mixtures and ablation studies.
- Experience in designing evaluation metrics and reward signals using automated metrics, LLM-as-judge evaluations, non-verifiable rewards, and incorporating human-in-the-loop feedback.
- A proactive approach to training models, conducting experiments, troubleshooting pipelines, and integrating machine learning systems into production environments, with a focus on delivering real-world impact and quality.
- Demonstrated ownership of significant research initiatives with effective execution, along with experience in mentoring team members in a dynamic, applied research environment.
- Proficient coding and experimentation skills in languages and frameworks such as Python, PyTorch, JAX, or TensorFlow, complemented by excellent communication skills to ensure alignment between research and product/engineering goals.
Nice to have
- Experience in fine-tuning and training large models at scale, including distributed or multi-node training techniques (e.g., FSDP, DeepSpeed, Megatron-style frameworks) and efficient training methodologies.
- Familiarity with fine-tuning existing reasoning models for specific tasks while maintaining their reasoning capabilities.
- Background in machine translation, multilingual natural language processing (NLP), or language quality assessment.
- Understanding of inference and serving at scale (e.g., vLLM, SGLang, TensorRT-LLM) and long-context modeling techniques.
- Publications in prestigious venues within the field.
Skills & tools
- Proficiency in Python and frameworks such as PyTorch, JAX, or TensorFlow.
- Strong understanding of machine learning concepts, particularly in the context of LLMs and translation models.
- Familiarity with data management and evaluation methodologies in machine learning.
Practical notes
DeepL offers a flexible hybrid work model, requiring employees to be in the office two days a week, along with adaptable working hours. Employees enjoy 30 days of annual leave in addition to public holidays. The company provides Virtual Shares, linking employee contributions directly to the growth of the organization. Benefits are competitive and tailored to the specific location of the employee. The team is diverse, representing over 90 nationalities, and regularly organizes in-person team events and monthly full-day hacking sessions to foster collaboration and innovation.