Senior Research Scientist | Multimodal Systems
Job description
About the role
This position centers on pushing the boundaries of how DeepL understands and translates complex visual documents through advanced multimodal reasoning. You will own the end to end lifecycle of next generation vision and layout understanding models, from exploratory research to robust production deployment. A core part of your work will involve designing systems that accurately interpret document structure, typography, and visual elements across diverse languages. You will collaborate closely with engineering and product teams to ensure research breakthroughs translate into tangible improvements for users. The role demands a strong partnership with the broader AI organization to align multimodal advances with company wide language and agent initiatives. You will have the opportunity to define technical standards and best practices for multimodal data within the translation pipeline. This is a hands on role where your contributions will directly shape the future of professional communication across different scripts and visual formats. Your work will influence both the immediate product roadmap and the long term research direction of the company.
Key facts
Location: UK
Engagement: Full-time
This is a hybrid workplace role.
What you'll do
Lead the design and implementation of novel multimodal architectures that process documents as combinations of text, images, and layout.
Perform rigorous experimentation to evaluate how visual context influences translation quality, and iterate on models based on empirical evidence.
Partner with data teams to curate and structure high quality document datasets that capture diverse layouts, fonts, and imaging conditions.
Define and maintain simulation environments that allow new multimodal ideas to be tested safely before exposure to production traffic.
Instrument production systems to capture multimodal specific signals, using these insights to drive continuous model refinement.
Develop evaluation frameworks that measure not only linguistic accuracy but also structural and visual fidelity in translated outputs.
Collaborate with product managers to translate user requirements for document understanding into concrete research objectives and success metrics.
Work with cross functional teams to ensure multimodal capabilities integrate seamlessly with existing translation and agent workflows.
Champion reproducible research by maintaining clear documentation, versioned experiments, and open communication of results.
Act as a technical leader in the community by sharing findings, contributing to open source components where appropriate, and mentoring junior researchers.
Establish benchmarks for multimodal performance on document translation tasks, enabling the company to track progress over time.
Identify failure modes in current systems and propose concrete research directions to mitigate risks before scaling.
Liaise with infrastructure teams to optimize training and inference pipelines for multimodal workloads and large model checkpoints.
Represent DeepL in relevant academic and industry forums to ensure the company remains at the forefront of multimodal research.
Requirements
You hold a PhD in computer vision, machine learning, or a closely related field with a strong publication record in top tier venues.
You have hands on experience building and training deep learning models, with a proven track record of delivering production grade systems.
You are proficient in Python and modern deep learning frameworks, with a solid understanding of transformer based architectures.
You have a strong grasp of computer vision techniques, including image classification, detection, segmentation, and layout analysis.
You understand the challenges of document understanding, including OCR errors, diverse layouts, and multilingual rendering issues.
You have experience working with large scale datasets and the computational constraints of training large models.
You are comfortable working in a fast paced environment where research priorities evolve with business needs.
You communicate complex technical concepts clearly to both technical and non technical stakeholders.
Nice to have
Experience with large language models or agentic systems is a strong advantage for aligning multimodal components with language workflows.
Familiarity with deployment pipelines for AI models, including containerization and monitoring in production, is valued.
Contributions to open source projects or participation in relevant research communities demonstrate active engagement beyond the immediate role.
Practical notes
This is a hybrid workplace role.
Employment is full time.
The position is based in London.