
Senior Research Engineer
Job description
About the role
Senior Research Engineer, Code World Models
JetBrains seeks a Senior Research Engineer to advance Code World Models: systems that learn how software behaves, evolves, executes, and interacts with developer tooling. This role focuses on foundational model development, with responsibility for data strategy, training pipelines, and evaluation for code-centric systems. You will directly influence how models perceive, reason about, and generate software. You will own the design and execution of research initiatives that push the state of code understanding and generation. Your work will span the full lifecycle of model development, from data curation to production-oriented evaluation. You will collaborate closely with engineers and scientists to translate research insights into reliable JetBrains products.
Key Responsibilities
- Architect and manage training workflows for code foundation models, including initial pre-training and iterative improvement phases.
- Build robust data ingestion and filtering systems for massive codebases, ensuring quality, diversity, and integrity.
- Design and maintain large-scale experiment tracking to compare training runs and validate improvements.
- Develop repository-level evaluation suites that test reasoning, synthesis, and interaction with real-world software artifacts.
- Partner with product and research teams to align model behavior with practical developer tool requirements.
- Implement execution-aware evaluation methods using tests, traces, and runtime feedback.
- Ensure reproducibility and contamination control across datasets and experiments.
- Drive infrastructure improvements that support efficient distributed training and debugging.
Qualifications
You must meet the following criteria to succeed in this position.
- Demonstrated hands-on experience with model training workflows for code, including continued and mid-training iterations.
- Expert-level Python engineering skills and fluency with mainstream deep learning frameworks.
- Deep understanding of large-scale training pipelines, including data processing, checkpointing, and distributed execution.
- History of working with complex data sources such as code repositories, test suites, and execution logs.
- Formal background in NLP, machine learning for software, or a closely related discipline.
Preferred Experience
- Prior work on code generation, understanding, program repair, or test synthesis.
- Exposure to execution-based signals such as compiler output, runtime states, and sandboxed execution.
- Familiarity with evaluation challenges like benchmark leakage and long-horizon task completion.
- Contributions to ML platform code, libraries, or open-source training systems.
- Experience managing models with very large parameter counts in distributed settings.
Work Environment
This role is based in one of our primary locations: Amsterdam, Netherlands; Belgrade, Serbia; Berlin, Germany; Limassol, Cyprus; London, United Kingdom; Madrid, Spain; Munich, Germany; Paphos, Cyprus; Prague, Czech Republic; Warsaw, Poland; Yerevan, Armenia.
Compensation and Benefits
The target annual compensation for this role is 95,000 EUR, reflecting our commitment to rewarding high-impact technical leadership. Employment is full-time, with comprehensive benefits aligned with local regulations.
Equal Opportunity
JetBrains is an equal opportunity employer. We welcome applicants from all backgrounds, identities, and experiences. Your unique perspective helps us build better tools for developers around the world.
Please review official application materials for specific instructions and submission details.
About the company
JetBrains is a cutting-edge software vendor specializing in the creation of intelligent development tools, including IntelliJ IDEA - the leading Java IDE, and the Kotlin programming language.
The role of a centers on advancing the frontier of Code World Models, with a core mandate to architect, execute, and iteratively refine systems that comprehend and generate software at scale. You will own the complete lifecycle of model development, translating abstract research hypotheses into concrete, production-aligned training pipelines and evaluation frameworks. This position requires deep ownership of data strategy, from constructing massive, high-quality code corpora to designing rigorous filtering mechanisms that preserve integrity and diversity across programming languages and ecosystems. You will be responsible for implementing robust experiment tracking infrastructures that enable precise comparison of training runs, facilitating data-driven decisions on architecture, hyperparameters, and learning dynamics. A critical aspect of this role involves developing repository-level evaluation suites that move beyond static benchmarks to test reasoning, synthesis, and interaction with real-world software artifacts in realistic scenarios. Collaboration with product and research stakeholders will be essential to ensure that model capabilities directly address practical needs within developer tooling, aligning technical advances with user value. You will implement execution-aware evaluation methods that leverage compiler output, runtime states, and sandboxed execution to assess correctness, efficiency, and robustness. Ensuring reproducibility and preventing contamination across datasets and experiments will be a fundamental discipline, requiring meticulous versioning, provenance tracking, and rigorous validation protocols. You will also drive infrastructure improvements that support efficient distributed training, debugging, and monitoring, optimizing resource utilization and accelerating iteration cycles.
In terms of qualifications, the role demands demonstrated hands-on experience with model training workflows specifically for code, including the ability to manage continued and mid-training iterations that adapt models to evolving data and objectives. Expert-level Python engineering skills are non-negotiable, as is fluency with mainstream deep learning frameworks that power large-scale neural network training and inference. A deep understanding of large-scale training pipelines is essential, encompassing data processing, efficient checkpointing strategies, and reliable distributed execution across complex hardware configurations. Candidates must have a history of working with complex data sources such as code repositories, test suites, and execution logs, extracting meaningful signals from heterogeneous and often noisy inputs. A formal background in NLP, machine learning for software, or a closely related discipline provides the theoretical foundation necessary for success in this role.
Preferred experience highlights prior work on code generation, understanding, program repair, or test synthesis, where direct impact on software engineering tasks has been demonstrated. Exposure to execution-based signals such as compiler output, runtime states, and sandboxed execution is crucial for building models that interact meaningfully with running software. Familiarity with evaluation challenges like benchmark leakage and long-horizon task completion will enable you to design experiments that reflect real-world complexity and avoid common pitfalls. Contributions to ML platform code, libraries, or open-source training systems are valued, as they demonstrate an ability to build tools that empower broader research communities. Experience managing models with very large parameter counts in distributed settings is preferred, given the scale and complexity of modern code foundation models.
This role operates within a distributed network of primary locations, offering flexibility in work environments while maintaining strict standards for collaboration and output. Locations include Amsterdam, Netherlands; Belgrade, Serbia; Berlin, Germany; Limassol, Cyprus; London, United Kingdom; Madrid, Spain; Munich, Germany; Paphos, Cyprus; Prague, Czech Republic; Warsaw, Poland; and Yerevan, Armenia. Compensation for this position is structured around a target annual rate of 95,000 EUR, reflecting the scope, responsibility, and impact associated with leading code world model research. Employment is full-time, with comprehensive benefits aligned with local regulations to ensure stability and support across all locations. JetBrains operates as an equal opportunity employer, welcoming applicants from all backgrounds, identities, and experiences, recognizing that diverse perspectives drive innovation in developer tools. Applicants are directed to review official application materials for specific instructions and submission details, ensuring a standardized and transparent hiring process.