
Staff Research Engineer
Job description
AI.
About the role
You design and iterate production grade training pipelines that power the next generation of JetBrains AI coding tools. This role owns model behavior from raw data ingestion through deployment, defining technical direction grounded in measurable accuracy and latency targets. You prefer hands on shipping over endless discussion, taking direct responsibility for end to end subsystems while collaborating closely with distributed teams across research and product. You track every experiment with rigorous metrics and clear comparisons, translating findings into contributions that advance the state of the art for software engineering AI.
Key facts
What you'll do
Designing training regimes and data strategies that scale across thousands of GPU hours in demanding production environments.
Curating and processing massive datasets for model pre training and ongoing specialization, ensuring quality and consistency across sources.
Building and tuning core model architectures to balance peak capability with strict efficiency and latency constraints.
Partnering with product managers and engineers to align model releases with the JetBrains AI roadmap and customer workflows.
Shipping reliable inference paths that serve millions of developers globally through JetBrains IDE integrations and cloud services.
Supporting and extending existing training and deployment subsystems to maintain robustness as models and tooling evolve.
Evaluating experiments with detailed metrics, error analysis, and ablation studies to guide research decisions and prioritize work.
Contributing to scientific publications and technical blogs that showcase JetBrains AI advances to the broader research community.
Driving standard setting for data quality, model performance, and operational reliability across the AI organization.
Championing best practices in reproducibility, monitoring, and debugging of large scale language model training systems.
Requirements
You have hands on experience designing and running production ML systems that survive real world scale and variability.
You possess a solid theoretical foundation in NLP and transformer based modeling approaches, including attention mechanisms and training dynamics.
You are proficient with PyTorch and common libraries for NLP processing, writing clean and maintainable training code.
You have worked on distributed training of models with many billion parameters, tuning communication and computation for efficiency.
You pay attention to detail and communicate clearly with both technical and non technical stakeholders, documenting decisions and tradeoffs.
Nice to have
Experience with LLM inference frameworks such as vLLM, DeepSpeed, TensorRT to optimize deployment and throughput.
Experience with LLM alignment techniques such as RLHF or RLAIF for shaping model behavior to user intent.
Familiarity with MLOps tools and CI CD practices for machine learning, automating experiments and deployments.
Experience working on Kubernetes and Kubeflow platforms for orchestrating large scale training workloads.
A track record of scientific publications in NLP related fields, demonstrating impact on the research community.
Practical notes
This is a full time position open in multiple locations across Europe, with remote options available in Germany.
No specific visa or travel deadlines are listed in the source material.
How we develop JetBrains AI
We rely on a cluster of hundreds of NVIDIA GPUs as core training infrastructure, managed through Kubernetes and Kubeflow.
Our stack centers on Python, PyTorch, and HuggingFace, integrated with experiment tracking via Weights & Biases and CI automation through TeamCity.
Git serves as the source control backbone for code, data pipelines, and configuration that drive reproducible experiments.
You will train LLMs from scratch on large GPU clusters, collect and process pre training and fine tuning datasets, and support and improve critical subsystems.
In this role, you will work with stakeholders to convert business requirements into technical specifications, own an entire subsystem end to end, and prioritize tasks based on customer needs.
We value engineers who can plan projects and make decisions independently, consult with others when needed, start with the simplest solutions, and gradually add complexity as understanding deepens.
If you have experience in design, deployment, and support of production ML systems, a strong theoretical background in NLP and transformer based approaches, attention to detail, and great communication skills, we'd be happy to have you on our team.
We'd be especially thrilled if you have experience with LLM inference frameworks including vLLM, DeepSpeed, TensorRT, LLM alignment techniques such as RLHF or RLAIF, MLOps tools and CI CD for machine learning, Kubernetes and Kubeflow, and a record of scientific publications in the NLP field.
JetBrains builds cutting edge development tools, including IntelliJ IDEA, the leading Java IDE, and the Kotlin programming language, enabling teams to deliver software with confidence and joy.