
Senior Research Engineer
Job description
About the role
Senior Research Engineer, Kotlin AI Value Stream
JetBrains centers its mission on code. Since 2000, the company has pursued the creation of the most powerful developer tools worldwide. Today, AI-assisted coding agents are integral to how developers write Kotlin, and JetBrains is dedicated to ensuring that these interactions produce high-quality results. The Kotlin AI Value Stream team leads this effort, concentrating on how AI agents comprehend, generate, and refine Kotlin code across all deployment targets, including Android, Kotlin Multiplatform, server-side, web, and desktop. The team is responsible for the evaluation frameworks, error diagnosis tools, and post-training workflows that assess and elevate agent behavior during real-world Kotlin development.
This role is a research position focused on the complete feedback cycle. You will analyze agent failures in Kotlin contexts, design evaluations that expose these failures, research corrective methods, and validate improvements. The work directly influences the experience of millions of developers using Kotlin through AI-assisted coding agents.
Core Responsibilities
You will build specialized tooling for understanding agent mistakes. This involves designing systems that capture, categorize, and examine errors produced by AI coding agents when they generate Kotlin code. You will also construct observability mechanisms that study agent session traces from JetBrains IDE plugins, Junie, Claude Code, Cursor, and other coding platforms.
You will design and maintain evaluation frameworks that assess Kotlin code generation across multiple quality dimensions. These dimensions include technical correctness, adherence to idiomatic Kotlin standards, build reliability, appropriate use of frameworks, and test completeness. You will establish simulation environments where agents can be tested on realistic Kotlin development scenarios. These scenarios range from initiating multiplatform projects and managing Gradle dependencies to migrating Java codebases to Kotlin. Ownership of the evaluation infrastructure includes metrics definition, experiment tracking, automated regression testing, and reproducible benchmarking.
You will research methods to refine agent and model behavior specific to Kotlin. This includes experimenting with post-training techniques such as supervised fine-tuning, direct preference optimization, and reinforcement learning to improve handling of Kotlin-specific patterns and libraries. You will explore context engineering strategies, including the use of configuration files, compiler feedback loops, Language Server Protocol integrations, and toolchains built for multi-component platforms. You will run controlled experiments to measure impact through A/B testing, benchmark analysis, and comparative studies on actual codebases. Collaboration with model providers such as Anthropic, OpenAI, and Google will help translate Kotlin-specific insights into model training and agent behavior adjustments.
You will also create public benchmark suites that set industry standards for Kotlin AI coding evaluation. These open-source benchmarks will cover a wide spectrum of Kotlin usage, including server-side applications with Spring and Ktor, multiplatform projects, build systems with Gradle, Android development, and library creation. The datasets will blend mined real-world tasks with carefully designed synthetic examples that test specific Kotlin capabilities. Benchmarks will be maintained and updated as models evolve to ensure they stay challenging, relevant, and resistant to data contamination.
Qualifications
You bring hands-on experience constructing evaluation or analysis pipelines for large language models or AI coding agents in research or production environments. You demonstrate strong Python engineering skills with at least three years of professional practice, writing clean and maintainable code within data-intensive and machine learning adjacent systems. You are proficient in querying large datasets using SQL or Athena and in performing statistical analysis of experimental outcomes. You can manage projects from initial problem identification through evaluation design, experimentation, and deployed fixes. You understand how developers use agents and can convert real-world failure patterns into evaluation and training tasks. You are fluent in Kotlin or are highly motivated to achieve deep Kotlin expertise, as you will work with Kotlin code on a daily basis.
Additional Assets
Experience contributing to open-source benchmarks or data-centric machine learning workflows is valued.
Tools and Skills
Proficiency in Kotlin, Python, SQL, Athena, Git, JetBrains IDEs, Junie, Claude Code, and Cursor.
Practical Information
Please verify all information on the official application page.
What you'll do
- Meet the bar Practical notes