Senior ML Researcher
Job description
About the role
You will own the research direction and implementation of temporal ontology extraction methods that turn messy software artifacts into a coherent, living knowledge structure. You will design contradiction detection systems that surface inconsistencies across code, documentation, and issue trackers through careful reasoning. You will build and validate semantic alignment techniques that keep a living spec synchronized with the reality of the software system. You will collaborate closely with JetBrains Research and external academic advisors to translate theoretical ideas into validated research contributions. You will define datasets, benchmarks, and metrics that prove the measurable impact of your methods on the overall system. You will prototype novel ideas rapidly, then partner with ML and software engineers to ship robust production solutions. You will represent Spectrum in the research community by publishing findings at top venues and articulating its value to diverse stakeholders.
Key facts
What you'll do
- Develop methods for LLM-based ontology construction from heterogeneous software artifacts, including code, documentation, and issue trackers, with a focus on contradiction detection and reasoning over the resulting knowledge graph.
- Create datasets, metrics, and benchmarks that drive measurable improvements across the system and provide empirical evidence for semantic alignment quality.
- Prototype and validate research ideas in experimental settings, then collaborate with ML and software engineers to integrate validated approaches into production workflows.
- Collaborate with JetBrains Research and external academic advisors to shape the research agenda and ensure alignment with前沿 knowledge graph and ontology learning communities.
- Publish research findings at top-tier academic venues and represent Spectrum in relevant research forums and practitioner discussions.
- Design and run rigorous experiments, including hypothesis formulation, ablation studies, and statistical evaluation, to assess the impact of novel extraction and alignment techniques.
- Implement strong Python and PyTorch-based solutions to perform structured knowledge extraction, enabling robust ontology construction from real-world software data.
- Communicate complex technical concepts clearly and effectively to both technical and non-technical audiences, ensuring stakeholders understand trade-offs and benefits.
- Maintain proficiency in English, both written and verbal, to facilitate seamless collaboration across global teams and research publications.
- Explore ontology engineering practices, semantic web technologies such as OWL, RDF, and SPARQL, and ontology alignment methods to enrich the living spec paradigm.
- Contribute background in code analysis, developer tools, or software engineering research to ensure that extracted ontologies reflect practical engineering realities.
- Apply knowledge graph embedding methods and graph neural networks to improve representation learning and reasoning over large software knowledge bases.
- Engage with early-stage startup dynamics, thriving in the zero-to-one phase where novel ideas are iterated quickly and ownership is high.
- Actively contribute to relevant open-source projects, sharing insights and building reusable components that advance the state of the art in ontology learning.
Requirements
- Hold a PhD (or equivalent research experience) in NLP, knowledge graphs, ontology learning, information extraction, or a closely related quantitative field.
- Maintain a strong publication record in at least one of the following areas: knowledge graph construction, ontology learning, information extraction, or natural language processing.
- Demonstrate experience applying LLMs to structured knowledge extraction tasks, showing concrete results in transforming unstructured text into usable schemas.
- Show competence in designing and running rigorous experiments, including hypothesis formulation, ablation studies, and robust statistical evaluation.
- Exhibit strong Python and PyTorch skills, with a proven ability to implement research ideas from scratch and iterate based on empirical findings.
- Possess excellent communication skills, with the capacity to explain intricate technical concepts to diverse audiences and stakeholders.
- Demonstrate proficiency in English, both written and verbal, to enable effective collaboration and clear dissemination of research outcomes.
- Have deep familiarity with ontology engineering principles and semantic web technologies, including practical exposure to OWL, RDF, and SPARQL.
- Bring background in code analysis, developer tooling, or software engineering research that aligns the work with real development challenges.
- Have hands-on experience with knowledge graph embedding methods or graph neural networks to model complex relationships in software artifacts.
- Show a history of working in early-stage startups, where you are comfortable navigating ambiguity and driving projects from initial concept to implementation.
- Contribute to open-source communities relevant to NLP, knowledge graphs, or ontology learning, demonstrating a commitment to shared progress.
Practical notes
LENGTH: 700-900 words. No HTML, no markdown, no em dashes.
Output the page only.