Machine Learning Engineer, Evaluation
Job description
About the role
You will own the design and execution of evaluation frameworks that determine how effectively humans and AI systems collaborate on software engineering tasks. This role requires you to build measurement infrastructure that is both scientifically rigorous and operationally reliable at massive scale. You will be responsible for translating ambiguous questions about skill into concrete, auditable evaluation pipelines that withstand scrutiny from enterprise customers. The work involves close collaboration with product teams, researchers, and customers to iteratively refine what success looks like in assessment. You will also define the guardrails and experiments that keep evaluation honest as models evolve and new forms of AI assistance emerge. Ultimately, you will shape the methodology that the industry uses to judge AI-era engineering talent.
Key facts
What you'll do
Design and implement LLM-powered evaluation pipelines that assess AI usage skills consistently, fairly, and at production scale across thousands of daily assessments.
Own the evaluation methodology end to end, including rubric definition, model application logic, measurement of evaluation correctness, and ongoing auditing for bias and drift.
Design and run controlled experiments to determine what good evaluation looks like, generating insights where the answer is unknown and the methodology must be invented.
Build RAG pipelines and fine-tuning workflows that ensure evaluation models adhere reliably to defined rules and produce explainable, reproducible scores.
Define and maintain benchmarking infrastructure that tracks evaluation quality over time, detects regressions before candidates do, and provides empirical evidence of improvement.
Translate complex model behavior into clear narratives and artifacts that product managers, enterprise customers, and candidates can understand and trust.
Collaborate closely with product and engineering partners to align evaluation capabilities with real-world hiring workflows and customer expectations.
Investigate and prototype new evaluation techniques as the tooling landscape evolves, ensuring HackerRank remains at the forefront of assessment innovation.
Establish monitoring systems that surface anomalies in evaluation outcomes, enabling rapid response to potential fairness or consistency issues.
Document evaluation procedures, assumptions, and limitations so that methodologies can be audited, reviewed, and improved by diverse stakeholders.
Work with longitudinal datasets to understand trends in AI-assisted development and how evaluation approaches must adapt over time.
Champion a research mindset within the evaluation team, encouraging exploration, hypothesis-driven work, and evidence-based decision-making.
Requirements
You have shipped LLM-powered systems in production where consistency and reliability were hard constraints, not nice-to-haves, and you understand the operational challenges that come with scale.
You think as rigorously about how you measure your model as about the model itself, recognizing that a poorly constructed eval is a worse outcome than a weaker model.
You have a research mindset and are comfortable operating in a space where the right methodology does not yet exist and needs to be invented through experimentation.
You think in systems, understanding how data, models, and human workflows interact to produce evaluation outcomes that influence real hiring decisions.
You are fluent in modern ML frameworks and tooling, with hands-on experience building data pipelines that are robust, observable, and maintainable.
You have strong written communication skills and can break down complex model behavior into clear explanations for both technical and non-technical audiences.
You are comfortable working with ambiguous problems where success criteria are defined iteratively and must be discovered through exploration.
You care deeply about fairness and bias in AI systems and are diligent about designing evaluation processes that minimize unfair advantages across diverse candidates.
Nice to have
Experience with large language models, including fine-tuning, RAG, and prompt engineering in evaluation contexts.
Background in educational assessment, psychometrics, or measurement theory that can inform rigorous evaluation design.
Exposure to hiring technology platforms and the challenges of skills-based evaluation in enterprise environments.
Practical notes
This role involves significant work with live candidate data and model outputs, requiring strong judgment around privacy, security, and ethical evaluation practices.