Machine Learning Scientist
Job description
About the role
You will design and execute open research programs that measure how AI models perform in real-world workflows, translating human experience into rigorous evaluation protocols. You will own the design, execution, and open release of large-scale experiments that dissect model capabilities across reasoning, style, robustness, and user preference dimensions. You will serve as the scientific spokesperson for Arena Intelligence, articulating the vision and methodology of open evaluation to researchers, engineers, and the broader community. You will curate high-impact datasets and benchmarks that power the public leaderboard and underpin community tools used by leading labs and independent researchers. You will collaborate deeply with engineering and product teams to translate experimental insights into production-ready evaluation frameworks. You will prototype novel evaluation methodologies, balancing scientific rigor with the speed required to address emerging model behaviors. You will champion transparency by releasing code, data, and analyses that empower the entire research ecosystem to build upon your work.
Key facts
What you'll do
- Design and conduct experiments that probe AI model behavior across reasoning, style, robustness, and user preference dimensions using diverse evaluation paradigms.
- Develop new metrics, evaluation methodologies, and statistical frameworks that move beyond traditional benchmarks to capture real-world model utility and trustworthiness.
- Analyze large-scale human voting and interaction data to extract actionable insights about model performance, user satisfaction, and preference patterns.
- Communicate research findings to the broader community through academic papers, educational content, conference talks, and open-source documentation.
- Collaborate with engineers to translate experimental results into scalable production systems that integrate evaluation workflows into the Arena platform.
- Prototype and iterate on research ideas rapidly, balancing methodological rigor with the speed needed to evaluate emerging model capabilities.
- Partner with model providers to define evaluation questions, align testing protocols, and support responsible model assessment practices.
- Contribute to the scientific integrity and transparency of the LMArena leaderboard, ensuring that public benchmarks reflect rigorous and human-centered evaluation standards.
- Curate and maintain open datasets that capture diverse human preferences, enabling reproducible research and fostering innovation across the AI ecosystem.
- Implement and refine preference modeling techniques, including reinforcement learning from human feedback, direct preference optimization, and contrastive learning approaches.
- Establish reproducible experimental pipelines that support systematic ablation studies, sensitivity analysis, and robustness checks across model architectures.
- Act as an ambassador for Arena Intelligence in the research community, building relationships that strengthen collaboration and data sharing.
Requirements
- Holds a PhD or equivalent research experience in Machine Learning, Natural Language Processing, Statistics, or a closely related quantitative field.
- Demonstrates a track record of designing novel training objectives, evaluation schemes, or statistical frameworks that improve model reliability and alignment.
- Possesses strong fluency in Python and experience with the full experimental stack, from dataset design and large-batch training to rigorous evaluation and ablation.
- Has hands-on experience training large-scale models, including reward models, preference models, and fine-tuning LLMs using methods such as RLHF, DPO, and contrastive learning.
- Shows deep understanding of modern LLMs and deep learning architectures, including Transformers, diffusion models, and reinforcement learning with human feedback.
- Exhibits a collaborative mindset, working closely with engineers to productionize research insights and iterating with product teams to align research with user needs.
- Communicates complex technical concepts clearly to both technical and non-technical audiences through writing, speaking, and open-source contributions.
- Is comfortable being a visible representative of Arena Intelligence, engaging openly with the research community and building a strong personal brand.
- Thrives in an interdisciplinary environment, coordinating with engineers, product managers, marketing, and the broader research community to advance model evaluation practices.
- Is committed to openness, transparency, and scientific integrity, ensuring that methods, data, and results are shared responsibly with the full ecosystem.
Nice to have
- Experience contributing to open-source evaluation frameworks, leaderboards, or benchmark datasets used by the research community.
- Background in human-computer interaction or user studies related to AI evaluation and preference modeling.
- Familiarity with production-scale data pipelines and experiment tracking systems used in large AI labs.
- Prior work publishing at top-tier conferences or journals in machine learning, NLP, or human-centered AI evaluation.
Practical notes
- This role is based in the Bay Area.
- The engagement is full-time.
- Only the hours, travel, visa, or deadlines explicitly stated in the source material are applicable.