Machine Learning Scientist
Job description
About the role
Arena Intelligence is seeking a variety of Machine Learning Scientist to help advance how we evaluate and understand AI models. You will own the design and analysis of experiments that uncover what makes models useful, trustworthy, and capable through human preference signals at scale. Your work will directly shape the scientific foundations of understanding AI across diverse domains such as agentic coding, creative generation, and professional productivity. You will move beyond static leaderboards to decompose what human experience reveals about model behavior in real workflows. This role is deeply interdisciplinary, requiring you to translate research insights into practical evaluation tools that guide both public leaderboard methodology and internal product decisions. If you are excited by open-ended questions, rigorous evaluation, and research grounded in real-world impact, you will find a meaningful home within our mission-driven team. You will contribute to building the foundation for everyone to understand, shape, and benefit from AI in a transparent and human-centered way.
Key facts
What you'll do
- Design and conduct experiments to evaluate AI model behavior across reasoning, style, robustness, and user preference dimensions using novel protocols.
- Develop new metrics, methodologies, and evaluation frameworks that move beyond traditional benchmarks toward real-world performance indicators.
- Analyze large-scale human voting and interaction data to uncover systematic insights into model performance and latent user preferences.
- Collaborate with engineers to implement and scale research findings into production systems that support reliable evaluation at Arena Intelligence.
- Prototype and test research ideas rapidly, balancing scientific rigor with iteration speed to deliver actionable evaluation insights.
- Author internal reports and external publications that contribute to the broader machine learning and AI evaluation research community.
- Partner with model providers to shape evaluation questions, define responsible testing practices, and support transparent model assessment.
- Contribute to the scientific integrity, transparency, and reproducibility of the Arena Intelligence leaderboard and associated evaluation tools.
- Iterate on dataset designs, large-batch training pipelines, and ablation studies to ensure that evaluation protocols scale effectively.
- Work closely with product teams to align modeling goals with user needs and translate evaluation findings into product improvements.
- Engage with the research community by sharing open datasets, methodologies, and insights that advance evaluation best practices across the field.
- Explore creative generation, agentic coding, and professional productivity workflows to decompose what human experience reveals about AI capabilities.
- Maintain fluency in the full experimental stack from dataset curation through large-scale evaluation and production integration.
- Drive curiosity and craftsmanship in evaluation work by questioning assumptions and refining methods based on empirical evidence.
Requirements
- Hold a PhD or equivalent research experience in Machine Learning, Natural Language Processing, Statistics, or a closely related quantitative field.
- Possess hands-on experience training large-scale models, including reward models, preference models, and fine-tuning LLMs using methods such as RLHF, DPO, and contrastive learning.
- Demonstrate a strong foundation in machine learning and statistics, with a track record of designing novel training objectives, evaluation schemes, or statistical frameworks to improve model reliability and alignment.
- Show proven ability to design and analyze experiments with statistical rigor while maintaining clarity in interpretation and communication.
- Have experience publishing research or contributing to open-source projects within machine learning, natural language processing, or AI evaluation.
- Exhibit comfort working with real-world usage data and designing metrics that extend beyond standard benchmarks to capture practical performance.
- Display the ability to translate research questions into practical systems and collaborate effectively with engineering and product teams.
- Demonstrate a passion for open science, reproducibility, and community-driven research that prioritizes transparency and shared progress.
- Thrive in a fast-paced, interdisciplinary environment where thoughtful curiosity and rigorous analysis are essential.
- Bring strong written and verbal communication skills to convey complex technical concepts to both technical and non-technical stakeholders.
- Be comfortable operating with ownership and initiative in a mission-driven setting where impact depends on methodological rigor and real-world relevance.
- Commit to maintaining the highest standards of scientific integrity in all evaluation work conducted for Arena Intelligence.
- Be prepared to work from the Bay Area location to ensure close collaboration with the research, engineering, and product teams.
Practical notes
- Hours: Full-time commitment is required for this role.
- Travel: No specific travel requirements are stated in the source.
- Visa: Candidates must be eligible to work in the Bay Area location.
- Deadlines: No explicit application deadline is provided in the source; interested candidates should apply as soon as possible.