Task Development Engineer
Job description
About the role
Metr is a nonprofit research organization that builds scientific methods for measuring AI capabilities, risks, and possible mitigations. The organization concentrates on dangers tied to AI research and development automation as well as misalignment problems. Metr has a track record of setting standards in catastrophic AI risk assessments across many independent evaluations. The Task Development Engineer will contribute to designing and refining the tasks used in these evaluations, shaping the practical work that AI systems are asked to perform as part of Metr assessment protocols.
Key facts
What you'll do
Design and develop evaluation tasks that test AI systems on specific capabilities and risk dimensions.
Create task specifications that align with Metr scientific methodology for AI safety assessments.
Collaborate with researchers to define the scope and structure of each evaluation task.
Iterate on task designs based on feedback from completed evaluations and pilot runs.
Ensure tasks are clear, measurable, and reproducible across different AI systems and settings.
Document task requirements, expected behaviors, evaluation criteria, and success metrics in thorough detail.
Work with the team to identify new task areas that fill gaps in current evaluation coverage.
Maintain and update existing task libraries to keep them current with evolving AI capabilities.
Coordinate with external partners who may provide access to AI systems for testing purposes.
Support the broader Metr mission by contributing to assessments that inform policymakers and civil society.
Review and refine task designs with the research team to improve evaluation quality over time.
Help establish best practices for task development within Metr and contribute to team knowledge sharing.
Requirements
Familiarity with AI safety concepts, including R&D automation risks and alignment challenges.
Experience designing or developing structured tasks, evaluations, or benchmarks of any kind.
Strong written communication skills for writing clear task specifications and documentation.
Ability to work independently and manage your own workflow as a contractor.
Comfort with iterating on task designs based on empirical results and team feedback.
Understanding of how AI systems behave differently across various settings and configurations.
Attention to detail when defining evaluation criteria and expected outcomes for each task.
Willingness to engage deeply with the scientific methods used in AI risk assessment.
Comfort working in a remote, asynchronous environment with minimal supervision and clear communication.
Genuine interest in AI safety and the societal implications of advanced AI systems.
Nice to have
Prior experience in AI research, AI safety, or AI governance organizations.
Background in software engineering, technical program management, or a related technical field.
Familiarity with frontier AI labs and their evaluation practices or safety commitments.
Experience working with nonprofit or research organizations on scientific assessment or evaluation projects.
Knowledge of AI incident investigation processes or familiarity with third-party assessment frameworks.
Interest in contributing to nonprofit research that shapes how AI risk is understood publicly.
Skills & tools
Proficiency in writing clear technical specifications and well-structured documentation for evaluation work.
Experience with task design frameworks or evaluation methodology in any domain.
Familiarity with AI model APIs and how to interact with AI systems programmatically.
Knowledge of Python or other scripting languages for task automation and data handling.
Understanding of scientific evaluation methods, reproducible research practices, and rigorous testing approaches.
Ability to work with version control and collaborative document editing tools.
Careful attention to precision and clarity when writing instructions or evaluation rubrics.
Curiosity about how different AI models respond to structured tasks and evaluation scenarios.
Practical notes
This is a contractor position, so the engagement is not full-time employment with benefits.
The role is fully remote, allowing you to work from any location without a fixed office.
As a contractor, you will set your own schedule within the expectations of the Metr team.
Metr is a nonprofit, so compensation structures may differ from for-profit technology companies.
The organization values ambitious and excellent individuals who are drawn to hard scientific problems.
Metr has been consulted by policymakers, civil society groups, frontierlabs, and governments on risk assessment.
Metr is generally referenced as the canonical third-party assessor for AI capability and risk evaluations.
The organization has conducted the first independent safety evaluations, the first loss-of-control evaluations, and the first agentic dangerous capability evaluations.