Staff+ Software Engineer, Safeguards Evals
Job description
About the role
In this position, you will be responsible for the design and implementation of systems and methodologies aimed at assessing model safety and identifying potential misuse. This role serves as a critical link between applied machine learning research and engineering, ensuring that our evaluation processes accurately reflect real-world threats and challenges.
Key facts
What you'll do
- Design and conduct experiments to enhance the quality of evaluations, focusing on aspects such as data generation, behavior simulation, and validation of grading systems.
- Develop and oversee evaluation harnesses for systems that investigate agentic behaviors, establishing metrics and test cases for agents that operate over extended timeframes.
- Generate high-quality datasets that reflect real-world harm scenarios, including but not limited to cyber threats, biological weapons, and influence operations.
- Seamlessly integrate research insights into production workflows for model training, updates to agents, and the gating of releases.
- Create tools that empower policy experts to independently manage and refine evaluations without needing extensive engineering support.
- Collaborate with cross-functional teams to ensure that evaluation methodologies align with organizational goals and safety standards.
- Analyze evaluation results to derive actionable insights that can inform future development and safety measures.
- Stay updated on the latest advancements in machine learning and safety protocols to continuously improve evaluation processes.
Requirements
- A minimum of 8 years of experience in software engineering or machine learning within an industry setting.
- A Bachelor's degree or an equivalent combination of education and relevant experience in a related field.
- Proven experience in constructing and maintaining data pipelines that support large-scale operations.
- Familiarity with the capabilities and limitations of large language models (LLMs), particularly in relation to agentic systems, tool utilization, and complex reasoning tasks.
- Strong data analysis skills, enabling the extraction of meaningful insights from extensive datasets.
- Demonstrated ability to shift between research-oriented prototyping and the development of production-ready code.
- A knack for transforming vague challenges into structured, testable experiments that yield valuable results.
Nice to have
- Experience in developing frameworks for evaluating LLMs or agents, including automated grading systems.
- A background in trust and safety, content moderation, or detection of abusive behaviors.
- Familiarity with red teaming, adversarial testing, or research focused on jailbreak methods.
- Knowledge of synthetic data generation techniques or data augmentation strategies.
- Experience working with distributed systems or in environments that require large-scale data processing.
- Skills in prompt engineering or the development of applications powered by LLMs.
Skills & tools
- Proficient in LLM evaluation frameworks and methodologies.
- Experienced in designing data pipeline architectures that support robust data processing.
- Knowledgeable in large-scale data processing techniques and distributed systems.
- Skilled in the design of agentic systems that require complex interactions and evaluations.
Practical notes
- Our company follows a hybrid work model based on location, requiring employees to spend at least 25% of their time in the office.
- We offer visa sponsorship and have immigration counsel available to assist candidates throughout the process.
- We encourage individuals to apply even if they do not meet every qualification listed in the job description.
- To avoid recruitment scams, please ensure that all communications come from an @anthropic.com email address.
At Anthropic, we are committed to fostering a diverse and inclusive work environment. We believe that a variety of perspectives enhances our ability to innovate and address complex challenges in the field of AI safety. If you are passionate about making a difference and meet the qualifications outlined above, we would love to hear from you. Join us in our mission to create safe and beneficial AI systems for everyone.
About the company
Anthropic is an AI safety company that builds reliable, interpretable, and steerable AI systems. Founded in 2021 by Dario and Daniela Amodei, former VP of Research at OpenAI, Anthropic created the Claude family of AI assistants. The company has raised over $13 billion from investors including Google, Salesforce, and Amazon.