Senior Research Engineer
Job description
About the role
This position targets an experienced machine learning engineer or a software engineer shifting toward AI safety research within a non-profit institute. The role focuses on conducting safety research, building experimental prototypes, and scaling solutions for frontier model risks. Collaboration with a multidisciplinary team is required to produce high-impact research and tools. You will own the design and execution of safety experiments that directly quantify risks posed by frontier models and translate findings into concrete mitigation strategies for external partners. A core part of your work will involve providing rigorous pre- and post-release adversarial evaluations for models such as Claude 4 Opus, ChatGPT Agent, and GPT-5 to materially strengthen deployment safety. You will systematically conduct red-teaming and develop novel attacks to surface critical vulnerabilities and guide defense strategies for high-stakes deployments in real-world settings. Another key responsibility is to build and study robustness mechanisms using scaling laws to determine whether robustness reliably improves with scale and to evaluate superhuman AI capabilities for latent security weaknesses. You will also perform mechanistic interpretability work, including Sparse Autoencoders, to probe deception in models such as AmongUs and to understand learned planning behaviors in environments like SokoBan. Finally, you will support the growth of the AI safety field by publishing findings at NeurIPS, ICML, and ICLR and by actively contributing to roadmaps and open technical reports that shape the field.
Key facts
What you'll do
Design and execute safety experiments on frontier models to quantify risks and inform mitigation strategies for external partners.
Provide pre- and post-release adversarial evaluations for models such as Claude 4 Opus, ChatGPT Agent, and GPT-5 to strengthen deployment safety.
Conduct red-teaming and develop novel attacks to surface vulnerabilities and guide defense strategies for high-stakes deployments.
Build and study robustness mechanisms using scaling laws to demonstrate whether robustness improves with scale and evaluate superhuman AI capabilities for security weaknesses.
Perform mechanistic interpretability work, including Sparse Autoencoders, to probe deception in models such as AmongUs and understand learned planning in SokoBan.
Support the growth of the AI safety field by publishing findings at NeurIPS, ICML, and ICLR and contributing to roadmaps and open technical reports.
Mentor engineers and scientists to elevate engineering practices and accelerate high-quality research outcomes across teams.
Work closely with multidisciplinary collaborators to prototype and iterate on safety tools that address emergent risks from powerful AI systems.
Translate empirical results from safety evaluations into actionable recommendations for model developers and deployment teams.
Maintain and extend Python-based toolchains and experiment infrastructure to support scalable, reproducible safety research on large models.
Requirements
Hold a Bachelor's degree in a relevant technical field.
Demonstrate significant software engineering experience through prior work history and open-source contributions.
Write fluent Python code for research engineering tasks and production-grade systems.
Work effectively in a fast-paced, high-impact research environment with strong results orientation.
Mentor peers and guide engineering skill development for scientists and engineers.
Bring deep expertise in at least one core competency aligned with current research focus areas.
Possess US work authorization or citizenship as required by law for this position.
Commit to defined onsite days in the Berkeley office to enable close collaboration and team cohesion.
Be prepared for occasional travel to research events and engagements that advance the mission.
Nice to have
Experience with large-scale training, fine-tuning, and post-training of transformer models.
Track record of publications at premier machine learning venues.
Hands-on work with GPU clusters and infrastructure for large model experimentation.
Practical notes
US work authorization or citizenship is required for this role.
The Berkeley office requires defined onsite days as part of team collaboration.
Some travel may be required for research events and engagements.
Typical interview steps
Data interviews commonly include a SQL or coding exercise, a statistics question, and a case study. Candidates may be asked to design a metric, interpret an experiment, or build a small model. Some companies give a take-home analysis. Expect questions about past projects and the business impact of your work. Interviewers often evaluate how you communicate uncertainty and business impact, not only the math. Bringing a clean write-up of a past analysis to the interview is well received.
Good to know
AI safety research relies on empirical experimentation on real-world models and infrastructure.
The role uses Python-based toolchains and GPU-heavy workflows common in modern ML research.
Collaboration across subfields such as robustness, interpretability, and deception is typical.
Publications and open technical reports drive impact in this field.
Mentoring peers helps scale engineering best practices and research velocity.