Agent Harness Engineer
Job description
About the role
You will own the core agent harness that transforms frontier models into long-horizon scientific reasoning systems for drug discovery. This role sits at the intersection of infrastructure, evaluation, and research, where your work directly determines how reliably Axiom's models can analyze complex biological data. You will design the scaffolding and tooling that turns raw model outputs into trustworthy, repeatable scientific analysis. Your responsibilities will span building the data backbone, sandboxed execution environments, and the eval systems that validate model behavior against real-world pharmaceutical requirements. You will encode the instincts of domain experts into measurable rubrics and golden sets, turning subjective judgment into reliable signals. You will ensure every agent run is observable, replayable, and debuggable across long-running trajectories and fragile environments. You will also own the guardrails, retry logic, and budget controls that keep agentic workflows aligned with safety and cost constraints. Finally, you will support ML research by providing robust environments and instrumentation for RL on agentic tasks, closing the loop between evaluation and training.
Key facts
What you'll do
- Own the agent harness, including the scaffolding, tooling, and infrastructure that convert frontier models into reliable agents for long-horizon scientific analysis.
- Build and maintain the data backbone that moves everything to the right place, including pipelines, storage, runtime context, agent trajectories, evaluation results, and training data.
- Create sandboxed execution environments that spin up and tear down instantly, while remaining reproducible and deterministic enough to support rigorous evaluation.
- Design and operate eval systems, covering offline suites, test cases derived from production traces, LLM-as-judge pipelines, and regression gates.
- Collaborate closely with domain experts to encode their taste and judgment into rubrics, golden sets, and review workflows, converting "I know it when I see it" into measurable standards.
- Build the tools and interfaces agents need to perform better work, while actively tracking and adopting state-of-the-art methods, protocols, and patterns from the broader AI community.
- Make every agent run fully observable and replayable, tracing model calls, tool calls, and state transitions, and building debugging tooling to interpret those traces.
- Engineer context for memory, compaction, retrieval, and recovery so that long-horizon agent runs remain coherent across hours of execution and unexpected crashes.
- Own the control loop of the system, including retries, budget caps, stop conditions, output verification, permissions, and safety guardrails.
- Support Axiom's ML research efforts by providing robust environments, reward instrumentation, and rollout infrastructure for reinforcement learning on agentic tasks.
- Maintain strong software engineering practices, with infrastructure, platform, data, or devtools depth, and produce production systems you are genuinely proud of.
- Instinctively ask how you would know if a system is working, and build measurement alongside every feature through evaluation and observability.
- Read agent failure traces the way other engineers read stack traces, diagnosing issues quickly and precisely from complex telemetry.
- Care deeply about reliability, because environments that fail silently can poison evaluations and training data in subtle, damaging ways.
- Surface the most important questions about what agents actually need to do better work, and drive clarity between product, research, and domain teams.
- Operate comfortably in a highly technical and sophisticated environment, keeping up with cutting edge developments in AI and life sciences.
- Thrive in a discipline with no playbook, because Axiom is inventing the role and the systems you will build in parallel.
- Demonstrate curiosity about how things work, with an engineering and tinkering mindset that excels at scavenging and integrating the state of the art.
- Show a passion for learning what "good" looks like from deep domain experts and translating their standards into robust systems and tests.
Requirements
- Can tackle deep technical challenges and own end-to-end outcomes by shipping simple, clean, and maintainable code.
- Possesses high ownership, taking responsibility for results across the full stack without waiting for a complete specification to begin moving.
- Allergic to unnecessary complexity, consistently reaching for the simplest system that works and keeping it that way as scale increases.
- Functions as a strong software engineer first, with demonstrated depth in infrastructure, platform, data, or devtools and a track record of production systems you are proud of.
- Instinctively measures their work, asking "how would we know if this is working?" and building evaluation and instrumentation into every feature.
- Reads agent failure traces with the same instinct and rigor that other engineers apply to stack traces.
- Understands that unreliable environments silently corrupt evaluations and training data, and designs systems to prevent that.
- Has a knack for clarifying what agents actually need to do, surfacing critical questions that enable better tool design and workflows.
- Is sharp and confident, able to keep up with a highly technical and sophisticated peer group while remaining comfortable getting in over their head and figuring it out.
- Thrives in a role with no established playbook, comfortable inventing the right approach as the problem space evolves.
- Is deeply curious about engineering and tinkering, capable of scavenging and repurposing cutting-edge ideas and tools.
- Is passionate about learning from domain experts in biology and chemistry, and turning qualitative notions of quality into concrete systems and metrics.
Nice to have
- Built bespoke evaluation, monitoring, and reinforcement learning environments with observable debuggability.
- Experience contributing to or building tooling around agentic workflows in scientific or pharmaceutical contexts.
Practical notes
This role is full time based in San Francisco at our global headquarters. The position requires local authorization to work in the United States. Candidates must be eligible to work without sponsorship in this role. No visa sponsorship is available for this position. Travel is not required for this role.