Software Engineer, Agentic AI Infrastructure
Job description
About the role
You will design and build the distributed backbone that powers Anrok's AI agent capabilities, owning the core infrastructure for memory, state management, and execution. In this role, you will make critical architecture decisions that define how agentic systems behave in production across global tax jurisdictions. You will partner with product and tax data teams to turn complex compliance workflows into reliable, scalable primitives for AI agents. You will own the evaluation and observability tooling that ensures agents are dependable, observable, and safe to release. This is foundational work where your contributions will directly shape the platform that other products are built upon. You will drive engineering standards for reliability, testing, and operational health across the team. If you thrive in ambiguous, high-impact environments and care deeply about system design, you will own the problems that make agentic AI infrastructure trustworthy.
Key facts
What you'll do
Design and build the distributed, fault-tolerant infrastructure that powers Anrok's AI agent capabilities - covering memory, state management, and execution for autonomous systems.
Own technical architecture decisions for new agent infrastructure services end-to-end, from initial design through production deployment and iteration.
Build the evaluation, observability, and reliability tooling that lets the rest of the team ship agents confidently on top of your infrastructure.
Collaborate closely with product engineers and our tax data team to translate real compliance workflows into reliable, scalable agent primitives.
Drive engineering best practices across the
team: code quality, testing discipline, system design, and operational health.
Identify and fix performance, reliability, and scalability bottlenecks as our agent systems grow.
Help grow the team - mentoring engineers, contributing to technical culture, and raising the bar on how we build.
Evaluate candidate agent architectures and contribute to hiring for roles that strengthen the long-term reliability of the platform.
Prototype new agent capabilities to validate infrastructure assumptions and surface edge cases before platform rollout.
Instrument systems to provide deep insight into agent behavior, resource usage, and failure modes in production.
Balance tradeoffs between flexibility, latency, and correctness when designing agent workflows that interact with tax data.
Work with security and compliance partners to ensure agent infrastructure meets enterprise readiness standards.
Support on-call responsibilities and lead incident response for critical agent infrastructure when needed.
Champion open-source-friendly patterns where appropriate to enable integration and extensibility downstream.
Requirements
You have 3+ years of experience building and operating distributed, production-grade systems that handle real-world scale.
You have designed services that are scalable, resilient, and fault-tolerant - and you have done it more than once in your career.
You have built and deployed AI agents or LLM-powered systems in production, and you understand what makes them reliable and maintainable at scale.
You are familiar with the modern AI stack - including prompt engineering, eval frameworks, RAG pipelines, tool calling, and agent orchestration - and you have a point of view on what the infrastructure beneath it should look like.
You think carefully about architecture and tradeoffs, and you are comfortable leading technical decisions without a complete playbook.
You have experience with the full software development lifecycle: design, code review, testing, deployment, and on-call ownership.
You are energized by ambiguous problems and can break them down into actionable steps that lead to shipped solutions.
You are curious and motivated to learn new tools, and you have built something with AI recently - whether an application, workflow, or creative solution - that you can walk us through.
You communicate clearly and enjoy collaborating with engineers who share high standards for system design and execution.
You are comfortable working in a fast-paced environment where priorities can shift as product and compliance needs evolve.
Nice to have
Experience with tax, finance, or compliance systems is a plus given our domain complexity.
Hands-on work with Kubernetes, service meshes, and infrastructure-as-code in large-scale deployments.
Contributions to or familiarity with open-source agent frameworks or observability tools.
Background in building eval frameworks for agent behaviors or automated testing for LLM workflows.
Experience with compliance workflows in digital commerce or billing systems.
Practical notes
This role is full-time based in San Francisco.
Hybrid work policy requires in-person collaboration three days per week at our San Francisco hub.
Employment eligibility to work in the United States is required.
No sponsorship is available at this time.
Travel is not required for this role.
Deadlines are not specified; applications will be reviewed on a rolling basis.