
Staff Software Engineer, Inference
Job description
About the role
We are seeking a Staff Software Engineer to own the software state machines that provision hardware, bring it into service, and manage its full lifecycle, turning racks of GPUs into running inference clusters without a human touching a runbook. The role works closely with the Research and Inference team, transforming today's ticket-based workflow into a future where a single API call stands up, scales, or tears down a cluster entirely automatically. The platform is manifest-driven, so teams declare the desired state of a cluster or host in terms of shape, topology, and software stack, and the system continuously reconciles reality to that manifest across every lifecycle stage. You will design the engines that manifest the schema, the engines that execute against it, and the workflows that carry hardware or a cluster from bare metal to a fully functioning AI cluster for training or inference. Success is measured by eliminating manual provisioning work rather than by documenting it better, and by providing a self-service API that feels like a product rather than a collection of scripts. A product mindset is essential: you have built internal platforms or APIs consumed by other engineering teams and care deeply about the developer experience of what you ship, embodying the principle of building it, owning it, and operating it in production.
Key facts
What you'll do
Design and implement the provisioning state machine that models the full lifecycle of a physical host, from discovery and inference bring-up through GPU driver and CUDA stack installation, health validation, and decommission or RMA, expressed as explicit, versioned states and transitions.
Build self-service APIs and a control plane that allow the inference team to request, scale, and tear down inference clusters with a single API call, removing the need for tickets or human intervention.
Automate self-healing by detecting degraded or failed nodes, draining them safely, triggering repair or replacement workflows, and automatically reintroducing healthy capacity into the shared pool.
Own end-to-end reliability of the provisioning pipeline through idempotency, retries, rollback mechanisms, and drift detection, ensuring the system behaves as predictably as any other production service.
Partner closely with the inference and ML platform teams to understand cluster shapes, topology, and interconnect requirements, and encode these as first-class abstractions in the platform.
Engineer infrastructure like software by applying strong typing, automated testing, rigorous code review, versioning, and CI/CD practices to infrastructure code, recognizing it as a product rather than a collection of Ansible playbooks.
Define and evolve declarative manifests that describe desired cluster and host states, and build the reconciliation engines that continuously steer real infrastructure toward that desired state.
Implement workflow engines that manage state transitions across large-scale hardware deployments, ensuring progress, consistency, and recoverability even in the face of partial failures.
Instrument and observe the full provisioning lifecycle so that operational signals expose bottlenecks, failure modes, and opportunities for automation.
Contribute to the architectural decisions that determine how racks of GPUs become running inference clusters, balancing tradeoffs between flexibility, performance, and operational simplicity.
Drive improvements in developer experience by reducing the time and steps required for researchers to access and manage inference capacity.
Act as an operational on-call owner for the provisioning platform, responding to incidents, diagnosing root causes, and guiding remediation.
Lead best practices for durable workflow orchestration, event-driven design, and infrastructure-as-software across the team.
Mentor engineers on modeling complex infrastructure as state machines and writing robust, testable platform code.
Requirements
Strong software engineering background in Go, Python, Rust, or similar, with a track record of writing and testing real software that runs in production.
Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent, to run long-lived, manifest-driven workflows that survive failures and can resume mid-execution.
Experience building software control planes or orchestration systems that model state and reconcile it over time, for example Kubernetes controllers or operators, custom reconciliation loops, or workflow engines.
Experience with event-driven systems, designing and building software around message queues, event streams, or pub/sub, rather than polling or cron-driven scripts.
A product mindset demonstrated by having built internal platforms or APIs that are consumed by other engineering teams, and a genuine concern for the developer experience of what you ship.
Exposure to bare-metal provisioning technologies such as PXE/iPXE, Redfish, IPMI, or BMC, and familiarity with networking fundamentals including VLANs, BGP, and fabric design, or experience with GPU and accelerator infrastructure.
Prior work with GPU cluster software stacks including NCCL, CUDA, and InfiniBand or RoCE.
Previous employment at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
Strong proficiency in systems programming using Rust or Go.
Practical notes
This is a full-time position based in San Francisco.
The role may require limited travel to data center sites as needed.
Candidates must be authorized to work in the United States without sponsorship for this role.