AI Infrastructure Systems Engineer
Job description
About the role
The position focuses on replacing manual operations with software-driven automation at massive scale. Success means the platform deploys, monitors, diagnoses, optimizes, and heals GPU infrastructure largely by itself.
Software engineers turn product ideas into working code. Engineers work in small teams, review each other's work, and ship in small batches. Most teams follow agile practices such as sprints and daily standups. Engineers also write tests, fix bugs, and improve performance. The field values clear communication as much as technical skill. Engineers spend part of every week on planning, code review, and debugging, not just writing new code. The ability to explain a technical decision in plain words separates strong engineers from the rest.
Key facts
Requirements
3+ years building distributed systems, infrastructure platforms, or large-scale backend software to ensure foundational experience.
Strong software engineering skills in Python, Go, or Rust to implement reliable infrastructure components.
Experience building platforms, automation systems, or developer infrastructure to reduce operational burden.
Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies to manage complex environments.
Strong systems thinking with the ability to understand problems across hardware and software to design coherent solutions.
A passion for solving complex infrastructure challenges through software instead of manual processes.
An automation-first mindset where repeated tasks trigger systematized solutions that eliminate manual work.
Nice to have
GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch expertise to accelerate model training workflows.
InfiniBand or RoCE networking knowledge to optimize high-throughput communication.
Bare-metal provisioning and lifecycle management experience for direct hardware control.
Large-scale AI training or inference clusters operations to ensure reliability under load.
Hardware health monitoring and predictive failure detection systems to prevent outages.
Distributed storage systems understanding for scalable data throughput.
AI agents and autonomous infrastructure operations to enable self-healing platforms.
Practical notes
This role operates in a research-driven environment co-designing software, hardware, algorithms, and models. The position requires onsite presence and participation on the San Francisco team. Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
Good to know
The role centers on software-defined infrastructure for AI compute at scale.
Core tools include automation frameworks, fleet intelligence platforms, and infrastructure-as-software approaches.
Engineers in this role solve hard operational problems through code rather than manual intervention.
Success depends on an automation-first mindset and strong systems thinking across hardware and software layers.
Questions to ask
Worth asking in any interview: how the team measures success, who the role works with daily, what the onboarding looks like, and what the company is trying to achieve this year. Asking what past hires did well is a strong final question. Keep the list short and pick the questions that matter most to you.
Career growth
Engineering careers usually progress from individual contributor to senior, staff, and principal levels. Some engineers move into management and lead teams of five to twenty people. Others stay on the technical track. Growth follows demonstrated impact, not tenure alone. A typical engineering ladder has clear levels with defined expectations for scope, quality, and mentorship. Moving up usually requires owning outcomes end to end rather than completing assigned tickets.