Staff Software Engineer, DC Infrastructure
Job description
About the role
You will architect and implement the software stack that directly manages tens of thousands of GPUs and the critical infrastructure that powers them. This role centers on owning the entire lifecycle of diagnostic, observability, and automation tools that keep our data centers online and performing at peak efficiency. You will be responsible for developing the systems that detect, diagnose, and remediate faults before they impact our customers' AI workloads. Your work will directly influence the reliability and performance of the hardware that trains and runs the world's most advanced models. You will collaborate closely with data center operations to build intelligent agents that manage the critical environment, from power to cooling. Ultimately, you will ensure that our fleet delivers maximum uptime and performance, driving customer success and advancing our mission.
Key facts
What you'll do
Develop and implement deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems.
Develop troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
Develop automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware.
In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.
Develop tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.
Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
Develop automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems.
Contribute to the design and implementation of scalable, resilient software systems that operate at the edge of our data centers.
Collaborate with cross-functional teams to define requirements, prioritize backlogs, and deliver solutions that meet strict uptime and performance SLAs.
Write clean, maintainable, and well-documented code that adheres to our engineering standards and best practices.
Participate in on-call rotations to respond to critical incidents and drive resolution from development through production.
Evaluate and integrate new technologies to solve complex infrastructure problems and improve operational efficiency.
Work closely with hardware engineers to bridge the gap between firmware, low-level drivers, and high-level orchestration.
Continuously iterate on existing tools to improve performance, reliability, and developer experience.
Requirements
Software engineering experience building and operating large-scale distributed systems.
The ability to identify a problem, rapidly develop a scalable solution and ship it into production environments.
Experience working with infrastructure components and understanding of how software impacts physical systems.
Strength in at least one programming language
Go, Python, Java, Rust.
Strong analytical and problem-solving skills to diagnose complex issues in distributed GPU compute environments.
Excellent communication and collaboration skills to work effectively with cross-functional teams including hardware and operations.
Ability to work independently and within a team in a fast-paced, ambiguous environment with a sense of urgency.
Experience with infrastructure as code principles and tools is required.
A proven track record of writing reliable, performant, and secure software.
Nice to have
Experience with Temporal and Kubernetes.
Experience working directly with hardware vendors.
Background in large-scale GPU fleet operations or hyperscale data center environments.
Practical notes
This role may require occasional travel to US worksites.
Candidates must be authorized to work in the United States without sponsorship now or at any time during the employment term.
- Industry competitive pay
- Restricted Stock Units in a fast growing, well-funded technology company
- Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
- Employer contributions to HSA accounts
- Paid Parental Leave
- Paid life insurance, short-term and long-term disability
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Generous paid time off and holiday schedule
- Cell phone reimbursement
- Tuition reimbursement
- Subscription to the Calm app
- MetLife Legal
- Company paid commuter benefit; $300 per month
Compensation Range
Compensation is competitive and aligned with market rates, with a range designed to reflect the value of the role and the expertise required.