Senior Software Engineer, Cluster Orchestration
Job description
About the role
You will design and maintain the systems that manage large-scale GPU clusters for AI workloads. This position focuses on building the control plane infrastructure that ensures high performance and reliability across our data centers. You will be responsible for the architecture and implementation of the foundational components that allow our Mission Control platform to operate at scale. The role requires a deep understanding of how distributed systems interact with physical hardware in demanding AI environments. You will own the critical pathways that translate computational demand into precise resource allocation and scheduling decisions. This position is centered on creating robust systems that handle the lifecycle of thousands of nodes and GPUs efficiently. You will work closely with infrastructure teams to ensure that the software stack aligns with the operational needs of our data center fleet. The work you do will directly impact the reliability and performance of AI training and inference workloads running on our platform.
Key facts
What you'll do
- Architect and implement software for large-scale fleet and node lifecycle management, ensuring automation and consistency across all data center locations.
- Design and build orchestration tools that efficiently manage massive GPU compute resources, optimizing for utilization and performance.
- Enhance the reliability, scalability, and visibility of our AI infrastructure through the development of critical control plane services.
- Partner with cross-functional engineering teams to analyze and optimize cluster performance, driving improvements in automated operations and deployment pipelines.
- Develop and maintain systems that provide deep observability into the health and status of GPU clusters, enabling rapid troubleshooting and insight.
- Implement infrastructure automation frameworks that reduce manual intervention and increase the resilience of our Mission Control platform.
- Collaborate on the definition and enforcement of infrastructure standards for node provisioning, configuration, and decommissioning across the fleet.
- Contribute to the evolution of our software stack by researching and integrating new technologies relevant to distributed systems and hardware abstraction.
- Ensure that the systems you build adhere to strict security and audit visibility frameworks, maintaining the integrity of our infrastructure.
- Drive the execution of projects that scale our cluster management capabilities to meet the rapidly growing demands of AI workloads.
- Take ownership of complex operational problems, diagnosing issues at scale and implementing long-term solutions that prevent recurrence.
- Work with product and infrastructure teams to translate high-level requirements into robust technical designs for cluster orchestration features.
Requirements
- You possess professional experience in distributed systems or cluster orchestration, with a proven track record of building and operating complex software.
- You demonstrate proficiency in managing Kubernetes environments at scale, including deep knowledge of its APIs and extension mechanisms.
- You have substantial experience with infrastructure automation and fleet management, showing the ability to manage large groups of machines efficiently.
- You are able to work on-site in Sunnyvale or Bellevue, as this role requires physical presence in our primary locations.
- You have a strong understanding of Linux operating systems, networking, and storage systems as they relate to high-performance computing environments.
- You bring experience with writing scripts and programs in systems-level languages to automate and manage infrastructure tasks.
- You have a history of troubleshooting intricate production issues, analyzing logs, and using monitoring data to drive decisions.
- You are comfortable working in a fast-paced environment where priorities can shift based on the needs of the business and our infrastructure.
Nice to have
- You have a background in high-performance computing or large-scale AI infrastructure, understanding the unique challenges these domains present.
- You have hands-on experience with Go or similar systems-level programming languages, allowing you to contribute to the core infrastructure codebase.
- You are familiar with observability tools and secure audit visibility frameworks, enhancing our ability to monitor and secure our systems.
Practical notes
- CoreWeave is a provider of cloud AI infrastructure specializing in GPU compute.
- The role involves supporting the Mission Control platform.
- This is a full-time position based in the United States.
- Candidates must be eligible to work in the United States without requiring work authorization sponsorship for this position.
- The on-site work location requirements for this role are either Sunnyvale, California, or Bellevue, Washington.