Sr GPU Infrastructure Software Engineer
Job description
About the role
CoreWeave is seeking a senior engineer to build and maintain the infrastructure powering large-scale AI workloads. You will focus on the systems that manage GPU clusters, ensuring reliability and high performance for complex computing tasks. In this capacity, you will own the design and execution of critical infrastructure components that directly influence the stability and throughput of AI training platforms. The role requires a proactive approach to solving intricate problems that emerge at the intersection of hardware and distributed systems. You will be responsible for translating operational demands into robust software solutions that scale efficiently across global data centers. Your work will involve close partnership with hardware specialists and application teams to align infrastructure capabilities with evolving computational requirements. Ultimately, you will drive the technical vision for the GPU infrastructure stack, ensuring it meets the stringent demands of next-generation AI workloads.
Key facts
What you'll do
- Architect and implement low-level software components that orchestrate GPU resources across distributed environments.
- Engineer high-throughput data paths to maximize the utilization of NVIDIA Hopper, Blackwell, Ada Lovelace, and Vera Rubin architectures.
- Construct fleet management systems that automate the provisioning, updating, and decommissioning of node pools with zero downtime.
- Develop observability frameworks that provide granular insights into hardware performance, thermal conditions, and power utilization metrics.
- Optimize scheduling logic to ensure efficient placement of AI workloads based on real-time cluster state and resource availability.
- Build resilient communication layers that maintain integrity and latency targets across high-speed interconnect fabrics in large data centers.
- Create diagnostic tools that accelerate root cause analysis for infrastructure failures impacting training and inference pipelines.
- Collaborate with security teams to enforce compliance and isolation policies across multi-tenant GPU clusters.
- Streamline the integration of bare-metal servers with managed Kubernetes platforms to offer a unified control plane for all compute resources.
- Implement automation for firmware and driver rollouts, ensuring compatibility and stability across heterogeneous hardware generations.
- Design interfaces that enable application developers to efficiently leverage massive parallelism without deep infrastructure expertise.
- Monitor system-level telemetry to predict and prevent hardware failures before they impact production workloads.
- Partner with product teams to define infrastructure roadmaps that support emerging AI model architectures and training methodologies.
- Document operational procedures and architectural decisions to maintain institutional knowledge and facilitate smooth on-call transitions.
Requirements
- Demonstrate extensive experience in infrastructure software engineering with a portfolio of systems that operate at scale in production environments.
- Show a deep understanding of GPU architectures, including memory hierarchies, compute units, and interconnect technologies specific to modern NVIDIA hardware.
- Exhibit proficiency in building distributed systems that handle concurrency, fault tolerance, and state management under extreme loads.
- Prove your ability to write high-performance code in systems programming languages, ensuring low latency and efficient resource utilization.
- Provide evidence of managing complex software deployments in cloud-native environments, including containerization and orchestration platforms.
- Highlight your experience with automated lifecycle management for large-scale server fleets, including imaging, provisioning, and decommissioning processes.
- Illustrate a history of collaborating effectively with cross-functional teams to resolve critical infrastructure challenges under tight deadlines.
- Confirm your capability to obtain and maintain government clearances or handle sensitive data as required by the specific role and client needs.
Nice to have
- Experience contributing to open-source projects related to container orchestration, scheduling, or virtualization.
- Familiarity with high-performance networking protocols and switch configurations used in data center environments.
- Knowledge of machine learning workflows and common frameworks used for model training and inference.
Practical notes
Full-time
Candidates must be eligible to work in the United States without sponsorship for this role.
This position may require on-site presence in Sunnyvale, CA, or Bellevue, WA, depending on team needs.
Travel requirements may include domestic trips to customer sites or regional data centers.
Applicants must meet all criteria outlined in the Requirements section to be considered for the role.
The posting will remain open until the position is filled, and early application is strongly encouraged.
Employment is contingent upon successful completion of background checks and verification of provided information.
Relocation assistance is not provided for this position.