Senior Platform Engineer II, Compute Services
Job description
About the role
You will design, build, and maintain the critical infrastructure that powers high-performance AI workloads. This Senior Platform Engineer II role within the Compute Services team is centered on architecting and ensuring the reliability of scalable compute systems specifically for large-scale GPU deployments and complex machine learning environments. You are responsible for the end-to-end lifecycle of the platform, from initial architecture and deployment to ongoing optimization and governance. The position demands a deep commitment to infrastructure as code, automation, and operational excellence to deliver the stable, high-throughput foundation required for AI innovation. Your work will directly determine the efficiency, scalability, and reliability of AI training and inference pipelines for our customers.
Location: USA
Engagement: Full-time
Team: Compute Services
What you'll do
- Architect and manage the compute infrastructure to support the most demanding AI training and inference tasks at scale.
- Develop and maintain automation for comprehensive fleet management and node lifecycle control to guarantee maximum availability and resilience.
- Collaborate closely with internal engineering teams to research, test, and optimize performance specifically for NVIDIA GPU cluster deployments.
- Drive improvements in the reliability, observability, and overall health of our globally distributed compute environment.
- Design, implement, and manage solutions encompassing both bare metal and virtualized compute resources to meet diverse workload requirements.
- Design and deploy scalable network configurations and security policies to robustly support data-intensive workloads and ensure efficient data flow.
- Troubleshoot complex infrastructure issues using advanced diagnostic methodologies and perform in-depth root cause analysis.
- Evaluate and integrate emerging technologies and tools to continuously enhance the capabilities, performance, and efficiency of the compute platform.
- Partner with security and compliance teams to ensure all infrastructure implementations adhere to strict governance, compliance standards, and security best practices.
- Lead initiatives to streamline deployment pipelines, automate manual processes, and significantly reduce operational overhead and toil.
- Monitor, analyze, and report on key system performance metrics to drive data-informed decisions for continuous improvement and capacity planning.
- Facilitate knowledge sharing sessions, code reviews, and technical documentation to elevate the technical proficiency and capabilities of the entire engineering team.
- Contribute actively to the definition, development, and enforcement of best practices for infrastructure management, operational runbooks, and disaster recovery procedures.
- Provide technical leadership and guidance in technical discussions related to long-term platform strategy, scalability planning, and architectural evolution.
Requirements
- Proven professional experience as a Platform Engineer or in architecting and managing distributed systems at significant scale.
- Deep technical expertise in Linux operating systems, kernel parameters, and infrastructure automation using tools such as Terraform, Ansible, or similar.
- Demonstrated ability to manage the full lifecycle of large-scale compute fleets and complex hardware infrastructure, including deployment, maintenance, and decommissioning.
- Proficiency in designing, building, and operating high-performance computing (HPC) environments or highly available cloud infrastructure architectures.
- A proven track record of successfully delivering critical, production-grade infrastructure components in fast-paced, high-availability environments.
- Strong foundational knowledge of core operating systems, networking protocols and implementation, and distributed storage systems.
- Experience with scripting and programming (e.g., Python, Go, Bash) to automate complex operational tasks and develop tooling.
- An unwavering commitment to maintaining the highest possible standards for system reliability, performance, security, and operational integrity.
Nice to have
- Production-level experience with Kubernetes and other container orchestration platforms for managing distributed workloads.
- Hands-on background in designing, deploying, and managing infrastructure specifically for NVIDIA GPU-based compute workloads.
- Practical familiarity with the integration of networking and storage subsystems tailored for demanding AI and machine learning workloads.
Skills & tools
- Linux
- Kubernetes
- Infrastructure Automation
- Distributed Systems
- GPU Compute Architecture
- Fleet Management
Practical notes
CoreWeave is an AI-native cloud provider. This role is based in one of our specified US office locations. Please submit your application through our official careers portal to be considered for this position.