Principal Engineer - Perf and Benchmarking
Job description
About the role
CoreWeave is seeking a technical leader to drive performance analysis and benchmarking initiatives across our AI infrastructure. You will define methodologies to measure and optimize the performance of our GPU compute clusters, ensuring our systems meet the demands of large-scale AI training and inference. In this capacity, you will act as the central authority on performance characterization, establishing the standard practices that inform both engineering and product decisions. You will own the end-to-end lifecycle of performance investigation, from hypothesis generation through data collection to final recommendation. The role requires deep collaboration with hardware, software, and operations teams to ensure alignment between low-level optimizations and business objectives. You will be responsible for maintaining the integrity and consistency of our performance data as it scales across global data centers. This position is critical for validating that our infrastructure delivers on its promises to customers and internal stakeholders. Your work will directly influence the architectural roadmap and the competitive positioning of our AI cloud platform.
Key facts
What you'll do
- Design and execute comprehensive benchmarking suites for GPU-accelerated workloads, establishing repeatable and statistically sound methodologies.
- Analyze system performance data to identify bottlenecks in training and inference pipelines, isolating issues at the kernel, framework, or hardware level.
- Collaborate with engineering teams to optimize software and hardware configurations for peak efficiency, balancing tradeoffs between throughput, latency, and resource utilization.
- Develop internal tools and automated frameworks to track performance metrics across our global data centers, ensuring consistent telemetry and traceability.
- Translate technical performance data into actionable insights for product development and capacity planning, presenting complex findings to diverse audiences.
- Define and drive the adoption of internal performance standards aligned with industry benchmarks such as MLPerf, adapting them to our specific infrastructure and use cases.
- Partner with product managers to establish performance goals and service-level objectives, creating benchmarks that validate success at scale.
- Investigate new GPU architectures and emerging compute technologies to evaluate their potential impact on CoreWeave's workload performance.
- Lead post-mortem analyses for significant performance regressions, coordinating cross-functional efforts to implement durable fixes.
- Mentor senior engineering staff in performance analysis techniques, fostering a data-driven culture of optimization and accountability.
- Design experiments that isolate variables in large-scale distributed training jobs, ensuring results are reliable and reproducible.
- Maintain a living knowledge base of performance best practices, anti-patterns, and tuning guidelines for internal consumption.
- Interface with vendors and partners to validate hardware specifications and drive alignment on performance-critical features.
- Represent CoreWeave in external technical discussions, contributing to industry dialogues on AI infrastructure performance and standards.
Requirements
- Extensive experience in performance engineering or systems architecture within high-performance computing or cloud environments, with a track record of delivering measurable improvements.
- Deep understanding of GPU architectures, specifically NVIDIA Hopper, Blackwell, and Vera Rubin series, including their memory hierarchies, instruction sets, and execution models.
- Proficiency in benchmarking methodologies for AI training and inference, including familiarity with MLPerf standards and the ability to adapt them to internal needs.
- Strong background in Linux systems, networking, and distributed computing, with the ability to diagnose issues across complex, multi-node environments.
- Ability to lead technical projects and mentor senior engineering staff, demonstrating clear communication and influence without direct authority.
- Experience contributing to or maintaining performance testing frameworks that are used to validate production infrastructure at scale.
- A history of working with large datasets and high-throughput I/O systems, understanding how storage performance interacts with compute.
- Commitment to maintaining rigorous data collection and analysis practices, ensuring that performance decisions are based on evidence rather than intuition.
Nice to have
- Experience with Kubernetes-based infrastructure and container orchestration at scale, particularly in multi-tenant AI workloads.
- Familiarity with low-level profiling tools and runtime acceleration techniques, including CUDA, Nsight, and related ecosystems.
- Contribution to open-source performance benchmarking projects, demonstrating a commitment to community standards and transparency.
Practical notes
CoreWeave is an AI-native cloud provider focused on high-performance infrastructure. Applicants must be authorized to work in the United States. This role may require travel between office locations as needed for team collaboration and project reviews. The engagement is full-time, and successful candidates are expected to align with CoreWeave's standard working hours to support global collaboration.