Staff Applied Research Engineer
Job description
About the role
CoreWeave is seeking a technical expert to advance our AI infrastructure and model development capabilities. You will work on high-performance compute environments to optimize how large-scale models are trained and deployed. In this role, you will own the design and execution of performance experiments that validate infrastructure changes for next generation AI workloads. You will act as a deep technical consultant to both internal engineers and external customers who rely on our platform for demanding AI tasks. Your work will directly influence the roadmap for GPU compute orchestration and scheduling across our global data center footprint. You will collaborate closely with research scientists and software engineers to turn novel ideas into robust and scalable production systems. This position requires a strong sense of ownership to troubleshoot complex issues that span hardware, software, and networking layers. You will contribute to the creation of reference architectures that define best practices for running large language and diffusion models efficiently.
Key facts
What you'll do
- Architect and implement low level kernel optimizations and custom CUDA kernels for specific model architectures.
- Instrument end to end training pipelines to capture fine grained telemetry and identify performance bottlenecks.
- Partner with product teams to define and track key service level objectives for AI infrastructure reliability and throughput.
- Design experiments to measure the impact of new hardware features such as memory bandwidth and interconnect topology.
- Build and maintain reusable infrastructure components that abstract complexity for data scientists using our platform.
- Lead investigations into numerical stability and precision issues that arise during large scale distributed training runs.
- Evaluate and integrate emerging libraries and frameworks that can leverage the latest NVIDIA GPU architectures.
- Work with sales and solutions engineering to prototype custom solutions for strategic customer use cases.
- Document methodologies, findings, and implementation details to ensure knowledge transfer and repeatability.
- Drive the adoption of infrastructure best practices through code reviews, technical specifications, and cross team collaboration.
Requirements
- Demonstrated experience in applied research or machine learning engineering at a staff level in a technical leadership capacity.
- Show a deep understanding of GPU accelerated computing paradigms and the challenges of large scale model training.
- Prove proficiency in optimizing AI inference and training pipelines for latency, throughput, and resource utilization.
- Exhibit the ability to thrive in a fast moving high growth environment where priorities evolve based on customer needs.
- Have a track record of delivering production grade systems that operate reliably at extreme scale.
- Bring strong written and verbal communication skills to convey complex technical concepts to diverse audiences.
- Show ownership for driving projects from initial hypothesis through to final production deployment and postmortem analysis.
- Display intellectual curiosity and a commitment to continuous learning in rapidly evolving AI infrastructure domains.
Skills & tools
- Hands on experience with NVIDIA GPU architectures including Hopper, Blackwell, Vera Rubin, and Ada Lovelace.
- Expertise in machine learning frameworks and distributed training techniques across multiple nodes.
- Competence with Kubernetes based infrastructure management and container orchestration for AI workloads.
- Proficiency with performance profiling and benchmarking tools to analyze and optimize system behavior.
- Familiarity with infrastructure as code principles and configuration management for reproducible environments.
- Understanding of networking protocols and storage systems that impact AI workload performance.
- Experience with scripting and programming to automate repetitive tasks and validate infrastructure changes.
- Knowledge of observability platforms and monitoring tools to maintain visibility into production systems.
Nice to have
- Contributions to open source projects related to machine learning, high performance computing, or infrastructure tooling.
- Experience with cloud provider specific services and their integration into AI platform offerings.
- Background in compiler technologies or low level optimization for GPU workloads.
- Participation in academic research or industry conferences that demonstrate thought leadership in the field.
Practical notes
CoreWeave is a recognized leader in AI cloud infrastructure, recently named a Visionary in the Gartner Magic Quadrant for Cloud AI Infrastructure. We provide purpose built hardware and software to support advanced AI research and production workloads. CoreWeave roles are based in the United States and require candidates to be eligible for employment without sponsorship at the time of hire. This position is full time and may require occasional travel to customer sites or partner conferences as needed. CoreWeave reserves the right to adjust roles, responsibilities, and team assignments to align with evolving business priorities and operational needs. All employees are expected to adhere to company policies and code of conduct while contributing to a collaborative and inclusive work environment.