Staff Software Engineer, Inference
Job description
About the role
You will design and build high-performance systems to support large-scale AI inference workloads on our specialized GPU infrastructure. This role focuses on optimizing model serving and runtime acceleration to ensure low-latency performance for complex AI applications. You will own the end-to-end lifecycle of critical inference infrastructure, from initial architecture and design through implementation, testing, and large-scale deployment. The position demands close collaboration with hardware, platform, and product teams to define and deliver the performance and reliability required for next-generation AI workloads. You will be responsible for translating abstract performance and scalability goals into concrete, production-ready software components. This role operates at the intersection of distributed systems, high-performance computing, and cloud infrastructure, requiring deep technical rigor and ownership. You will play a key role in shaping the developer experience and operational characteristics of our inference platform.
Key facts
What you'll do
- Architect and implement systems for high-throughput, low-latency model inference.
- Optimize runtime environments to maximize the efficiency of GPU compute resources.
- Develop tools and services that integrate with our managed Kubernetes and fleet lifecycle controllers.
- Collaborate with engineering teams to refine performance benchmarks for AI inference workloads.
- Build scalable solutions that support the deployment of large-scale models across our cloud platform.
- Design and implement high-performance networking protocols and communication primitives to minimize latency and maximize throughput between GPU nodes.
- Create observability and debugging tools specifically tailored for AI inference workloads to diagnose performance bottlenecks and system issues.
- Partner with hardware engineering teams to validate and optimize software stacks for new GPU architectures and accelerator technologies.
- Implement advanced scheduling and resource allocation strategies to improve utilization and reduce queue times for inference jobs.
- Lead the design of stateful and stateless components within the inference serving stack to ensure robustness and scalability under load.
- Collaborate with product managers to define and deliver technical capabilities that meet evolving customer requirements for AI deployment.
- Conduct in-depth performance analysis and profiling to identify and resolve inefficiencies in the inference runtime and application stack.
- Mentor junior engineers by providing technical guidance and code review focused on performance, reliability, and scalability best practices.
- Drive the adoption of infrastructure-as-code practices for deploying and managing inference workloads across hybrid cloud environments.
Requirements
- Extensive experience in software engineering with a focus on distributed systems or high-performance computing.
- Proven track record of designing and deploying production-grade inference services.
- Proficiency in optimizing workloads for GPU-accelerated environments.
- Deep understanding of container orchestration and infrastructure management.
- Strong experience with low-latency, high-throughput network programming and protocols.
- Demonstrated ability to work effectively in a fast-paced, rapidly evolving technical environment.
- Excellent problem-solving skills and a methodical approach to debugging complex, distributed performance issues.
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Nice to have
- Experience with MLPerf benchmarking or similar performance evaluation frameworks.
- Background in developing runtime acceleration technologies or model serving stacks.
- Familiarity with Tensorizer or similar model optimization tools.
Practical notes
CoreWeave is an AI-native cloud provider focused on high-performance infrastructure. Applicants should be prepared to work in a fast-paced environment centered on specialized hardware and large-scale AI deployments. CoreWeave reserves the right to adjust roles, responsibilities, and requirements based on evolving business and technical needs. This position is eligible for CoreWeave's comprehensive benefits package, including health, dental, and vision insurance, as well as retirement planning options. CoreWeave is an equal opportunity employer and encourages applications from all qualified individuals regardless of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status. The roles and locations listed may be subject to change based on project and business needs. CoreWeave may conduct background checks as part of the hiring process. This job description is not an employment contract, and CoreWeave reserves the right to modify or terminate employment at any time at its discretion.