Staff Software Engineer, Applied Training
Job description
About the role
You will design and implement systems that enable large-scale machine learning model training on our high-performance GPU infrastructure. This role focuses on optimizing the intersection of software engineering and model development to ensure high efficiency for complex AI workloads. You will own the architecture of critical infrastructure components that directly impact the performance and stability of training clusters. The position requires a deep partnership with research and product teams to translate demanding use cases into robust software solutions. You will be responsible for driving improvements that reduce training job failure rates and increase resource utilization. This role demands a proactive approach to identifying systemic issues before they impact large-scale training runs. You will play a key role in shaping the technical roadmap for the Applied Training organization.
Key facts
What you'll do
- Architect and develop high-throughput control plane services that orchestrate the lifecycle of massive machine learning training jobs.
- Instrument distributed systems to capture granular telemetry, enabling rapid diagnosis of performance anomalies during long-running training executions.
- Engineer scalable storage interfaces to handle the high I/O demands of modern dataset pipelines used in large model training.
- Implement resilient job scheduling algorithms that account for heterogeneous GPU clusters and dynamic workload priorities.
- Streamline the developer experience by building internal tools that abstract the complexity of managing multi-tenant GPU resources.
- Optimize network topology configurations to minimize data transfer latency across nodes in distributed training rings.
- Partner with data scientists to prototype new execution backends that validate novel approaches to model parallelism.
- Establish rigorous testing frameworks to ensure infrastructure changes do not introduce regressions in critical training pathways.
- Lead incident response efforts for platform outages, coordinating with cross-functional teams to restore service swiftly.
- Contribute to open source ecosystems where appropriate to maintain alignment with industry standards in AI infrastructure.
- Analyze utilization metrics to right-size infrastructure allocations, balancing cost efficiency with training performance goals.
- Mentor junior engineers by providing code reviews and technical guidance on best practices for scalable system design.
- Drive the adoption of infrastructure-as-code practices to ensure consistency and reproducibility across environments.
- Act as a technical liaison between platform engineering and customer-facing teams to align on strategic priorities.
Requirements
- Bachelor's degree or equivalent practical experience in Computer Science or a related technical field.
- Demonstrated experience of eight or more years in software engineering roles with increasing responsibility.
- Extensive hands-on expertise in building and maintaining software for large-scale GPU compute environments.
- Profound understanding of distributed systems principles, including consensus, replication, and fault tolerance.
- Mastery of container orchestration platforms, specifically Kubernetes, within the context of AI and machine learning workloads.
- Strong programming proficiency in at least one systems-level language such as Go or C++, combined with scripting ability in Python.
- Significant background in high-performance networking, including TCP/IP stack optimizations and RDMA technologies.
- Deep familiarity with storage systems, including parallel file systems and object storage, tailored for high-throughput AI workloads.
- Experience designing systems that meet stringent reliability and uptime requirements for production-grade services.
- Comfort working in fast-paced environments where requirements evolve based on technological advancements in AI.
- A track record of contributing to complex codebases that require strict security and compliance standards.
- Willingness to engage on-call rotation schedules to support critical infrastructure issues across time zones.
Nice to have
- Prior experience working with NVIDIA hardware stacks and related software ecosystems such as CUDA and cuBLAS.
- Background in optimizing deep learning frameworks like PyTorch or JAX for distributed training at scale.
- Hands-on experience with low-latency inference platforms or reinforcement learning environments.
- Familiarity with high-performance computing (HPC) paradigms and their application to AI workloads.
- Contributions to the Linux kernel or device driver development relevant to compute or networking subsystems.
Practical notes
This role is based in Sunnyvale, California, or Bellevue, Washington. The position is full-time and eligible for CoreWeave's comprehensive benefits package. Compensation details are not specified in the source material. This position requires engagement with on-call duties and the ability to travel as necessary for business operations. There are no specific deadlines for application outlined in the provided source data.