Senior Software Engineer, Compute Architecture
Job description
About the role
CoreWeave is seeking a Senior Software Engineer to join the Compute Architecture team. You will focus on building and scaling our specialized GPU infrastructure to support high-performance AI workloads. In this capacity, you will architect solutions that bridge the gap between cutting-edge silicon and demanding computational workloads. Your contributions will directly influence the performance and reliability of the platforms that power advanced AI research and deployment. You will operate at the intersection of cloud infrastructure and hardware acceleration to solve complex engineering challenges. This role requires a proactive approach to problem-solving and a commitment to excellence in software design. You will be responsible for ensuring that our systems meet the rigorous demands of modern machine learning workflows.
Key facts
What you'll do
- Architect and implement scalable systems to orchestrate massive GPU compute clusters across distributed environments.
- Engineer high-throughput data pathways to optimize infrastructure performance for AI training and inference tasks at scale.
- Construct robust software frameworks designed to improve the reliability and efficiency of hardware resource allocation and scheduling.
- Partner with cross-functional engineering teams to seamlessly integrate novel GPU architectures into our evolving cloud platform.
- Analyze intricate performance metrics to identify bottlenecks and refine the low-level operations of compute-intensive applications.
- Pioneer the development of infrastructure components that abstract complex hardware details for higher-level software services.
- Lead the design of resilient systems capable of maintaining stateful operations and fault tolerance in large-scale deployments.
- Mentor junior engineers by providing technical guidance on best practices for writing clean, maintainable code in a fast-paced collaborative environment.
- Evaluate emerging hardware technologies and determine their viability for integration into our core infrastructure stack.
- Drive the execution of proof-of-concept projects to validate new ideas for enhancing GPU utilization and workload isolation.
Requirements
- Demonstrate professional experience in software engineering with a focus on distributed systems or infrastructure, showcasing a track record of delivering reliable large-scale systems.
- Exhibit proficiency in developing software for high-performance computing environments, including an understanding of concurrency, networking, and I/O optimization.
- Possess a solid understanding of GPU architecture and its application in cloud computing, including memory hierarchies and execution models.
- Show the ability to write clean, maintainable code that adheres to strict standards and functions effectively within a collaborative engineering culture.
- Have hands-on experience with systems programming in languages commonly used in performance-critical infrastructure such as C++ or Rust.
- Bring a history of debugging complex issues in production environments using observability tools and log analysis.
- Illustrate strong communication skills necessary for collaborating with product managers, researchers, and infrastructure teams.
- Maintain a disciplined approach to testing and validation to ensure software quality and stability before deployment.
Nice to have
- Bring experience working with Kubernetes and container orchestration at scale to manage the lifecycle of distributed compute workloads.
- Offer a background in low-level systems programming or hardware-software integration, particularly involving accelerators and custom silicon.
- Show familiarity with AI/ML infrastructure stacks and performance benchmarking methodologies to guide optimization efforts.
Skills & tools
- Expertise in GPU compute technologies including but not limited to CUDA, ROCm, or corresponding frameworks.
- Deep knowledge of distributed systems architecture encompassing consensus, replication, and load balancing strategies.
- Mastery of cloud infrastructure management techniques utilizing major public cloud providers and virtualization platforms.
- Proficiency in software development for high-performance computing involving parallel algorithms and memory management.
- Experience with infrastructure-as-code tools to provision and manage complex networked environments.
- Understanding of containerization technologies and image optimization for rapid deployment cycles.
- Familiarity with monitoring and observability platforms to track system health and performance indicators.
- Competence with version control systems and modern CI/CD pipelines to automate testing and deployment workflows.
Practical notes
CoreWeave provides specialized infrastructure for AI and high-performance workloads. Candidates must be authorized to work in the location they are applying for, with roles potentially requiring on-site presence. Please submit your application through the official careers portal to be considered for this position. Only applications submitted via the designated portal will be reviewed. The company is an equal opportunity employer and encourages applicants of all backgrounds to apply. Specific role details and team assignments will be confirmed following initial screening stages.