Software Engineer - Compute
Job description
About the role
Join a team focused on developing the underlying infrastructure for both public and private cloud services. This role is integral to enabling the provisioning of both bare-metal and virtual machine environments that power our platform. You will be responsible for enhancing the developer and customer experience by streamlining software delivery workflows and improving our internal platform tooling. The position demands a deep collaboration with cross-functional partners to ensure compute resources are reliable, scalable, and efficient. You will contribute directly to the core systems that drive our infrastructure capabilities. This is an opportunity to build foundational software that touches a wide range of users and services. Your work will ensure that our compute layer can handle demanding workloads with consistency and performance.
Key facts
What you'll do
- Design, build, and maintain software for GPU and CPU compute infrastructure, prioritizing performance, scalability, and reliability above all other metrics.
- Implement and develop services that provision and manage bare-metal and virtual machine instances throughout their lifecycle.
- Create distributed systems responsible for managing and orchestrating compute resources across a wide array of product configurations and customer demands.
- Troubleshoot and debug intricate issues that span both production and development environments with a high degree of independence.
- Assume ownership of on-call responsibilities and lead incident resolution efforts from detection through remediation.
- Collaborate with multiple internal and external teams to resolve ambiguities in requirements or proposed solutions during design review sessions.
- Analyze complex system behaviors to identify root causes and implement robust, long-term fixes that prevent recurrence.
- Partner with product and infrastructure teams to translate high-level requirements into concrete technical specifications for compute platforms.
- Contribute to the evolution of our operational playbooks and runbooks to improve reliability and automation.
- Evaluate new technologies and tools to assess their applicability to our compute infrastructure roadmap.
Requirements
- Possess a minimum of 3 years of hands-on experience writing Go (Golang) or Python in production-grade software environments.
- Bring at least 3 years of direct experience managing and configuring bare metal servers and virtualization hardware in data center settings.
- Demonstrate comfort working within Linux environments and the ability to debug issues that span the operating system, hardware, and networking layers.
- Exhibit the ability to independently troubleshoot highly complex distributed systems and service architectures.
- Communicate effectively and professionally with software engineers, infrastructure specialists, and vendor representatives to resolve technical challenges.
- Hold the right to work authorization in the United States for this position located in California.
- Thrive in a fast-paced environment where priorities can shift rapidly and adaptability is essential.
- Maintain a strong sense of ownership for the systems you build and the outcomes they deliver to end users.
Nice to have
- Familiarity with GPU infrastructure or high-performance computing environments and the associated workload patterns.
- Hands-on experience with cluster management tools such as Slurm or Kubernetes in large-scale deployments.
- Deep understanding of core public cloud internals, including virtualization technologies, KVM, QEMU, security best practices, and fleet health monitoring.
- Practical experience with durable execution platforms and workflow engines such as Temporal to manage long-running operations.
Skills & tools
- Proficiency in Go (Golang) for building high-performance backend services.
- Expert-level knowledge of Python for scripting, automation, and integration tasks.
- Extensive experience with Linux operating systems, distributions, and shell environments.
- Strong understanding of virtualization technologies and how they intersect with physical hardware.
- Working knowledge of bare metal hardware components, provisioning, and lifecycle management.
Practical notes
This role requires working from our San Francisco, San Jose, or Bellevue office four days per week, with Tuesday designated as a work-from-home day to allow for focused work. We offer generous cash and equity compensation, comprehensive health, dental, and vision coverage, wellness and commuter stipends for select roles, a 401k plan with a 2% company match for USA employees, and a flexible paid time off plan. The compensation range specified for this role varies based on location, with specific bands for San Francisco and San Jose provided in the key facts section to ensure transparency.