Principal Systems Software Engineer
Job description
About the role
You will architect the fluid fabric that unifies Bare-Metal-as-a-Service, Intelligent IaaS, and Elastic CaaS into a single high-performance pool of intelligence. You will bridge the gap between silicon and software while advising executive leadership on critical hardware and software co-design decisions. You will lead elite R&D teams in shipping production-grade kernel and orchestration code that pushes massive-scale training workloads to the theoretical limits of hardware. You will serve as the visionary lead for Crusoe's next-generation AI infrastructure and redefine the I/O path for the age of generative AI. You will draft white papers and RFCs that define the next two years of our compute and networking stack. You will represent Crusoe in open-source communities and industry forums to influence the global direction of cloud-native AI infrastructure.
Key facts
What you'll do
- Unifying Infrastructure Pillars: Architect systems that deliver raw GPU throughput via zero-latency InfiniBand/RDMA fabrics for massive-scale training while designing highly optimized thin virtualization layers using KVM or custom micro-VMs to provide enterprise-grade isolation without the virtualization tax.
- Elastic CaaS and Orchestration: Build a high-performance container substrate utilizing Kubernetes or Slurm that allows AI workloads to burst and scale across heterogeneous GPU nodes with minimal scheduling latency.
- Mastering the I/O Path: Lead the architectural design of our internal cloud fabric, drawing on experience from top-tier hyperscalers to drive the technical roadmap for SR-IOV, RDMA, and virtualized GPU scheduling.
- Advanced R&D Leadership: Lead elite workstreams to prototype and productionize novel methods for managing memory, networking, and compute that do not yet exist in standard cloud distributions.
- Technical Strategy and Documentation: Draft white papers and RFCs that define the next two years of Crusoe's compute and networking stack while ensuring alignment with energy-first infrastructure goals.
- High-Level Debugging and Optimization: Work alongside Staff and Senior engineers to resolve complex race conditions in the I/O path and optimize kernel-level memory pinning for GPU clusters at scale.
- Industry Influence and Collaboration: Represent Crusoe in open-source communities and industry forums to influence the global direction of cloud-native AI infrastructure and drive best practices.
Requirements
- 12+ years of experience designing and shipping core infrastructure at a major hyperscaler such as OCI, AWS, Azure, or GCP, or at a specialized HPC cloud with a track record of high-scale operations.
- Authoritative knowledge of the Linux kernel, virtualization internals including KVM, QEMU, and Firecracker, and high-performance networking technologies such as RoCE v2 and InfiniBand.
- Proven ability to design software that maximizes the performance of NVIDIA or AMD GPUs and high-speed NICs through hardware-software co-design.
- Experience leading cross-functional R&D teams through high-ambiguity projects and delivering production-ready, mission-critical systems on demanding timelines.
- A portfolio of significant contributions to the field, which may include patents, major open-source contributions, or published research in distributed systems and infrastructure.
- The rare ability to communicate technical nuances of memory-mapped I/O to engineers and articulate the business value of new fabric architectures to boards and executive stakeholders.
- A Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related analytical field, or equivalent professional experience that demonstrates deep systems mastery.
- Full-time authorization to work in the United States and eligibility to perform contract duties in San Francisco, CA, without sponsorship.
- Willingness to work from the Crusoe San Francisco office and travel as required for team collaboration, manufacturing visits, or customer engagements.
Nice to have
- Patent Holder: Possession of patents related to network virtualization, GPU scheduling, or distributed file systems that align with our infrastructure innovations.
- Open Source Leadership: Maintained significant open-source contributions to virtualization, networking, or container orchestration projects that are widely adopted in the cloud-native ecosystem.
Practical notes
This is a full-time position based in San Francisco, California. Employees are expected to work from the Crusoe San Francisco office. Travel may be required for team collaboration, manufacturing visits, or customer engagements. Visa sponsorship is not available for this role.