Software Engineer - Platform Core
Job description
About the role
You will own the design and evolution of the lean, high-reliability Linux-based operating system that serves as the foundational layer for our supercomputer network fabric. You will write or rewrite high-performance device drivers to extract the maximum physical capability from our network and compute hardware in demanding environments. You will manage the full deployment, operation, and debugging of our platform in production, driving improvements in reliability through rigorous root-cause analysis. You will collaborate directly with hardware teams and external partners to design the next generation of supercomputer hardware and bring up our OS and software stack on new platforms. You will develop and continuously improve observability and configuration management tools to maintain clear insight into system behavior. You will ensure strong communication of technical concepts, concisely and accurately sharing knowledge with teammates across roles. You will exercise initiative and ownership, leading projects that demonstrate impact on the company mission. Work ethic and strong prioritization skills will guide your execution in a fast-moving, high-stakes environment.
Key facts
What you'll do
- Architect and maintain the minimal, high-reliability Linux operating system that underpins our distributed supercomputer infrastructure.
- Engineer high-performance device drivers that push network and compute hardware to its physical limits.
- Oversee deployment pipelines, runtime operations, and incident response for our platform in production environments.
- Diagnose and resolve anomalies at scale, performing deep root-cause analysis to drive lasting system improvements.
- Enhance observability tooling and configuration management systems to provide reliable insight and control.
- Partner with hardware engineering and external collaborators to define next-generation supercomputer hardware and software integration.
- Bring up operating systems and software stacks on novel hardware platforms from bare metal through production-ready states.
- Optimize device and peripheral drivers, focusing on high-speed interfaces, DMA, and cache coherency challenges.
- Evaluate, debug, and integrate protocols and hardware across Ethernet, PCIe, I2C, SPI, and other critical interfaces.
- Implement and refine security hardening, product security measures, and secure boot schemes for system integrity.
- Utilize advanced debugging tools spanning kernel, network, and userspace domains, including ftrace, perf, Wireshark, tcpdump, eBPF, and gdb.
- Document system behavior and interfaces to support collaboration and long-term maintainability.
Requirements
- Perform hands-on systems programming in C or C++ as a core part of your daily work.
- Demonstrate strong operating systems fundamentals, including scheduling, memory management, and I/O, with the ability to explain them in extreme detail.
- Show firm understanding of computer networking, both within host operating system stacks and across common network devices like switches and routers.
- Communicate clearly in written and verbal form, conveying complex technical ideas accurately and concisely.
- Exhibit ownership and initiative, thriving in a flat organizational structure where leadership is earned through results.
- Maintain a strong work ethic and the ability to prioritize tasks effectively in a mission-critical environment.
- Collaborate productively with small, highly motivated teams focused on engineering excellence.
- Apply curiosity and problem-solving skills to challenging problems that directly impact the company mission.
- Meet the demands of a fast-paced environment where continuous learning and adaptation are essential.
- Align with a culture that values direct contribution to product, infrastructure, and long-term technical impact.
Nice to have
- Deep knowledge of the Linux kernel and its networking stack, or transferable experience with another production-grade kernel.
- Proficiency with kernel, network, and userspace debugging tools such as ftrace, perf, Wireshark, tcpdump, eBPF, and gdb.
- Experience bringing up new hardware from scratch, including flashing bare-metal software and bootstrapping operating systems on new platforms.
- Demonstrated history of developing and optimizing device and peripheral drivers, especially for high-speed interfaces.
- Understanding of major peripheral hardware standards and the ability to debug them using schematics and signal analysis.
- Familiarity with product security practices, OS hardening techniques, and secure boot implementations.
Practical notes
- Full-time employment is expected with on-site presence in Palo Alto, California, or Seattle, Washington as applicable.
- Compensation details are provided as a range reflecting experience and location.
- Employment is contingent on eligibility to work in the United States and compliance with all visa requirements.
- Applicants are encouraged to move quickly through the process if interested, as roles are filled on a rolling basis.