Staff System Software Engineer
Job description
About the role
Graphcore is seeking a technical expert to join our system software department. You will focus on the design and implementation of kernel drivers and user-space libraries that power our AI compute technology. In this capacity, you will own the end-to-end lifecycle of critical system software, ensuring it meets the stringent performance and reliability standards required by our hardware. You will act as a foundational engineer, translating architectural intent into robust, production-grade components that directly enable AI workloads. The role demands a deep partnership with hardware teams to co-design solutions that maximize silicon capabilities. You will be responsible for debugging complex interactions across the hardware and software stack. Success in this position will be measured by your ability to deliver stable systems software that empowers machine intelligence at scale.
Key facts
What you'll do
- Architect, build, and test kernel and device driver software using C, C++, and Python, ensuring high integrity and performance.
- Implement low-level system components that interface directly with silicon, focusing on stability and efficiency.
- Diagnose and resolve intricate software defects using advanced debug and performance analysis tools to meet strict timelines.
- Collaborate within an agile scrum team to deliver software components on schedule, adapting to evolving technical requirements.
- Engage with hardware and silicon engineering groups during product development cycles to align on integration strategies.
- Coordinate with architects and internal stakeholders to resolve technical challenges and define system-level interfaces.
- Develop and maintain user-space libraries that abstract hardware complexity for application developers.
- Optimize software pipelines to leverage advanced hardware features such as PCIe topologies and SoC architectures.
- Contribute to the upstream Linux kernel where applicable, ensuring compatibility and adherence to best practices.
- Support the integration of firmware components and validate their behavior in system-level contexts.
- Participate in the design of data center software stacks, considering aspects of cloud operations and deployment scenarios.
- Ensure all deliverables comply with security, reliability, and operational standards required for global customer deployments.
Requirements
- Professional background in software development using C, C++, or Python, with a strong emphasis on systems-level programming.
- Experience developing and deploying OS kernel and device drivers for Linux or Windows environments in production scenarios.
- Understanding of low-level software stacks and hardware-layer interactions, including memory management and interrupt handling.
- Proficiency with debug and performance analysis tools such as profilers, tracers, and kernel debuggers.
- Experience with at least one of the following: PCIe, CPU, SoC, firmware, or hardware/software integration methodologies.
- Ability to manage your own workload and proactively seek input from engineering leadership to align priorities.
- Strong communication skills for working within a multinational team and supporting global customers across multiple time zones.
- A methodical approach to problem-solving, with the patience to investigate deeply nested software and hardware issues.
- Commitment to writing clean, documented, and maintainable code that can be reviewed and extended by peers.
- Willingness to adhere to strict coding standards and version control practices essential for collaborative development.
- Readiness to engage in on-call responsibilities to support critical production incidents as part of a global team.
Nice to have
- Experience developing drivers or firmware for GPUs, providing insight into high-throughput compute devices.
- Familiarity with CUDA or OpenCL, understanding how applications interact with hardware via these frameworks.
- History of contributing to the upstream Linux kernel, demonstrating familiarity with community processes and code reviews.
- Exposure to data center or cloud operations, such as Kubernetes or OpenStack integration, for managing distributed workloads.
Skills & tools
- C, C++, Python
- Linux/Windows Kernel Development
- PCIe, SoC, Firmware
- Debugging and Performance Analysis
Practical notes
This is a full-time position based in Cambridge, UK. The compensation package includes a competitive salary, flexible working arrangements, and a generous annual leave policy. The total rewards package also features private medical insurance, a health cash plan, a dental plan, pension matching up to 5%, life assurance, income protection, and parental leave. Applicants must be eligible to work in the United Kingdom without sponsorship. The role involves collaboration with global teams, requiring occasional travel and participation in meetings across different time zones. Graphcore is committed to building a diverse team and encourages applications from all qualified individuals.