Staff Engineer
Job description
About the role
Graphcore is a SoftBank Group company focused on building the hardware and software infrastructure required for advanced artificial intelligence. This role sits within the System Management team, where you will develop interfaces that bridge the gap between physical hardware and customer AI workloads. You will lead the engineering of rack management solutions while ensuring high-performance operation across internal and external datacenters. The position requires deep collaboration with cross-functional partners to define and deliver resilient system management capabilities. You will own the design decisions that impact large-scale AI infrastructure deployments. This role is instrumental in translating operational requirements into robust software platforms. Your work will directly influence the reliability and efficiency of AI training and inference workloads. You will play a key role in the evolution of Graphcore's datacenter infrastructure stack.
Key facts
What you'll do
- Manage the entire software development life cycle for rack management systems, from initial design through to production deployment and automated testing.
- Collaborate across internal departments to resolve infrastructure issues and maintain system reliability.
- Utilize Infrastructure-as-code and Continuous Deployment to configure and validate new AI hardware in production environments.
- Partner with Datacenter Operations to monitor fleet performance and execute corrective maintenance.
- Lead technical planning and scoping within an Agile and Scrum framework, identifying project risks and constraints.
- Design and implement software interfaces for hardware telemetry, control, and lifecycle management.
- Drive the standardization of deployment patterns and operational runbooks for rack-scale systems.
- Analyze performance metrics and system logs to guide optimizations in hardware and software integration.
- Coordinate with security and compliance teams to ensure infrastructure solutions meet organizational policies.
- Mentor junior engineers on best practices for system software development and operations.
- Implement automation to reduce manual effort and increase consistency in managing AI datacenter infrastructure.
- Evaluate emerging technologies and tools to future-proof the rack management platform.
- Facilitate cross-team alignment to ensure software deliverables support broader product and business goals.
- Oversee the integration of new hardware features into existing management and monitoring workflows.
Requirements
- Bachelor degree in a relevant field or equivalent professional experience.
- Proficiency in Go programming and Linux systems engineering, including Bash and Python scripting.
- Experience developing RESTful APIs.
- Background in managing production Kubernetes clusters and containerized workloads using Docker or Podman.
- Hands-on experience with Infrastructure-as-code tools such as Terraform, OpenTofu, or Ansible, along with CI/CD platforms like GitLab or GitHub Actions.
- Knowledge of Redfish for datacenter hardware provisioning, telemetry, and control.
- Ability to manage project work plans, priorities, and technical documentation.
- Strong understanding of systems architecture and the ability to translate requirements into technical solutions.
- Demonstrated experience working in fast-paced, iterative software development environments.
- Capability to independently own complex components and drive them to production readiness.
- Excellent written and verbal communication skills for collaborating with technical stakeholders.
- Willingness to adapt to evolving priorities while maintaining a focus on system stability.
- Commitment to following coding standards, conducting code reviews, and improving software quality.
- Understanding of datacenter networking concepts and the implications for system management software.
Nice to have
- Familiarity with Kubernetes operator development and custom resources.
- Experience with High Performance Computing environments, specifically SLURM.
- Knowledge of virtualized deployments using KVM, QEMU, or Open vSwitch.
- Experience with distributed storage systems like Ceph.
- Proficiency in monitoring and observability stacks including Prometheus, Grafana, OpenSearch, Loki, Mimir, OpenTelemetry, Fluentd, or Kafka.
- Experience configuring managed switches such as SONiC, EOS, or DNOS.
- Familiarity with PyTorch for AI workloads.
- Experience using AI coding assistants.
- Background in end-to-end pipeline automation for build, test, and deployment.
Practical notes
Benefits include medical, dental, and vision coverage, 401(k) retirement plans, disability and life insurance, commuter benefits, and wellness services. Flexible spending accounts and health savings accounts are available. The company provides an equal opportunity hiring process and supports reasonable adjustments for candidates during the interview stage.