Systems Engineer
Job description
About the role
This role monitors infrastructure and coordinates responses to protect AI technology and applied AI initiatives. The position bridges hardware and software teams to ensure reliable compute operations in Milpitas, California. You will act as the operational link between development and deployment, focusing on the stability of the infrastructure that powers AI workloads. Success in this position requires a proactive approach to identifying potential issues before they impact research or production systems. You will be responsible for maintaining the integrity and performance of the environments where AI models are trained and inferred. The role demands a balance between deep technical investigation and coordination with distributed teams. You will translate high-level requirements into concrete operational tasks that keep the infrastructure online. Ultimately, you will safeguard the systems that drive Graphcore's applied AI initiatives and technology roadmap.
Key facts
What you'll do
Document configurations, procedures, and runbooks to support operational consistency for global operations and infrastructure teams.
Evaluate deployment risks and execute changes in production environments to protect system integrity and support continuous operations for AI technology and applied AI initiatives.
Plan capacity and infrastructure changes in collaboration with cross-functional stakeholders to align resource allocation with evolving product and research demands.
Monitor the health and performance of infrastructure to identify anomalies and initiate coordinated responses.
Maintain operational resilience by implementing safeguards and automating routine tasks to reduce manual overhead.
Partner with hardware and software teams to troubleshoot complex issues that span multiple layers of the stack.
Analyze trends in system usage and performance to recommend improvements that enhance efficiency and reliability.
Support the implementation of security and compliance controls relevant to infrastructure and operations.
Serve as a point of contact for incident response, ensuring clear communication during critical events.
Drive standardization efforts to ensure that operational practices are consistent across global teams.
Investigate root causes of infrastructure failures and contribute to long-term solutions that prevent recurrence.
Champion best practices for logging, monitoring, and documentation to improve the maintainability of environments.
Requirements
You have hands-on experience with Linux operating systems and administration.
You communicate effectively with both technical and non-technical audiences.
You are comfortable working in a fast-paced environment where priorities can shift quickly.
You have the ability to understand complex systems and explain them clearly to others.
You demonstrate strong problem-solving skills and a methodical approach to troubleshooting.
You show reliability in following through on tasks and meeting operational commitments.
You have experience collaborating with cross-functional teams to resolve shared challenges.
You adhere to strict standards for documentation and procedural accuracy.
Practical notes
Confirm eligibility before proceeding.
Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
About the company
Graphcore has built a new type of processor for machine intelligence to accelerate machine learning and AI applications for a world of intelligent machines
Career growth
Engineering careers usually progress from individual contributor to senior, staff, and principal levels. Some engineers move into management and lead teams of five to twenty people. Others stay on the technical track. Growth follows demonstrated impact, not tenure alone. A typical engineering ladder has clear levels with defined expectations for scope, quality, and mentorship. Moving up usually requires owning outcomes end to end rather than completing assigned tickets.
Questions to ask
Useful questions for the interview: what a typical week looks like, how work is assigned, what tools the team uses, and how feedback works. Asking how the role has changed recently and what the team wishes it had known when joining is also reasonable. Questions about the manager's priorities are especially valued.