Senior Systems Engineer - Performance & Reliability
Job description
Senior Systems Engineer - Performance & Reliability at Graphcore.
About the role
You will determine if hardware racks meet production standards by managing measurement processes for large-scale Linux clusters. This role balances infrastructure development, workload testing, and data analysis to ensure system reliability and performance. You will design and execute performance and reliability experiments on distributed systems to validate hardware readiness for deployment. The position requires you to analyze measurement results to provide evidence-based conclusions for production readiness and support critical business decisions. You will maintain and improve the infrastructure used to run tests at scale, ensuring the testbed remains robust and efficient. Measuring compute, network, and machine learning workload performance across rack-level and datacenter-scale systems is a core responsibility of this role. Evaluating result variability and repeatability will be essential to ensure accurate decision-making and system reliability over time. You will work closely with cross-functional teams to translate reliability requirements into measurable test criteria and procedures.
Key facts
What you'll do
- Design and execute performance and reliability experiments on distributed systems to validate hardware scalability.
- Analyze measurement results to provide evidence-based conclusions for production readiness and to guide optimization efforts.
- Maintain and improve the infrastructure used to run tests at scale, focusing on automation, observability, and resilience.
- Measure compute, network, and machine learning workload performance across rack-level and datacenter-scale systems to identify bottlenecks.
- Evaluate result variability and repeatability to ensure accurate decision-making and reduce uncertainty in deployment approvals.
- Automate test workflows and experiment orchestration to increase coverage, reduce manual effort, and improve consistency.
- Collaborate with hardware and software teams to correlate system-level performance with individual component behavior.
- Develop dashboards and reporting mechanisms that communicate complex performance data to both technical and non-technical stakeholders.
- Support the development of methodologies for workload characterization and scenario modeling for future product generations.
- Investigate anomalies in test data, diagnose root causes, and implement corrective actions to improve test integrity.
- Document testing procedures, configurations, and findings to ensure reproducibility and institutional knowledge retention.
- Contribute to the evolution of testing standards and best practices across the organization to align with industry benchmarks.
Requirements
- Significant software engineering experience across multiple projects or systems, demonstrating a broad understanding of software lifecycle and architecture.
- Proficiency in Python, including scripting, debugging, and writing maintainable code for test automation and data analysis.
- Experience working in Linux environments, ideally with high-performance or distributed systems, including command-line fluency and system troubleshooting.
- Practical knowledge of automation and CI/CD systems such as GitLab CI, Jenkins, or GitHub Actions to implement and monitor test pipelines.
- Ability to design experiments and communicate technical findings clearly through both written documentation and verbal presentations.
- Capacity to work independently in ambiguous environments where requirements are evolving, making sound decisions with incomplete information.
- Strong attention to detail and the ability to interpret complex data sets to draw meaningful conclusions about system behavior.
- Eligibility to work in Poland is mandatory, and applicants must hold the right to work in the country without the need for sponsorship.
Nice to have
- Experience with large-scale distributed systems, cloud platforms, or HPC environments to better understand complex deployment scenarios.
- Background in performance, reliability, or systems-level testing to contribute effectively from the early stages of test design.
- Familiarity with pytest or similar structured measurement frameworks to integrate testing into existing workflows.
- Experience analyzing system behavior under load, including compute, network, or ML workloads, to correlate performance with reliability outcomes.
- Knowledge of containerization, orchestration, or provisioning tools like Docker, Kubernetes, or OpenStack to manage test environments.
- Proficiency in C++ or other application programming languages to collaborate with low-level system components and performance-critical modules.
- Exposure to statistical data analysis to evaluate variability, confidence intervals, and trends in experimental results.
Practical notes
- Applicants must hold the right to work in Poland. Visa sponsorship is not available for this position.
- Benefits include annual leave, medical and dental plans, a gym card, and a company pension with up to 4% matching.
- Benefits are reviewed annually to ensure they remain competitive and aligned with employee needs.
- The role follows a hybrid schedule, requiring 2-3 days of in-office work per week to facilitate collaboration and knowledge sharing.
- No specific deadlines for application submission are provided, so interested candidates are encouraged to apply as soon as possible.
- Travel requirements are not specified, indicating that the position is primarily based at the assigned work location in Gdańsk.
- The engagement is full-time, aligning with standard working hours and expectations for consistent availability during business operations.