Staff Software Engineer
Job description
About the role
Graphcore is developing the next generation of AI compute hardware and software. We are seeking a Staff Software Engineer to join our team in building high-performance libraries that power our AI ecosystem. In this role, you will own the design and implementation of critical C++ components that form the backbone of our AI compute stack. You will drive technical direction for complex projects, ensuring that our libraries meet the stringent demands of performance and correctness. Your work will directly influence how efficiently our hardware executes the most demanding neural network workloads. You will collaborate closely with hardware engineers to align software behavior with silicon capabilities and constraints. This position requires a proactive approach to solving ambiguous technical challenges at the intersection of compilers, runtime systems, and linear algebra.
Key facts
What you'll do
- Develop and implement C++ kernels for tensor operations and linear algebra, including convolutions, GEMM, batched GEMM, reductions, and fused operations.
- Provide technical leadership by guiding engineering decisions, reviewing designs, and defining project approaches for complex software initiatives.
- Maintain product quality and performance through the creation of regression tests, microbenchmarks, and numerical validation methodologies.
- Optimize code for upcoming AI hardware by focusing on memory layout transformations, kernel launch efficiency, cache locality, and threading strategies.
- Resolve technical bugs and improve overall system functionality across the software stack affecting performance and stability.
- Participate in Agile workflows and mentor team members to support collective growth and elevate the engineering standard.
- Partner with compiler and framework teams to ensure seamless integration and optimal execution of graph operations.
- Analyze performance profiling data to identify bottlenecks and devise solutions that maximize throughput and minimize latency.
- Contribute to the definition and enforcement of coding standards, best practices, and architectural principles for long-term maintainability.
- Explore and prototype new algorithms or optimizations to evaluate their feasibility for integration into production systems.
Requirements
- Proficiency in Python and C++ programming is essential for interacting with frameworks and implementing high-performance logic.
- Experience with Linux-based profiling and knowledge of processor architectures is required to effectively analyze and improve performance.
- Proven history of technical leadership, including mentoring others and driving high-quality engineering outcomes, is mandatory.
- Ability to function at a Staff level, balancing hands-on coding with high-level technical ownership and strategic impact.
- Strong communication skills and a collaborative approach to teamwork are necessary to work effectively with cross-functional partners.
- Demonstrated capability to work independently with minimal supervision while ensuring alignment with team and company goals.
- A strong commitment to maintaining high standards of code quality, testing, and documentation is expected.
- Willingness to engage in deep technical investigations and root cause analysis for challenging performance issues.
Nice to have
- Expertise in algorithmic performance, including vectorization, lock-free patterns, threading, and memory hierarchy optimizations.
- Experience testing performance-sensitive or numerical code to ensure correctness under varied conditions.
- Familiarity with benchmarking, reproducibility, determinism, and tolerance design for numerical algorithms.
- Practical experience with at least one DNN or BLAS stack, such as oneDNN, cuDNN, or similar libraries.
- Knowledge of CPU micro-optimizations and numerical stability across FP8, BF16, FP16, and FP32 data formats.
- Experience integrating native code into frameworks like PyTorch via custom ops, dispatch keys, or extensions.
- Familiarity with Linux packaging, dependency management, and API/ABI stability considerations for library distribution.
Practical notes
This is a full-time position based in Gdańsk, Pomeranian Voivodeship, Poland. The salary range for this role is 350,700 - 474,400 PLN. The role is part of the ML Kernels & Runtime team, where collaboration with hardware and software teams is central. Applicants must be authorized to work in Poland without sponsorship. Relocation support is not available for this role. The position requires a high level of initiative and ownership to drive projects from conception to production. Successful candidates will engage in continuous learning to keep pace with rapidly evolving AI hardware and software landscapes. This role is critical to the performance and competitiveness of Graphcore's AI compute solutions.