Staff Platform Engineer
Job description
About the role
DeepL is building a global AI product and research company focused on secure, intelligent solutions for complex business problems, and this role sits at the heart of that mission. You will take ownership of the reliability, capacity, and cost efficiency of DeepL's compute infrastructure, including the hybrid platform serving products across AWS and on-prem, as well as the on-prem GPU clusters that power research. You will shape the platform's architecture by bringing strong technical positions, weighing them with peers, and committing to decisions that serve the best outcome for the organization. You will design, build, and operate production-grade Kubernetes clusters across cloud and on-prem infrastructure, guiding them from design through long-term operation and continuous improvement. You will drive the technical work of deepening the hybrid model, unifying how workloads run across on-prem hardware and cloud, especially AWS, while supporting research as it adopts cloud-native technology and ways of working. You will define infrastructure-as-code and platform standards that hold across the track, and strengthen observability and security so the platform's behavior is legible to the teams that depend on it. You will mentor engineers across the infrastructure track and raise the technical bar of the group through design reviews, architecture decisions, and everyday habits that keep the platform maintainable. You will build consensus across engineering teams and turn it into shipped change, shaping platform direction at the track level while leading incident response for hybrid infrastructure and driving remediation through to root cause.
Key facts
What you'll do
- Take ownership of the reliability, capacity and cost efficiency of DeepL's compute infrastructure, CPU and GPU: the hybrid platform serving our products across AWS and on-prem, and the on-prem clusters powering our research, including their hardware lifecycle.
- Shape the platform's architecture: bring strong technical positions, weigh them on their merits with your peers, and commit to decisions serving the best outcome.
- Design, build and operate production-grade Kubernetes clusters across cloud and on-prem infrastructure, taking them from design through to long-term operation.
- Drive the technical work of deepening the hybrid model - unifying how workloads run across on-prem hardware and cloud, especially AWS, and supporting research as it adopts cloud-native technology and ways of working.
- Define infrastructure-as-code and platform standards that hold across the track, and strengthen observability and security so the platform's behaviour is legible to the teams that depend on it.
- Mentor engineers across the infrastructure track and raise the technical bar of the group, through design reviews, architecture decisions and the everyday habits that keep the platform maintainable.
- Build consensus across engineering teams and turn it into shipped change, shaping platform direction at track level.
- Lead incident response for hybrid infrastructure, driving remediation through to root cause, and help sustain an on-call rotation the team can carry for the long term.
- Define and evolve platform standards that keep infrastructure secure, observable, and cost-effective across on-prem and cloud environments.
- Partner closely with product and research teams to understand their workloads and ensure the platform meets their needs without compromising stability or performance.
- Implement automation to reduce manual toil, improve system resilience, and enable faster, more predictable delivery of platform capabilities.
- Monitor platform health at scale, using data and observability to guide capacity planning, performance improvements, and reliability initiatives.
- Collaborate with cross-functional teams to plan and execute infrastructure upgrades, migrations, and disaster recovery preparations.
- Contribute to the broader engineering community by documenting patterns, sharing learnings, and proposing improvements to tools and processes.
Requirements
- Deep, hands-on Kubernetes expertise: you have designed, built and operated clusters at scale, through their full lifecycle.
- Depth at scale in either public cloud or on-prem infrastructure, plus the credibility to work in the other: AWS experience is preferred, and bare-metal or data centre experience is strongly valued.
- Networking and Linux depth you can debug with: you can follow a problem from a container through the host to the network edge.
- Infrastructure as code with Terraform or equivalent, and GitOps delivery with tooling such as ArgoCD, applied to environments you have owned in production.
- Software engineering skills in at least one major language, Go or Python preferred, at the level of building tooling other engineers rely on.
- A track record of technical ownership spanning several teams: you create clarity where a mandate is absent and you are comfortable making decisions with incomplete information.
- Comfort operating in an environment where ambiguity is common and change is the default, balanced by a bias for action and structured thinking.
- Effective written and verbal communication, with the ability to translate technical concepts to both technical and non-technical audiences.
Practical notes
- The role is based in London.
- Engagement is full_time.
- Hours, travel, visa, or deadlines are not specified in the source.