Staff Site Reliability Engineer
Job description
About the role
Staff Site Reliability Engineer, Platform Reliability
We are searching for a Staff Site Reliability Engineer to join the Platform Engineering reliability team at Doctolib. In this role, you will act as a technical leader responsible for platform stability and reliability at a European scale. Your work will guide reliability standards across more than 170 applications. You will directly support 520,000 health professionals and 90 million patients by ensuring that our platform remains dependable, debuggable, and resilient. This position sits at the intersection of infrastructure, developer experience, and product engineering. You will partner with SREs, software engineers, and product teams, providing technical guidance, mentorship, and leadership for cross-cutting reliability initiatives.
What you will do
You will design, implement, and champion large-scale reliability initiatives that span infrastructure automation, observability pipelines, and incident management workflows. You will define and evolve measurable service objectives, error budgets, and alerting standards. Your focus will be on reducing noise and improving telemetry clarity for every product team. You will steer on-call rotations and refine response processes to ensure incidents drive improvement and learning. Mentoring senior engineers will be a core responsibility, as you coach reliability practices and elevate engineering craft through reviews and feedback.
You will partner with software engineers early in the development lifecycle, embedding reliability practices during design and launch readiness for complex features. Your influence will extend to architecture reviews, where you will represent reliability concerns and align leadership on long-term platform health. You will build trusted relationships with feature teams, accelerating production readiness through guidance, checklists, and operational standards. You will champion observability architecture, improving instrumentation quality and scaling telemetry pipelines for regulated, high-stakes domains. You will operate within a regulated healthcare context, where privacy and compliance shape every technical decision and deployment practice. You will advance reliability enablement through golden paths, runbooks, and shared libraries that simplify operation at European scale.
Who you are
Before you continue, consider this: if you do not match every detail below exactly but your experience aligns with this role, we still encourage you to apply.
- Bring eight or more years of experience managing large-scale, multi-team production environments with clear ownership of infrastructure and reliability.
- Demonstrate extensive experience with cloud platforms such as AWS, GCP, or Azure, along with fluency in container orchestration and Kubernetes deployment and scaling strategies.
- Have designed, implemented, and operated service-level indicators, service-level objectives, and error budgets in production under demanding traffic patterns.
- Have led on-call rotations and directed incident response in high-stakes situations, driving thorough postmortems and corrective actions.
- Use at least one backend programming language, such as Go or Python, applying strong systems thinking to complex reliability challenges.
- Communicate clearly in writing and speech, aligning stakeholders and driving decisions amid ambiguity across distributed teams.
- Balance meticulous long-term architecture with fast, iterative improvements, maintaining quality while accelerating delivery.
- Partner with feature teams before go-live, providing hands-on guidance on reliability, conducting launch reviews, and ensuring operational readiness.
- Speak English fluently to collaborate across regions and contribute effectively to global standards and documentation.
Nice to have
Deep expertise in observability tooling, including logging, tracing, metrics architecture, and strategies for high-cardinality data. Experience designing telemetry pipelines and improving instrumentation quality while working closely with developers on signal quality. Backgrounds in regulated industries such as healthcare, fintech, or similar environments where privacy and compliance directly influence engineering decisions. Interest in reliability enablement, golden paths, runbooks, and shared libraries that scale across organizations.
Skills and tools
Proficiency with Kubernetes; experience on AWS and GCP; scripting and programming in Python or Go; observability platforms; and structured incident management processes.
Practical notes
Please What you'll do
- Meet the bar