Site Reliability Engineer
Job description
About the role
You will own the stability and availability of production infrastructures for Theodo's clients on a daily basis. You will react to incidents in production and conduct thorough investigations when technical issues arise. You will own post-mortem creation and ensure clear, professional communication with clients about outages and resolutions. You will answer questions from our customers and provide actionable recommendations to enhance their infrastructure quality and resilience. This role involves on-call duties on a voluntary basis, offering flexible participation in incident response outside regular hours. The position is ideal for individuals who are deeply curious and eager to accelerate their learning through diverse challenges. You will have the opportunity to work across a wide range of infrastructures and tools, ensuring that no two days are the same.
Key facts
What you'll do
- React to incidents in production environments and coordinate rapid mitigation actions.
- Conduct in-depth investigations to identify root causes of technical failures and prevent recurrence.
- Own the creation of post-mortem documentation and present findings to both technical teams and clients.
- Communicate clearly with clients to explain complex technical issues in understandable terms.
- Provide recommendations to clients for improving the reliability and performance of their infrastructure.
- Monitor system health and performance metrics to identify potential issues before they escalate.
- Collaborate closely with development and operations teams to streamline incident response processes.
- Participate in continuous improvement initiatives for monitoring, alerting, and runbooks.
- Maintain and optimize infrastructure components to ensure high availability and scalability.
- Support the design and implementation of resilient architectures for cloud-based solutions.
- Mentor junior team members on best practices for reliability engineering and incident management.
- Engage with new technologies and tools to evaluate their potential impact on operational efficiency.
- Document operational procedures and standards to ensure consistency across client environments.
- Contribute to the development of automation strategies to reduce manual operational overhead.
- Act as a technical liaison between Theodo's internal teams and external clients during critical incidents.
Requirements
- Hold a solid academic degree from an engineering school or equivalent institution.
- Bring a first professional experience with infrastructure and cloud technologies, including Kubernetes.
- Demonstrate rigor and strong stress management capabilities to thrive in production environments.
- Show the ability to react quickly and make sound decisions under pressure.
- Possess a strong sense of responsibility for the impact of infrastructure issues on business operations.
- Have excellent analytical skills to diagnose complex problems in distributed systems.
- Communicate effectively both in writing and verbally with technical and non-technical stakeholders.
- Be comfortable working in an agile and fast-paced delivery environment.
- Show initiative and ownership when addressing operational challenges.
- Adhere to best practices in security, reliability, and infrastructure as code.
- Willingness to learn continuously and adapt to evolving tools and technologies.
- Ability to work independently and as part of a collaborative team.
- Commitment to maintaining high standards of quality in all operational tasks.
Practical notes
- Le poste est proposé en CDI, basé depuis nos locaux à Paris.
- Nous permettons à nos équipes de télé-travailler jusqu'à 3 jours par semaine.
- Le salaire proposé dépend du niveau d'expérience de chacun : nous en parlons toujours dès le premier appel téléphonique.
- Nous répondons à toutes les candidatures dans un délai de 3 jours.
- un appel téléphonique de 30 minutes avec un membre de l'équipe recrutement, pour comprendre tes critères de recherche
- un entretien d'une heure avec le recruteur ou la recruteuse avec qui tu as échangé à la première étape, pour parler plus en détail de la culture Theodo
- un entretien de logique
- un entretien technique de débug sur Docker de 45 minutes
- un échange de 30 minutes avec un membre de l'équipe dirigeante
- Tu seras accompagné(e) par une personne de l'équipe recrutement, qui sera là pour te coacher tout au long du process.
Nous répondons à toutes les candidatures dans un délai de 3 jours.
- un appel téléphonique de 30 minutes avec un membre de l'équipe recrutement, pour comprendre tes critères de recherche
- un entretien d'une heure avec le recruteur ou la recruteuse avec qui tu as échangé à la première étape, pour parler plus en détail de la culture Theodo
- un entretien de logique
- un entretien technique de débug sur Docker de 45 minutes
- un échange de 30 minutes avec un membre de l'équipe dirigeante
Tu seras accompagné(e) par une personne de l'équipe recrutement, qui sera là pour te coacher tout au long du process.