Engineering Manager
Job description
About the role
You will lead the OREO (Observability Reliability Engineering Obsession) team within our Platform Engineering division. Your primary responsibility is to scale our observability infrastructure and guide a high-performing SRE team while managing critical services that support our entire technical organization. You will act as a strategic and operational leader, ensuring the reliability and performance of our core platform. This role requires a balance of technical depth and people management to drive sustainable practices. You will foster a culture of psychological safety and operational excellence across the team. The position involves significant ownership over the roadmap and execution of observability initiatives. You will be the key driver in evolving our infrastructure monitoring and resilience strategies.
Key facts
What you'll do
- Mentor and develop Site Reliability Engineers, managing their career growth, performance reviews, and professional development opportunities.
- Define the technical roadmap for observability, including metrics, logging, alerting, and tracing, ensuring alignment with industry best practices.
- Oversee the operation and strategic direction of transversal services like Terraform Enterprise and HashiCorp Vault within the infrastructure landscape.
- Foster a culture of psychological safety, operational excellence, and continuous improvement through coaching and process iteration.
- Manage the on-call experience, including scheduling, escalation procedures, and leading postmortem reviews to prevent recurring system issues.
- Collaborate with product and engineering leadership to align platform capabilities with business needs and long-term strategic objectives.
- Recruit and onboard new talent to expand the team, ensuring a high bar for technical and cultural fit.
- Architect and evolve observability architectures using tools such as OpenTelemetry, Fluent Bit, Prometheus, Thanos, Loki, Elasticsearch, or Datadog.
- Drive the adoption of infrastructure as code and secrets management practices, specifically focusing on Terraform and HashiCorp Vault.
- Analyze complex system interactions and dependencies to improve the overall resilience and performance of critical platform services.
- Lead incident response efforts, ensuring clear communication and effective resolution during critical outages.
- Translate high-level business objectives into concrete technical requirements for the platform team.
- Evaluate emerging technologies and tools to enhance the observability and reliability of our healthcare platform.
- Build and maintain strong partnerships with other engineering teams to ensure cohesive platform development.
Requirements
- Minimum 5 years of experience in software engineering or SRE, specifically within cloud-native environments like Kubernetes, AWS, or GCP.
- At least 3 years of experience in engineering management, leading infrastructure, platform, or SRE teams in a fast-paced environment.
- Technical expertise in observability architectures using tools such as OpenTelemetry, Fluent Bit, Prometheus, Thanos, Loki, Elasticsearch, or Datadog.
- Proficiency in infrastructure as code and secrets management, specifically Terraform and HashiCorp Vault.
- Ability to bridge the gap between high-level strategic planning and hands-on technical leadership.
- Strong understanding of monitoring, logging, and alerting principles and their practical application in production systems.
- Experience with cloud providers such as AWS or GCP, including networking, compute, and storage services.
- Demonstrated capability to manage multiple priorities and stakeholders in a dynamic, growing organization.
- Excellent communication skills, able to convey technical concepts to both technical and non-technical audiences.
Skills & tools
- Languages: Rails, TypeScript, Java, Python, Kotlin, Swift
- Mobile: React Native
- Infrastructure: AWS, GCP, Kubernetes, Terraform, OpenTofu, HashiCorp Vault, AWS Secrets Manager
- Observability: Fluent Bit, OpenTelemetry, Loki, Elasticsearch, Prometheus, Thanos, Datadog
Practical notes
We encourage applications even if your profile does not perfectly match every listed requirement. Our platform is cloud-native and modular, supporting diverse healthcare needs across multiple countries. The role is based in Paris, France, and requires full-time engagement. Candidates must be legally authorized to work in France. This position involves significant travel within the Paris metropolitan area for team collaboration and stakeholder meetings.签证 information will be discussed with successful candidates. The application process will have defined deadlines, and early submission is recommended to ensure full consideration.