Site Reliability Engineer (SRE)
Job description
About the role
You will own the reliability and performance of Scaleway's storage platforms through advanced automation and monitoring initiatives. You will design and implement infrastructure solutions that ensure high availability and fault tolerance across our global services. Your work will directly support ambitious companies by providing robust and efficient systems that scale with their needs. You will collaborate closely with development and product teams to embed SRE practices from the earliest design phases. This role places you at the heart of a team-driven environment where sharing knowledge accelerates project delivery. You will apply principles of load balancing and energy efficiency to optimize everyday operations. Your contributions will help shape the sovereign cloud of tomorrow by strengthening our network SRE products. You will continuously analyze incidents and drive root cause analysis to prevent future disruptions.
Key facts
What you'll do
- Develop automation tools and frameworks to streamline infrastructure management and reduce manual intervention.
- Build and maintain CI/CD pipelines using Infrastructure as Code best practices to ensure consistent deployments.
- Implement and refine monitoring and alerting systems leveraging OpenMetrics and OpenTelemetry for real-time insights.
- Ensure system reliability through structured incident response and thorough root cause analysis processes.
- Collaborate with developers and product teams to bake resilience into network systems from inception.
- Participate in architecture reviews and provide SRE perspective early in the design to prevent scalability issues.
- Apply principles of fault-tolerance, load balancing, and energy efficiency optimization to enhance performance.
- Share knowledge within the team and broader engineering organization via the SRE Guild to elevate collective expertise.
- Contribute to the reliability and performance of services that serve thousands of customers across multiple sectors.
- Drive operational excellence by minimizing downtime and optimizing fault tolerance in critical storage platforms.
- Support the integration of sustainable practices within infrastructure operations to align with company goals.
- Act as a technical leader during on-call rotations to ensure rapid response to production incidents.
- Evaluate emerging tools and technologies to future-proof our monitoring and automation ecosystems.
- Facilitate cross-team collaboration to align infrastructure strategy with product roadmaps and business objectives.
Requirements
- Strong experience with Infrastructure as Code (IaC) and CI/CD pipelines is essential for modern infrastructure management.
- Solid expertise in Linux systems and production troubleshooting to address complex operational challenges.
- Proficiency with monitoring and logging tools such as OpenMetrics and OpenTelemetry for observability.
- Programming skills in Python, Go, or Rust to build custom automation and tooling solutions.
- Good understanding of network systems is a great bonus including BGP, BGP EVPN, and VXLAN protocols.
- Ability to work autonomously while maintaining alignment with team goals and SRE objectives.
- Clear communication skills in both written and verbal forms to convey technical concepts effectively.
- Comfortable working in English and French to collaborate across multinational teams and offices.
- Willingness to engage in on-call responsibilities and incident response as part of operational duties.
- Commitment to continuous learning and adapting to evolving cloud and infrastructure technologies.
- Strong sense of ownership towards system reliability and performance improvements.
- Ability to collaborate in a fast-paced, international environment with diverse engineering professionals.
- Willingness to contribute to internal guilds and communities of practice focused on best practices.
Practical notes
- Hybrid work: We offer up to 3 days of remote work per week.
- Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
- Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
- Well-being commitments: Whether it's access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.
- International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.
- Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.
- A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.
- A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.
- Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy.
- This role is based in Paris and is a long-term full-time position.
- Candidates must be eligible to work in France without sponsorship.
- Travel is not required for this position.
- Visa sponsorship is not available for this role.
- Please apply only if you meet the hard requirements listed in the qualifications section.
Scaleway is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees regardless of personal characteristics.