Site Reliability Engineer
Job description
About the role
You will own the design and execution of reliability initiatives that directly impact the availability and performance of Scaleway's global cloud infrastructure. You will build and evolve automation that provides deep observability into production systems, turning raw data into actionable insights for faster incident response. This role requires you to troubleshoot complex, high-impact issues alongside product teams while maintaining strict service continuity standards. You will actively participate in on-call rotations to ensure rapid detection and remediation of production incidents across all environments. You will champion the implementation of resilient and scalable infrastructure patterns that align with Scaleway's commitment to technical excellence. You will maintain clear and up-to-date documentation for all tools and operational procedures to ensure long-term maintainability. Finally, you will collaborate closely with development teams to embed reliability practices throughout the entire product lifecycle.
Key facts
What you'll do
- Architect and implement robust monitoring and alerting frameworks using Prometheus and Grafana to detect anomalies before they impact customers.
- Automate the diagnosis and remediation of production incidents by developing custom tooling in Go and Python to reduce manual intervention.
- Lead complex troubleshooting efforts for critical production issues, coordinating with cross-functional engineering teams under pressure.
- Design and operate scalable observability pipelines that aggregate logs, metrics, and traces to provide a unified view of system health.
- Manage the full lifecycle of infrastructure components across development, staging, and production environments to ensure consistency.
- Drive the adoption of Infrastructure-as-Code practices using tools like Ansible and AWX to provision and manage cloud resources securely.
- Optimize the performance and resiliency of containerized workloads and bare metal infrastructures deployed on Scaleway's platform.
- Document all operational procedures and runbooks to ensure clarity, consistency, and efficient knowledge transfer.
- Partner with product teams to define service level objectives and ensure infrastructure readiness for new feature rollouts.
- Contribute to the evolution of core systems based on feedback from production incidents and regular post-mortem analysis.
- Implement security best practices and ensure that all infrastructure changes comply with organizational and regulatory requirements.
- Mentor junior engineers by sharing knowledge and fostering a culture of continuous learning and improvement within the SRE team.
Requirements
- Hold a Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field.
- Bring a minimum of three years of professional experience in Site Reliability Engineering or a similar operations role.
- Demonstrate expert-level proficiency in Bash and Python scripting to automate complex operational tasks.
- Show extensive hands-on experience with Linux systems, specifically Ubuntu and Debian, including performance tuning and troubleshooting.
- Possess a deep understanding of networking fundamentals such as TCP/IP, DNS, BGP, load-balancing, IPv6, and firewall management.
- Have direct experience with cloud environments and infrastructure models including bare metal servers, virtual machines, containers, and orchestration platforms.
- Exhibit strong familiarity with monitoring, logging, and tracing tools such as Prometheus, Grafana, and Elastic stacks.
- Display advanced competence in managing relational databases, particularly PostgreSQL, including replication and backup strategies.
- Prove experience with CI/CD pipelines and version control systems, with a focus on GitLab workflows and automation.
- Fluency in English is required for both written communication and verbal collaboration with international teams.
Nice to have
- Experience with container orchestrators such as Kubernetes and service mesh technologies.
- Familiarity with serverless platforms and edge computing infrastructures.
- Knowledge of compliance frameworks and disaster recovery planning.
Practical notes
This is a full-time, long-term position based in Paris. The role follows a hybrid work model with up to three days of remote work per week. Candidates must be willing to participate in rotational on-call schedules to support 24/7 service availability. Travel is not required for this position. Scaleway is committed to equal opportunity hiring and welcomes candidates from diverse backgrounds.