Release Engineer
SupabaseRemoteFull Time2w ago
Job description
About the role
Supabase is on the lookout for a Release Engineer who possesses a strong Site Reliability Engineering (SRE) mindset to join our dedicated Release Engineering team. This position is pivotal in safeguarding the integrity, observability, and recovery of our deployment processes and the underlying systems. You will play a crucial role in establishing reliable deployments as the norm, making it the straightforward choice for all teams involved.
Key facts
What you'll do
- Assume responsibility for the reliability of Supabase's deployment and release infrastructure, ensuring it meets established service level objectives and adheres to error budgets.
- Improve pre-production validation processes by standardizing and enhancing existing deployment workflows.
- Fortify disaster recovery strategies by ensuring that environments can be reliably redeployed from scratch when necessary.
- Create and maintain monitoring systems for critical user journeys through synthetic testing, allowing for early detection of regressions.
- Streamline the identification and recovery process for deployment-related incidents, minimizing downtime.
- Engage in on-call rotations, facilitate blameless postmortems, and convert insights into actionable enhancements such as runbooks, alerting mechanisms, and automation.
- Boost the visibility and auditability of deployments, ensuring a transparent record of what was deployed, when, where, and by whom.
- Develop and maintain comprehensive documentation for operational procedures, including emergency protocols and access guidelines.
- Define and monitor service level agreements, service level objectives, error budgets, and DORA metrics to ensure meaningful alerts are in place.
- Guarantee that deployments fail safely and promptly when health checks indicate potential issues.
- Strengthen access controls and emergency procedures to ensure a safe and effective incident response.
- Collaborate closely with product engineering and platform teams to align release practices with overarching reliability objectives.
Requirements
- A minimum of 5 years of experience in Site Reliability Engineering, production operations, platform engineering, or release engineering.
- Demonstrated experience in managing and being on-call for large-scale production systems.
- Strong understanding of service level agreements, service level objectives, error budgets, DORA metrics, and operational key performance indicators, along with experience using observability tools like Prometheus, Grafana, and Alertmanager.
- Proven track record in leading incident response efforts using tools such as incident.io, PagerDuty, or Opsgenie, conducting blameless postmortems, and effectively reducing mean time to detect and mean time to recover.
- Solid operational experience with AWS, including managing multiple accounts, IAM, and VPC in production settings.
- Familiarity with infrastructure-as-code tools such as Pulumi or Terraform, as well as Kubernetes.
- A strong propensity for scripting and automating tasks to eliminate repetitive processes.
- Excellent communication skills, capable of engaging effectively with both technical experts and product engineers.
- Experience working efficiently in asynchronous, globally distributed teams.
- Comfort with navigating uncertainty and continuously improving systems.
Nice to have
- None specified
Skills & tools
- Prometheus
- Grafana
- Alertmanager
- incident.io
- PagerDuty
- Opsgenie
- AWS
- IAM
- VPC
- Pulumi
- Terraform
- Kubernetes
Practical notes
- This position is entirely remote, with no physical office locations. A WeWork membership or co-working allowance is provided to facilitate a productive work environment.
- Employees are eligible for an Employee Stock Ownership Plan (ESOP), allowing them to have equity ownership in the company.
- A technology allowance is available to assist with setting up a comfortable home office.
- Supabase covers 100% of health insurance costs for employees and 80% for their dependents.
- The company organizes annual off-site gatherings for all employees to foster team bonding and collaboration.
- Work hours are flexible and designed to accommodate asynchronous communication.
- An annual education allowance is provided to support professional growth and development.
Join us at Supabase and contribute to our mission of making reliable deployments the standard for all teams. Your expertise will help shape the future of our deployment processes and enhance the overall reliability of our systems.