Senior Site Reliability Engineer
airbyteUSAFull Time3d ago
AWSGCPKubernetesTerraformCI/CDLLMAIETLOperationsSupportRecruitingGrowth
Job description
Senior Site Reliability Engineer at airbyte.
About the role
Join the Data Replication team, a full-stack product group managing millions of weekly sync jobs across various clouds and regions. This role focuses on infrastructure and reliability, ensuring stability for thousands of data use cases. You will establish reliability benchmarks, reduce incidents, and simplify deployment for engineers.
Key facts
What you'll do
- Manage the core infrastructure for the Data Replication platform, including Kubernetes clusters, CI/CD, secrets, networking, and cloud resources on AWS and GCP.
- Collaborate with product engineers to integrate new features reliably with existing infrastructure.
- Enhance observability, alerting, and anomaly detection systems, exploring LLM automation for improvements.
- Improve AI-assisted release and internal tools, such as canary deployments, progressive rollouts, automated release qualification, and rollback, with an emphasis on LLM automation.
- Raise the infrastructure bar for the team by developing self-serve tools, creating runbooks, and guiding engineers to take more ownership of their stack.
Requirements
- 7 or more years of experience in infrastructure, platform engineering, SRE, or DevOps.
- Direct experience managing Kubernetes, Helm, and Terraform in production settings.
- Extensive background with observability tools like Prometheus, Grafana, and Datadog, along with on-call duties.
- Experience owning CI/CD pipelines and developer tools.
- Capable of reading backend code to diagnose system failures and implement correct instrumentation.
- Proficient with AI tools, including LLMs and agentic frameworks, for automation, faster debugging, and reducing manual effort.
- A startup mindset, comfortable with uncertainty, rapid pace, and end-to-end problem ownership.
Nice to have
- Familiarity with data pipelines, replication systems, or ETL/ELT platforms.
- Experience with control plane/data plane architectures or internal developer platforms.
- Knowledge of Airbyte, CDKs, or connector-based architectures.
Skills & tools
- Kubernetes
- Helm
- Terraform
- Prometheus
- Grafana
- Datadog
- AWS
- GCP
- LLMs
- AI agentic frameworks
Practical notes
- Airbyte offers flexible PTO with a recommendation of at least 25 days off annually.
- Comprehensive benefits include 16 weeks of paid parental leave, medical, dental, and vision coverage, 401(k), professional development budget, and commuter benefits.
- Breakfast and lunch are provided in the San Francisco office.
- This role requires 4 days per week onsite in San Francisco.
- Airbyte is an equal opportunity employer and does not accept agency submissions for this position.