Site Reliability Engineer
Job description
About the role
Your work directly shapes every user experience at Gamma.
Software engineers turn product ideas into working code. Engineers work in small teams, review each other's work, and ship in small batches. Most teams follow agile practices such as sprints and daily standups. Engineers also write tests, fix bugs, and improve performance. The field values clear communication as much as technical skill. Engineers spend part of every week on planning, code review, and debugging, not just writing new code. The ability to explain a technical decision in plain words separates strong engineers from the rest.
Key facts
What you'll do
Production systems across AWS keep reliable, available, and perform under load, guided by platform metrics and alerts.
Automation reduces toil, makes deployments safer, and speeds recovery when incidents occur, implemented through infrastructure-as-code and operational runbooks.
Incident response leads through blameless post-mortems and systemic fixes that prevent repeated failures in production.
Compute, networking, databases, and managed services stay optimized and managed to meet workload demands cost-effectively.
Requirements
The posting states a bachelor's degree requirement. Five or more years of site reliability engineering, DevOps, or systems engineering with deep, hands-on AWS expertise.
Strong programming skills in Python, Go, or TypeScript/Node.js used to build real tools and automation for infrastructure and operations.
Solid experience with infrastructure-as-code tools such as Terraform or CloudFormation and end-to-end observability solutions.
A track record of making systems more reliable through automation, smarter monitoring, and architectural improvements.
Deep understanding of networking, distributed systems, containerization with Docker and Kubernetes, and database performance at scale.
Sharp incident management instincts and debugging skills for navigating complex production failures.
Nice to have
Experience scaling SaaS products to millions of users, or background with Kafka, chaos engineering, or service mesh technologies.
AWS certifications, or experience with security and compliance frameworks such as SOC 2 or ISO 27001.
Practical notes
The team maintains an in-office culture with 4-5 days per week in San Francisco and flexibility to work from home for focused work.
Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
Good to know
Site reliability engineers design and maintain the infrastructure that keeps services available and performant at scale.
Work often involves automation, observability platforms, and incident response to protect user experience.
Common tools in this field include infrastructure-as-code, container orchestration, distributed tracing, and cloud services.
Questions to ask
Good questions to ask the employer in the interview: what does success look like in the first six months, how is the team structured, what is the current biggest challenge, and how are decisions made. Asking about growth paths and the review process is also well received. Employers expect questions, and good ones show preparation.
Career growth
Engineering careers usually progress from individual contributor to senior, staff, and principal levels. Some engineers move into management and lead teams of five to twenty people. Others stay on the technical track. Growth follows demonstrated impact, not tenure alone. A typical engineering ladder has clear levels with defined expectations for scope, quality, and mentorship. Moving up usually requires owning outcomes end to end rather than completing assigned tickets.