Senior Site Reliability Engineer
Job description
About the role
Become a vital member of the Data Center Business team, where you will contribute to the advancement of Lambda's core cloud platform. This position is centered around enhancing the reliability, scalability, and operational efficiency of systems as the company's user base and infrastructure grow. You will play a crucial role in establishing a robust foundation for AI workloads by leveraging technologies such as Kubernetes, infrastructure automation, observability tools, deployment systems, and incident management strategies.
Key facts
What you'll do
- Manage and scale essential platform services across Lambda's data centers to ensure optimal performance.
- Enhance the reliability of systems involved in compute provisioning, instance lifecycle management, and regional orchestration.
- Create and implement monitoring, alerting, and tracing solutions to track service health, provisioning delays, and failures impacting customers.
- Define service level indicators (SLIs), service level objectives (SLOs), error budgets, and standards for operational readiness.
- Automate the identification and correction of configuration drifts, failed workflows, and orphaned resources to maintain system integrity.
- Develop secure deployment, rollback, and disaster recovery processes utilizing infrastructure as code and GitOps methodologies.
- Design mechanisms for fault isolation to minimize the impact of failures and prevent cascading issues.
- Lead the response to production incidents, conduct postmortem analyses, and implement effective long-term solutions.
- Collaborate closely with teams across Compute, Networking, Storage, Security, and Support to ensure cohesive operations.
- Engage in on-call rotations and work towards improving their sustainability through automation and enhanced tools.
- Mentor fellow engineers and promote higher reliability standards throughout the organization.
Requirements
- A minimum of 7 years of experience in site reliability engineering, infrastructure management, distributed systems, or production software development.
- Significant hands-on experience with Kubernetes in a production setting.
- In-depth knowledge of Kubernetes architecture, including scheduling, networking, resource management, upgrades, and common failure scenarios.
- Familiarity with physical data centers, private cloud setups, hybrid cloud environments, or systems that do not solely depend on managed services.
- Proficient in Terraform or similar infrastructure-as-code solutions.
- Experience in constructing CI/CD or GitOps workflows using tools such as Argo CD, Flux, Helm, or Kustomize.
- Knowledge of observability tools like OpenTelemetry, Prometheus, Grafana, or Datadog.
- Ability to develop production-grade tools using programming languages such as Go or Python.
- Understanding of distributed systems principles, including consistency, retries, idempotence, backpressure, and handling partial failures.
- Proven experience in defining and managing SLIs and SLOs.
- Capability to lead effectively during high-severity incidents.
- A proactive problem-solving mindset focused on engineering solutions and automation for recurring operational challenges.
- Strong communication skills and the ability to collaborate effectively across teams.
- A sense of ownership, sound judgment, and a humble approach to teamwork.
Nice to have
- Experience with AI infrastructure, GPU platforms, or high-performance computing environments.
- Background in managing distributed systems across multiple regions or data centers.
- Familiarity with Kubernetes controllers, operators, custom resource definitions (CRDs), admission control, or scheduler extensions.
- Knowledge of etcd performance, backup, restoration, or disaster recovery processes.
- Experience with Linux systems, container runtimes, cgroups, storage solutions, or networking.
- Exposure to chaos engineering, fault injection techniques, or automated remediation strategies.
- Understanding of Kubernetes role-based access control (RBAC), OpenID Connect (OIDC), workload identity, or certificate management.
- Awareness of compliance frameworks such as SOC 2, ISO 27001, or similar standards.
Skills & tools
Kubernetes, Terraform, Argo CD, Flux, Helm, Kustomize, OpenTelemetry, Prometheus, Grafana, Datadog, Go, Python
Practical notes
This role requires presence in the San Francisco, San Jose, or Bellevue office for four days each week, with Tuesday designated as a work-from-home day. The benefits package includes competitive cash and equity compensation, comprehensive health, dental, and vision coverage for you and your dependents, wellness and commuter stipends for select positions, a 401k plan with a 2% company match for U.S. employees, and flexible paid time off policies.
About the company
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our val