Lead Site Reliability Engineer
OptomiUSA3w ago
EngineeringReliabilityremotecurated-jd
Job description
Lead Site Reliability Engineer at Optomi.
About the role
Optomi is hiring a Lead Site Reliability Engineer to manage and scale a generative AI platform for a major entertainment company. You will guide technical strategy for cloud infrastructure and reliability while staying active in daily engineering tasks.
Key facts
What you'll do
- Architect and operate infrastructure for generative AI applications.
- Manage production Kubernetes clusters, including capacity and reliability.
- Create and standardize Helm charts for platform deployments.
- Build Infrastructure-as-Code solutions using Terraform.
- Design CI/CD pipelines with a focus on Harness.
- Execute deployment strategies like blue/green, canary, and GitOps.
- Enhance observability through logging, monitoring, and distributed tracing.
- Maintain a 99.99 percent uptime SLA for critical services.
- Troubleshoot distributed systems and cloud networking.
- Mentor team members and define engineering standards.
Requirements
- 7+ years of experience in SRE, DevOps, or Platform Engineering.
- Hands-on production Kubernetes administration.
- Proficiency in building and maintaining Helm charts.
- Expert-level Terraform and Infrastructure-as-Code skills.
- Professional experience with Harness for CI/CD.
- Multi-cloud background across GCP, AWS, and Azure, with a preference for GCP.
- Scripting proficiency in Python and Bash.
- Advanced YAML configuration skills.
- Production experience with PostgreSQL, Redis, Kafka, MongoDB, and Vault.
- Experience with CI/CD platforms like GitHub Actions, GitLab CI, Jenkins, or Azure DevOps.
- Background in observability tools such as Splunk, OpenTelemetry, Prometheus, and AppDynamics.
- History of leading technical projects and mentoring engineers.
- Experience working in Agile or Scrum environments.
Skills & tools
- Kubernetes, Terraform, Harness, GitOps, Python, Bash, YAML, Helm
- GCP, AWS, Azure
- PostgreSQL, Redis, Kafka, MongoDB, Vault
- Splunk, OpenTelemetry, Prometheus, AppDynamics
- GitHub Actions, GitLab CI, Jenkins, Azure DevOps
Practical notes
This role requires the ability to communicate complex technical concepts to both technical and non-technical stakeholders.