Senior Site Reliability Engineer
American Unit, IncUSA1w ago
EngineeringReliabilityremotecurated-jd
Job description
Senior Site Reliability Engineer at American Unit, Inc.
About the role
American Unit, Inc. is hiring a Senior Site Reliability Engineer to manage and scale the infrastructure supporting our web, mobile, and AI/ML platforms. You will work with cross-functional teams to improve system availability, automate manual tasks, and maintain high performance for our digital services.
Key facts
What you'll do
- Build and maintain scalable, resilient platform services.
- Establish and monitor SLOs, SLIs, and error budgets.
- Manage incident response, lead root cause analysis, and participate in on-call rotations.
- Develop observability solutions and dashboards to track system health.
- Create self-service tools and automation to increase developer productivity.
- Manage cloud infrastructure on AWS and Azure.
- Write and maintain Infrastructure as Code using Terraform, CloudFormation, Helm, and Kubernetes manifests.
- Optimize CI/CD pipelines and support GitOps workflows.
- Conduct disaster recovery exercises, resiliency testing, and chaos engineering.
- Implement automated remediation and maintain operational runbooks.
Requirements
- Bachelor degree in Computer Science, Engineering, IT, or equivalent professional experience.
- Minimum of 5 years in SRE, DevOps, Platform Engineering, or Cloud Operations.
- Experience managing large-scale distributed systems and customer-facing platforms.
- Proficiency in Linux administration and troubleshooting.
- Hands-on experience with Kubernetes and containerized environments.
- Experience with AWS or Azure cloud platforms.
- Scripting and automation skills in Python and Bash.
- Experience with Terraform, Git, Docker, and CI/CD pipelines.
- Background in managing APIs, microservices, and distributed architectures.
- Proficiency in incident management and performance engineering.
Nice to have
- Experience supporting iOS and Android mobile applications.
- Knowledge of API gateways, CDN technologies, and edge architectures.
- Experience with AI/ML platform operations.
- Familiarity with DevSecOps and cybersecurity standards.
- Relevant certifications such as AWS Certified DevOps Engineer, Solutions Architect, or Kubernetes.
- Experience working within Agile and Product-based operating models.
- Proficiency in Go programming.
Skills & tools
- Cloud: AWS, Azure
- Infrastructure: Kubernetes, Docker, Terraform, CloudFormation, Helm
- Observability: Splunk, Grafana, Prometheus, Datadog, New Relic, OpenTelemetry
- Languages: Python, Bash, Go
- Practices: CI/CD, GitOps, Chaos Engineering, Incident Management
Practical notes
This role requires a consistent onsite presence in Bellevue, WA. Applicants should be prepared to discuss their experience with high-availability production environments and incident leadership during the interview process.