
Staff Cloud Reliability Engineer
Job description
About the role
The Staff Cloud Reliability Engineer owns the design and execution of reliability initiatives for Viant's cloud native platform. You will architect guardrails and automation that enable scalable, secure, and observable operations across our infrastructure. This role demands deep systems engineering expertise to translate business requirements into robust technical solutions. You will act as a key architect for containerized and serverless workloads, ensuring best practices are embedded into the platform. You will partner closely with development teams to elevate code quality and operational maturity. The position requires active participation in oncall rotations to resolve incidents and drive continuous improvement. You will define and implement strategies that balance velocity with stability.
Key facts
What you'll do
Architect and deploy automation tools to manage cloud based infrastructure at scale.
Containerize legacy and modern applications to optimize portability and resource efficiency.
Curate and maintain a catalog of internal frameworks and templates to accelerate project delivery.
Perform level of effort analysis and scoping for complex initiatives involving multi team dependencies.
Act as a technical liaison supporting development and operations teams during critical milestones.
Debug and enhance existing codebases to improve performance, resilience, and maintainability.
Implement monitoring strategies to ensure comprehensive observability of application and system performance.
Participate in rotational oncall duties to address production incidents and support maintenance windows.
Champion the adoption of infrastructure as code patterns across engineering organizations.
Evaluate and integrate emerging technologies to future proof the platform and reduce technical debt.
Collaborate with security and compliance stakeholders to enforce governance across cloud resources.
Optimize CI/CD workflows to accelerate feedback loops and improve deployment frequency.
Document operational procedures and runbooks to ensure continuity and knowledge transfer.
Mentor junior engineers on cloud native principles and troubleshooting methodologies.
Requirements
8+ years of professional experience in DevOps or Site Reliability Engineering roles.
3+ years of hands on experience with Linux administration in production environments.
3+ years of experience managing workloads on at least one major cloud provider such as AWS or Google Cloud.
3+ years of designing and operating serverless architectures using platforms like AWS Lambda or Google Cloud Functions.
3+ years of practical experience with Docker and Kubernetes for container orchestration.
3+ years of writing and managing infrastructure as code with Terraform across multiple environments.
Demonstrated ability to create CI/CD pipelines using GitHub Actions and related tooling.
Proficiency in at least one high level programming language such as Python or GoLang.
Experience querying and analyzing data stores using SQL and Google BigQuery.
Strong understanding of networking concepts, identity management, and security best practices.
Ability to work independently and collaboratively in a fast paced, agile environment.
Excellent written and verbal communication skills for cross functional collaboration.
Willingness to adapt to evolving priorities and shifting project requirements.
Commitment to maintaining high standards of code quality and documentation.
Nice to have
Experience with monitoring and alerting platforms such as Prometheus, Grafana, or Datadog.
Knowledge of cost optimization techniques for cloud billing and resource utilization.
Familiarity with log aggregation tools like Elasticsearch, Loki, or Splunk.
Background in building internal developer platforms or self service tooling.
Understanding of compliance frameworks relevant to advertising technology.
Practical notes
This role requires participation in an oncall rotation for timely response to system incidents and maintenance needs.
Base compensation range is listed as $180,000 - $200,000 in accordance with California law. Final title and compensation will be based on several factors including work experience and education.
Employment is contingent upon eligibility to work in the United States.
This position is open to candidates located in Irvine, California, and Los Angeles, California.