Senior CloudOps Engineer
cloudzeroUSAFull Time2d ago
PythonGoAWSAzureGCPKubernetesTerraformGitHub ActionsLLMClaudeGeminiAI
Job description
Senior CloudOps Engineer at cloudzero.
About the role
CloudZero is rapidly expanding its platform to handle increasing data complexity and customer growth. This role focuses on ensuring the reliability and performance of our most demanding real-time data ingestion systems. You will be instrumental in maintaining the stability of a unique serverless architecture that processes billions of events daily.
Key facts
What you'll do
- Establish and maintain reliability practices for CloudZero's real-time data ingestion, including cross-team service level objectives (SLOs) and architectural improvements.
- Approve critical system paths before deployment and intervene when error budgets are exceeded.
- Implement system instrumentation to facilitate rapid failure detection and data-driven debugging.
- Develop reliability tools such as load generators, fault-injection frameworks, and SLO libraries.
- Write and maintain production-grade Python code for shared libraries, internal services, and automation.
- Design and manage CloudFormation and SAM modules for provisioning reliable and cost-efficient cloud resources.
- Automate operational tasks like deployments, scaling, and backups to minimize manual intervention.
- Enhance the effectiveness of existing autonomous agents and integrate systems with AI tooling.
- Collaborate with product engineering teams to design resilient services and optimize for cost and performance.
- Drive the adoption of best practices and instrumentation across engineering teams.
Requirements
- Demonstrated expertise in production Python as your primary programming language, including ownership, testing, and maintenance at scale.
- Experience defining and implementing a Service Level Objective (SLO), including decisions on what not to alert.
- Practical experience operating asynchronous, event-driven systems and understanding concepts like back-pressure, consumer lag, and partial failure.
- Proven ability to lead and implement a significant reliability or platform change across a team without direct authority.
- At least 5 years of experience building and operating distributed systems in AWS, with a focus on owning reliability outcomes.
- Hands-on experience with Infrastructure as Code using CloudFormation and SAM, or deep transferable knowledge of Terraform or Pulumi.
- Direct experience instrumenting systems with monitoring tools like Sumo Logic, Datadog, Prometheus, or Splunk.
- Ability to effectively debug production issues under pressure.
- Interest in working with frontier AI models such as Claude, Codex, or Gemini.
- Commitment to thoughtful, reliable system design over reactive problem-solving.
- Strong documentation habits to ensure long-term team clarity and system stability.
- Skill in explaining complex technical concepts to non-technical audiences.
- Eagerness to manage multiple areas simultaneously, deep dive into critical issues, and then standardize solutions.
Nice to have
- Experience building chaos engineering or load testing practices.
- Familiarity with internal developer portals like Cortex or Backstage.
- Background in test automation or ephemeral test environments.
- Experience with GitHub Actions at scale.
- Development of LLM-backed tools used by engineers daily.
Skills & tools
- Python
- AWS (MSK, Kinesis, SQS, Step Functions)
- CloudFormation
- AWS SAM
- Sumo Logic
- Datadog
- Prometheus
- Splunk
- Kafka
- Terraform (transferable)
- Pulumi (transferable)
- Claude
- Codex
- Gemini
Practical notes
CloudZero processes 14 trillion billing events annually and is a listed partner on Anthropic's cost and usage API. The company has raised over $119 million. This role is a hybrid position based in Boston. The on-call rotation is light, with a weekly schedule for shared infrastructure and feature teams handling their own services.