Lead Site Reliability Engineer
Job description
About the role
We are seeking a Lead Site Reliability Engineer to enhance the reliability, scalability, and performance of our complex systems. In this role, you will be responsible for designing and implementing strategies to improve system availability, resilience, and efficiency. You will lead incident management efforts, develop automation solutions, and collaborate closely with development, operations, and product teams to ensure our services meet high standards of reliability. Your expertise will help establish best practices in infrastructure management, monitoring, and incident response, fostering a culture of continuous improvement and operational excellence across the organization.
Key facts
What you'll do
- Establish and oversee Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to measure and improve system reliability.
- Design, build, and maintain resilient, scalable systems that achieve an uptime of 99.9% or higher for critical services.
- Lead incident response efforts, coordinate troubleshooting activities, and conduct detailed post-incident reviews to identify root causes and implement corrective actions.
- Automate incident detection, response, and recovery processes through the development of runbooks, scripts, and automation tools.
- Develop and enhance monitoring, alerting, and observability solutions across multiple systems using tools like Open Telemetry, Prometheus, Grafana, and ELK stack.
- Collaborate with development teams to implement reliable, scalable, and high-performance services, ensuring best practices in infrastructure design and deployment.
- Manage infrastructure using cloud services such as AWS, with expertise in Kubernetes, EKS, and Fargate to facilitate container orchestration and deployment.
- Promote Infrastructure as Code (IaC) practices by utilizing tools like Terraform or Pulumi to automate infrastructure provisioning, scaling, and management.
- Participate in chaos engineering initiatives to proactively test system resilience and identify potential failure points before they impact users.
- Conduct capacity planning, performance testing, and forecasting to ensure systems can handle increasing user demand and data volume.
- Engage in on-call rotations, providing 24/7 support for system reliability and incident resolution.
- Implement advanced alerting strategies, anomaly detection, and automated remediation workflows based on metrics and logs.
- Collaborate with cross-functional teams to improve deployment pipelines, automate workflows, and reduce manual intervention.
- Advocate for best practices in security, compliance, and operational procedures to ensure system integrity and data protection.
- Document incident management processes, system architecture, and reliability strategies to facilitate knowledge sharing and continuous learning.
- Stay current with industry trends, emerging tools, and methodologies related to site reliability engineering, cloud infrastructure, and automation.
Requirements
- 3 to 5 years of experience in Site Reliability Engineering or a similar role, with exposure to both cloud-based and on-premises environments.
- Strong understanding of Linux systems, networking, and systems administration fundamentals.
- Hands-on experience working with cloud platforms, particularly AWS, including services like EC2, S3, RDS, and VPC.
- Proficiency in container orchestration using Kubernetes, EKS, and Fargate, with a solid understanding of container lifecycle management.
- Experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki to monitor system health and troubleshoot issues.
- Proficiency in at least one programming language such as Python or Go for developing automation scripts and tooling.
- Strong shell scripting skills using bash or similar scripting languages.
- Familiarity with distributed tracing systems like OpenTelemetry, including integrating tracing, metrics, and logs for comprehensive observability.
- Experience with chaos engineering tools and methodologies, such as Chaos Mesh, Chaos Monkey, or AWS Fault Injection Simulator, to test system resilience proactively.
- Knowledge of CI/CD pipelines, automation frameworks, and deployment strategies including blue-green and canary deployments.
- Excellent problem-solving skills, with the ability to analyze complex issues and implement effective solutions quickly.
- Strong communication skills to coordinate across teams, document processes, and share knowledge effectively.
- Ability to work under pressure, manage multiple priorities, and participate in on-call rotations effectively.
Nice to have
- Experience working with distributed systems at scale, with an understanding of their unique challenges.
- Proficiency in statistical analysis and data-driven decision-making related to system metrics and performance.
- Familiarity with high-performance, low-latency systems and architectures.
- Knowledge of security best practices in cloud and infrastructure environments.
- Experience with writing production-level code, including developing APIs or automation tools.
- Practical experience running chaos engineering drills to validate system resilience.
- Background in designing and implementing resilient architectures using patterns like circuit breakers, retries, and failover mechanisms.
Skills & tools
- A reliability-focused mindset that balances rapid product development with system stability and robustness.
- Deep understanding of SLOs, SLIs, and error budgets to measure and improve system reliability.
- Practical experience with CI/CD pipelines, automation frameworks, and infrastructure as code (IaC).
- Proven track record in incident management, root cause analysis, and conducting effective postmortems.
- Familiarity with deployment strategies such as blue-green, canary, and rolling updates, along with resiliency patterns like circuit breakers and retries.
- Strong scripting and programming skills to develop automation and tooling solutions.
- Ability to analyze metrics and logs to identify issues and optimize system performance.
- Knowledge of security, compliance, and best practices for cloud infrastructure and services.
- Experience working in Agile teams, with a focus on continuous improvement and operational excellence.
Practical notes
Zeta Global is a data-driven marketing technology company that emphasizes innovation and leadership in the industry. Founded in 2007, Zeta combines a vast proprietary data set with artificial intelligence to improve consumer engagement and drive business growth. The company is publicly traded on the New York Stock Exchange under the ticker ZETA. Zeta offers a suite of advanced marketing solutions, including the Zeta Marketing Platform (ZMP) and Athena by Zeta™, a conversational agent designed to optimize marketing outcomes. The organization values technical excellence, collaboration, and continuous learning, providing an environment where engineers can grow their skills and contribute to impactful projects. This role involves working on-site in Bengaluru, Karnataka, India, and requires active participation in incident response and system improvement initiatives. The company encourages a proactive approach to reliability, automation, and innovation, supporting its mission to deliver best-in-class marketing technology solutions.
About the company
 Zeta Global (NYSE: ZETA) is the AI-Powered Marketing Cloud that leverages advanced artificial intelligence (AI) and trillions of consumer signals to make it easier for marketers to acquire, grow, and retain customers more efficiently. Through the Zeta Marketing Platform (ZMP), our vision is to make sophisticated marketing simple by unifying identity, intelligence, and omnichannel activation into a single platform â powered by one of the industryâs largest proprietary databases and AI.