Site Reliability Engineer
Job description
About the role
You will own the design and execution of infrastructure automation that directly underpins the reliability of our global SaaS platform serving millions of locations. You will implement and maintain observability frameworks that provide deep insight into system behavior and performance for the engineering organization. This role requires you to build and evolve CI/CD pipelines that enable rapid, safe, and predictable deployments across our environment. You will collaborate closely with product engineers to embed reliability principles into service designs from the outset of development. You will participate in a global on-call rotation to provide technical leadership and rapid response for critical production incidents. You will drive the elimination of operational toil through continuous process improvement and automation initiatives. You will manage and optimize cloud infrastructure to balance performance, resilience, and cost efficiency.
Key facts
What you'll do
- Architect and implement infrastructure using infrastructure-as-code (IaC) tools to ensure consistency and scalability across environments.
- Develop and deploy comprehensive observability solutions, including metrics, traces, and logs, to monitor service health and performance.
- Design, build, and maintain robust CI/CD pipelines that automate testing and deployment workflows for reliability engineering teams.
- Partner with cross-functional product teams to influence system architecture, ensuring operational resilience and scalability are core considerations.
- Execute and refine incident response procedures as part of a global on-call rotation to provide 24/7 coverage for production systems.
- Analyze operational workflows to identify repetitive tasks and automate them to reduce manual intervention and human error.
- Optimize cloud resource utilization and configurations to improve cost efficiency without compromising performance or availability.
- Define and track service-level objectives and indicators to measure and drive reliability goals across the platform.
- Collaborate with networking and security teams to ensure infrastructure configurations adhere to best practices and compliance standards.
- Evaluate and integrate new tools and technologies that enhance the stability and efficiency of the production environment.
- Mentor junior engineers on SRE practices, contributing to the growth of the team and the broader engineering culture.
- Conduct post-incident reviews to identify root causes and implement preventative measures to avoid future occurrences.
- Manage capacity planning activities to ensure infrastructure scales to meet current and future business demands.
- Serve as a technical liaison between development, operations, and infrastructure teams to streamline communication and execution.
Requirements
- Hold a Bachelor's degree in Computer Science, Engineering, or a related field, or possess equivalent practical experience.
- Bring 5+ years of professional experience in a Site Reliability Engineer, DevOps Engineer, or similar infrastructure-focused role.
- Demonstrate expertise in one or more general-purpose programming languages such as Python or Go.
- Show hands-on experience with major public cloud computing platforms like AWS or GCP.
- Exhibit proficiency in configuration management and infrastructure-as-code tools such as Terraform or Salt.
- Display proficiency in Kubernetes-based container orchestration environments and ecosystem management.
- Demonstrate experience with monitoring, logging, and visualization tools such as Prometheus, Grafana, and OpenSearch.
- Communicate technical concepts clearly and effectively to both technical and non-technical stakeholders.
- Have a strong understanding of networking fundamentals, including load balancing, DNS, and firewall concepts.
- Show familiarity with security best practices and compliance considerations for cloud environments.
- Understand database technologies, including both SQL and NoSQL databases, and their operational characteristics.
- Exhibit strong problem-solving skills and the ability to debug complex issues in distributed systems.
- Possess excellent written and verbal communication skills for collaboration across global teams.
- Show a proven track record of working in fast-paced, agile environments with ambiguous requirements.
- Demonstrate ownership and accountability for production systems and their reliability outcomes.
- Have experience with scripting and automation to streamline operational tasks.
- Show commitment to learning and adapting to new technologies in a rapidly evolving landscape.
Nice to have
- Hands-on experience with large-scale distributed systems spanning multiple regions and availability zones.
- Deep knowledge of networking and security best practices, including encryption and identity management.
- Familiarity with a wide range of database technologies, including relational, document-based, and time-series stores.
- Experience with chaos engineering practices and resilience testing methodologies.
- Background in software development lifecycle management and agile methodologies.
- Exposure to SaaS product delivery models and multi-tenant architecture patterns.
- Understanding of cost optimization strategies for cloud billing and resource allocation.
- Experience with logging aggregation platforms and log analysis techniques.
- Knowledge of service mesh technologies and their application in microservices architectures.
Practical notes
This role is based in Hyderabad, India, and operates from a work-from-office location. The position requires participation in a global on-call rotation, which may include nights and weekends to support critical systems. Travel requirements are minimal, though occasional in-person team gatherings may be possible based on business needs. Candidates must be eligible to work in India without visa sponsorship for this position. The role is full-time, and successful candidates are expected to commit to the established work schedule and on-call responsibilities.