Staff Site Reliability Engineer
Job description
About the role
Arcadia is the AI-powered energy intelligence platform for businesses. We replace fragmented tools and manual workflows with one platform to pay utility bills, buy energy, and advance sustainability - across every location, at enterprise scale. Trusted by Fortune 2000 companies, Arcadia combines unified data, AI-powered analytics, and expert advisory to help enterprise teams save money, mitigate risk, and cut carbon. We deliver this through three comprehensive solutions: Utility Bill Management, automating the entire utility bill lifecycle from data capture and validation to payment processing and auditing; Energy Procurement Advisory, bringing together comprehensive data, AI-powered analytics, market expertise, and a strong partner network to make sophisticated procurement options accessible to all; and Sustainability Reporting, providing verified emissions data with seamless integration into leading sustainability platforms. Tackling the world's most complex energy challenges requires diverse thinking, and we are building teams of people from different backgrounds, industries, and disciplines united by a belief that energy management should be simple, intelligent, and a genuine driver of business value. We are seeking a Staff Site Reliability Engineer (L4) to join our SRE/Platform Engineering team in India. This is a senior technical leadership role focused on engineering leadership through execution, mentorship, and architectural ownership rather than people management. Our India SRE team is growing, and this role is central to that growth as we scale.
Key facts
What you'll do
Own and deliver SRE projects end-to-end, transforming ambiguous requirements into resilient production systems through scoping, design, implementation, testing, rollout, and documentation. Serve as a technical anchor for the India SRE team, conducting design reviews, pairing on complex debugging sessions, and mentoring engineers to build the judgment required to work through ambiguous problems independently. Design and implement infrastructure solutions across AWS, including EKS, VPC, RDS, IAM, CloudWatch, CloudTrail, GuardDuty, S3, CloudFront, Lambda, and SQS, using Terraform and CloudFormation while making careful tradeoffs between speed, reliability, and cost. Lead Kubernetes operations, covering cluster upgrades, capacity planning, CNI troubleshooting, workload scaling, Helm chart packaging, and GitOps deployments, and build runbooks and automation so these practices become repeatable rather than one-off heroics. Evolve CI/CD pipelines across Jenkins with Groovy scripting, GitHub Actions, AWS CodePipeline, ArgoCD, and FluxCD, emphasizing the reduction of manual deployment steps and improvements to rollback safety. Drive observability stack enhancements, delivering the infrastructure and architectural direction necessary for engineering teams to leverage Prometheus, Grafana, and CloudWatch effectively for insight and reliability. Identify and execute FinOps initiatives, finding zombie resources, right-sizing instances, enforcing tagging standards, and presenting cost-reduction recommendations with supporting data. Champion reliability engineering practices by establishing postmortem processes, defining and tracking SLOs and error budgets, and guiding the organization on how to balance velocity with stability. Collaborate closely with US-based SRE leadership on roadmap priorities, incident response, and platform strategy to ensure alignment across time zones and consistent execution of platform vision.
Requirements
You bring depth and experience that enable autonomous execution in the India timezone while collaborating closely with US-based leadership on strategy and incident response. You have a strong background in infrastructure and platform engineering with proven ability to own multi-week SRE projects from problem statement through production deployment. You are skilled in AWS services and infrastructure as code, particularly Terraform and CloudFormation, and understand how to design systems that balance speed, reliability, and cost. You have hands-on Kubernetes experience, including cluster operations, networking, scaling, and GitOps workflows. You are fluent in CI/CD tools and practices, including Jenkins, GitHub Actions, and AWS-native pipelines, and you care about building safe, maintainable deployment processes. You have a solid understanding of observability concepts and tools such as Prometheus, Grafana, and CloudWatch, and you enjoy turning raw metrics into actionable reliability insights. You embrace FinOps principles and have a track record of identifying waste and advocating for cost-aware design in cloud environments. You communicate clearly and constructively in written and verbal form, and you are comfortable mentoring peers through design reviews, debugging escalations, and architectural discussions. You are comfortable working from the Chennai, Tamil Nadu, India location and aligning with the operational cadence of a globally distributed team.
Nice to have
Experience contributing to open source projects relevant to cloud infrastructure, observability, or CI/CD tooling. Familiarity with energy, sustainability, or utility domain concepts. Contributions to public repositories that demonstrate your infrastructure and automation work. A portfolio of runbooks, architecture diagrams, and automation scripts that illustrate your approach to reliability and platform engineering.
Practical notes
You will be based in Chennai, Tamil Nadu, India, and are expected to work during India standard time with occasional overlap into US hours for incident response and cross-team collaboration. This role requires the ability to work independently on complex, multi-week initiatives with minimal day-to-day direction. Travel is generally not required for this position. There may be occasional visa support needs for specific in-person collaboration or business travel, handled on a case-by-case basis. Deadlines for internal applications or specific project milestones, if any, will be communicated separately during the hiring process.