Senior Principal Site Reliability Engineer
Job description
About the role
This role designs and operates large-scale resilience and experimentation platforms for global trading and financial systems. You will lead architecture and delivery for enterprise chaos engineering across hybrid Kubernetes and cloud infrastructure. The position drives production safety, validation, and reliability improvements for critical products used by millions of users. Platform and DevOps engineers build the systems that run everything else, managing infrastructure, CI/CD pipelines, observability, and reliability. Systems thinking is the core skill required, and platform teams are measured by developer velocity and system reliability. Most companies run on-call rotations, and understanding incident response is an integral part of this position. Chaos engineering focuses on balancing aggressive validation with strict risk controls in production environments.
Key facts
What you'll do
Enterprise-grade chaos platforms are shaped through architecture and code, enabling multi-cluster, multi-region, and multi-environment resilience across test and production. Blast radius controls, automated rollback, and real-time monitoring form a closed loop where faults are injected, impact is observed, and pass or fail is determined. Daily and periodic experiments validate stability, while large-scale disaster recovery drills verify cross-AZ and cross-region failover strategies. Resilience metrics quantify system health, and findings drive improvement plans that technology teams execute. Technology foundations are evaluated and selected, and practices, playbooks, and mentorship equip SRE and developer teams to run chaos engineering safely and effectively. You will implement production-grade fault injection using tools such as Chaos Mesh or Litmus, with deep knowledge of CRDs and operators in Kubernetes. Backend systems are designed and architected using languages such as Go, while failure modes like network partitions, split-brain, and cascading failures are well understood. Observability platforms built on Prometheus, Grafana, Thanos, and OpenTelemetry integrate with chaos workflows, and technical documentation and solution designs clearly communicate safety controls and experiment outcomes.
Requirements
Eight or more years of backend or infrastructure engineering are required, with at least three years focused on chaos or stability practices in production. You must possess production-grade fault injection experience using tools such as Chaos Mesh or Litmus, with deep knowledge of CRDs and operators in Kubernetes. Backend systems design and architecture must be performed using languages such as Go, and you must have a thorough understanding of failure modes like network partitions, split-brain, and cascading failures. Observability platforms built on Prometheus, Grafana, Thanos, and OpenTelemetry must be integrated with chaos workflows in your previous roles. You must be capable of creating technical documentation and solution designs that clearly communicate safety controls and experiment outcomes. Professionals in this field typically rely on observability pipelines, automation frameworks, and infrastructure-as-code to manage experiments. Eight or more years of total experience are required, with at least three years in chaos or stability practices in production. A Bachelor's degree is required, and Malaysia-based roles are subject to local employment eligibility and visa requirements.
Practical notes
Malaysia-based roles are subject to local employment eligibility and visa requirements. Onsite expectations and team details are confirmed during the application process. Typical interview steps include platform interviews that usually consist of an infrastructure scenario, a scripting or coding exercise, and operational questions. Candidates may be asked to design a deployment pipeline or debug an outage, and incident experience along with an automation mindset will be tested. Interviewers often ask about a past outage and how you handled it, looking for structured post-incident thinking rather than heroics. Good questions to ask the employer in the interview include what does success look like in the first six months, how is the team structured, what is the current biggest challenge, and how are decisions made. Asking about growth paths and the review process is also well received, as employers expect questions and good ones show preparation. Career growth in platform roles progresses from engineer to senior, staff, and platform lead, with some individuals moving into SRE leadership or cloud architecture. Breadth across networking, storage, and reliability becomes more important at senior levels, and platform careers reward breadth and calm under pressure. Experience automating your own work is the strongest signal for senior roles.