
Engineering Lead - Site Reliability Engineer
Job description
Engineering Lead - Site Reliability Engineer at Heidi Health.
About the role
Heidi is actively seeking an Engineering Lead focused on Site Reliability to safeguard and scale the platform underpinning the world's finest AI Care Partner. In this capacity, you will own the end-to-end reliability and cloud strategy that ensures clinicians can start a patient session without friction at any moment. You will translate the clinical mission of doubling global healthcare capacity into resilient infrastructure decisions that are observable and actionable. This position demands a leader who can operate comfortably at 3am while simultaneously managing the long-term architecture required for 73 million patient visits monthly. You will be instrumental in ensuring that every clinician, whether in Australia or Europe, experiences the platform as consistently as the care they provide.
Key facts
What you'll do
- Lead and expand a platform engineering team responsible for cloud infrastructure and reliability, defining career paths and ensuring sustainable on-call practices.
- Architect and govern Heidi's multi-region cloud footprint, managing compute, networking, data stores, Kubernetes, identity, and infrastructure-as-code with strict clinical data residency requirements.
- Establish and enforce reliability standards by defining SLIs and SLOs that directly inform product roadmaps and error budget consumption.
- Command incident response lifecycles, including severity classification, stakeholder communication, and the execution of blameless post-incident reviews that yield concrete system improvements.
- Collaborate with the Release Manager to refine deployment strategies using feature flags, canary releases, and automated rollback procedures to reduce change failure rates.
- Optimize the developer experience by eliminating manual processes, ensuring engineers can deploy changes without creating operational debt or ticketing overhead.
- Partner with product and security leadership to align platform investments with the needs of health systems, ensuring the platform accelerates delivery rather than constraining it.
- Mentor engineers on observability practices, ensuring that dashboards and alerts reflect true system health and support rapid decision-making during critical events.
Requirements
- Demonstrate significant experience running cloud infrastructure or SRE teams at scale, with a proven ability to lead engineers through complex operational challenges.
- Maintain deep, current hands-on proficiency with a major cloud provider, Kubernetes orchestration, Terraform or equivalent infrastructure-as-code tools, and modern observability platforms.
- Have built and maintained a real reliability practice, including the definition of SLOs, enforcement of error budgets, and leadership of major incident responses.
- Show a track record of designing highly available systems across multiple regions, including management of data residency, failure modes, and disaster recovery testing.
- Exhibit a service-oriented mindset where platform success is measured by the velocity and confidence of product engineering teams, not by technical elegance for its own sake.
- Bring experience in regulated environments such as healthcare or clinical technology, where compliance, auditability, and data sensitivity are non-negotiable.
- Communicate clearly and calmly with executive stakeholders and technical teams alike, translating reliability risks into business impact and mitigation strategies.
- Commit to working in-office in either Sydney or Melbourne, recognizing that the energy of co-location strengthens collaboration and decision-making.
Nice to have
None specified.
Practical notes
- This position is based in-office in Sydney or Melbourne.
- You will be included in the on-call rotation you help design, meaning leaders must be prepared to carry the pager.
- The role requires close collaboration with the Head of Engineering, Security, and product teams, granting real influence over Heidi's global operations.
- No specific working hours are stipulated, but the role expects availability to respond to critical incidents when they arise.
- Travel is not a primary component of the role, though occasional in-person collaboration may be required.
- Visa sponsorship is not mentioned in the source material, so candidates should assume standard eligibility requirements apply.
- The position is full-time and permanent, aligning with the long-term growth trajectory of the company.
This role represents a pivotal opportunity to define the infrastructure backbone of a healthcare organization that is rapidly scaling to serve millions of patients weekly. The successful candidate will not merely manage systems but will lead the philosophy of reliability that allows clinicians to focus on what they do best.