Incident Management & Resilience Lead
Job description
About the role
Incident response and prevention define this role, with ownership across engineering, security, and support to set and enforce a consistent resilience standard. Success is measured by fewer and smaller incidents across the organization.
Most organizations organize this kind of work into small teams with a clear owner for each project. People are expected to set priorities, share progress, and ask for help early. Written communication matters because decisions are often recorded and revisited. Successful professionals in any field share a few habits: they take notes, keep commitments, and communicate in writing. Teams value people who make their work easy to review and build on.
Key facts
What you'll do
Ownership of incident management is defined, implemented, and maintained across engineering, security, and support to set a consistent resilience standard.
High-impact incidents are resolved in the moment with clarity, coordination, and communication maintained under sustained pressure.
Postmortems are executed with rigor so their findings lead to concrete changes that reduce future incident impact.
Observability tooling such as Datadog and Honeycomb is used to investigate live issues, supported by AWS, Terraform, Kafka, and Redis failure modes.
A consistent resilience bar is held across all product teams, with engineering managers aligned to the same standards.
Engineers are coached to run incidents, communicate transparently, and own outcomes without reliance on top-down authority.
Requirements
A degree is required as stated in the metadata.
You have built or materially elevated incident management capability in a production environment at scale.
You can read code and telemetry to lead real investigations, not just facilitate calls.
You can sustain composure during prolonged, messy incidents and reduce their frequency and severity.
You can enforce standards across teams you do not manage while preserving relationships and clarity.
You prefer a blank page where you define incident management rather than inheriting a mature function.
You gain more from preventing incidents than from heroic live resolutions.
You thrive with flat autonomy and minimal scaffolding, and you value cross-team persuasion as part of the work.
You are comfortable working fully remotely with asynchronous communication and scheduled overlap.
Nice to have
Experience with observability tooling such as Datadog and Honeycomb.
Familiarity with AWS, Terraform, Kafka, and Redis.
Exposure to Ruby on Rails, PostgreSQL, and similar stack components.
Practical notes
This role is based in the ANZ region with no sponsorship for positions outside this area.
All applications receive a response, and the process is driven by velocity and clarity.
Typical interview steps
Interviews in this field typically test judgment, communication, and fit as much as technical skill. Be ready to describe a challenge you faced, what you did, and what you learned. Most companies value honest answers over polished ones. Whatever the role, preparation is visible. Reviewing the company's product, the job description, and your own past work before the conversation is the strongest step.
Good to know
Incident management at this scale focuses on prevention, clear communication, and ownership.
The platform supports thousands of organizations that depend on continuous software delivery.
The role operates with high autonomy and flat structures to enable fast decisions.
Remote work is practiced asynchronously with deliberate overlap for collaboration.
The tech stack includes Ruby on Rails, PostgreSQL, Kafka, and Redis on AWS.
Career growth
Career growth comes from taking on harder problems and making your work visible. Look for opportunities to own outcomes, mentor others, and learn the business beyond your team. Growth in any career comes from scope, results, and reputation. Take on work that is slightly uncomfortable, and make sure others can see the outcomes.