
Incident Management
Job description
About the role
The Incident Commander will own the full lifecycle of major and high-severity incidents affecting trading and financial platforms. You will lead real-time response efforts during outages, coordinating technical and business stakeholders under strict time constraints. This role requires driving root cause analysis and ensuring that corrective actions are implemented to prevent recurrence. You will enforce SLA governance and ensure all incident documentation meets audit and regulatory standards. The position involves owning disaster recovery and business continuity exercises to ensure operational resilience. You will be responsible for maintaining a clear line of sight between technical incidents and business impact. This role will partner closely with engineering, infrastructure, and risk teams to align on priorities and improvements. You will serve as the primary escalation point and decision maker during critical service disruptions.
Key facts
What you'll do
- Lead and coordinate end-to-end incident management for critical fintech services, ensuring rapid stabilization and recovery.
- Serve as the primary Incident Commander during major outages, directing cross-functional response and communication activities.
- Own problem management activities, driving deep root cause analysis and tracking corrective and preventive actions to closure.
- Manage and monitor incident lifecycles, enforcing SLA adherence and maintaining comprehensive incident documentation.
- Oversee data integrity across ITSM repositories and tools, ensuring accurate configuration and change records.
- Provide detailed downtime and availability reporting, including impact analysis and recovery verification.
- Generate monthly Service Management Risk Reports to highlight operational risks and track remediation progress.
- Ensure service management processes align with the Service Catalog and evolving business expectations.
- Participate in audit readiness activities, including SOC Type 1 & Type 2 examinations and regulatory assessments.
- Prepare and deliver audit evidence, control testing results, and compliance dashboards to support external and internal reviews.
- Coordinate with Operations, Engineering, Risk, and Compliance teams to close audit findings and track corrective actions.
- Run weekly Release Washups and engage in Change Advisory Board meetings to discuss changes affecting stability and risk.
- Support automation initiatives aimed at reducing MTTR, enhancing monitoring, and streamlining incident workflows.
- Develop, test, and maintain Disaster Recovery Plans, leading periodic DR testing and validation exercises.
- Operationalize the Business Continuity Plan across IT and business units to ensure continuity readiness.
Requirements
- Must have 8+ years of experience in Incident Management, Problem Management, and IT Operations within fintech, banking, or trading environments.
- Must have proven experience as Major Incident Manager or Incident Commander handling real-time high-severity events.
- Must have strong knowledge of ITIL processes, including Incident, Problem, Change, Knowledge, and Configuration Management.
- Must have hands-on experience with ITSM tools such as ServiceNow, JSM, HPSM, and CMDB.
- Must have experience in reporting and analytics, including availability metrics, SLA compliance, and RCA trend analysis.
- Must have experience supporting audit activities, including process documentation and engagement with internal and external auditors.
- Must demonstrate excellent communication and stakeholder management skills, with the ability to lead under pressure.
- Must have a good understanding of high-level software architecture as well as infrastructure components supporting financial services.
Nice to have
- Knowledge of trading systems, order management platforms, or exchanges.
- Experience in running Splunk queries and basic troubleshooting.
- ITIL v3 or v4 certification or equivalent qualification.
- Prior experience in global 24×7 financial services environments.
- Familiarity with SOC audit frameworks and regulatory compliance in financial services.
- Preference for working night shifts to support continuous operations.
Practical notes
This is a contract role based in London with remote work options. The engagement requires availability during night shifts to support global operations. Candidates must be eligible to work in the United Kingdom. The role may involve participation in audit calls and walkthroughs at scheduled times. Travel is not expected as part of this role.