Sr. Incident Manager
Job description
About the role
The Senior Incident Manager at DoubleVerify is responsible for leading and managing the company's Major Incident Management program. This role involves overseeing the entire lifecycle of critical incidents, from initial detection and escalation through resolution and post-incident review. The incumbent will coordinate efforts across multiple technical and business teams to ensure swift and effective responses to high-severity incidents, minimizing customer impact and safeguarding revenue streams. The position requires a strategic thinker capable of making informed decisions under pressure, with a focus on continuous process improvement and operational excellence. The Senior Incident Manager plays a role in maintaining the reliability and stability of DoubleVerify's systems, ensuring that incident response procedures are efficient, well-documented, and aligned with business priorities.
Key facts
The role is based on-site at DoubleVerify's headquarters in New York City, where the company's main operations are located. This is a full-time position requiring on-site presence, with no mention of remote or hybrid work arrangements. The position is part of the Major Incident Management team, which is responsible for handling the most critical and impactful incidents affecting the company's services. The role involves working closely with cross-functional teams including engineering, product, legal, and customer support to coordinate incident response efforts and communicate effectively with stakeholders at all levels.
What you'll do
- Serve as the primary point of contact and accountability for Sev1, Sev2, and Sev3 incidents, ensuring rapid response and resolution.
- Lead real-time incident response efforts by coordinating cross-team activities, ensuring clear communication channels, and facilitating decision-making processes.
- Manage incident communications, including providing timely updates to executive leadership, internal teams, and external stakeholders or clients when necessary.
- Translate complex technical issues into clear, understandable business impacts to facilitate informed decision-making and stakeholder alignment.
- Own and continuously improve the Major Incident Management process, including developing playbooks, escalation procedures, and response workflows.
- Conduct thorough post-incident reviews to identify root causes, lessons learned, and areas for process improvement, ensuring that corrective actions are implemented effectively.
- Monitor key performance indicators such as mean time to recovery (MTTR), incident recurrence rates, and incident trend analysis to drive operational improvements.
- Collaborate with Product, Commercial, and Legal teams to prepare and deliver client communications during major incidents, ensuring transparency and maintaining trust.
- Align incident response activities with business priorities, focusing on minimizing customer impact and protecting revenue streams.
- Develop and implement automation, tooling, and workflow enhancements to streamline incident detection, escalation, and resolution processes.
- Provide leadership and mentorship to junior team members, fostering a culture of continuous learning and operational excellence.
- Participate in capacity planning, risk assessments, and resilience testing to proactively identify vulnerabilities and improve system robustness.
- Maintain detailed incident documentation and reporting to support compliance, audit requirements, and knowledge sharing.
- Stay current with industry best practices, emerging technologies, and incident management methodologies to enhance the company's incident response capabilities.
Requirements
- Over 7 years of experience in Site Reliability Engineering, DevOps, Technical Operations, or Incident Management roles, with a proven track record of managing high-severity incidents.
- Extensive experience leading Sev1 and Sev2 incidents in environments that demand high availability and rapid resolution.
- Demonstrated ability to coordinate cross-functional teams during critical outages, with strong leadership and decision-making skills.
- Deep understanding of distributed systems, cloud platforms such as AWS or GCP, and infrastructure components.
- Proficiency with monitoring, alerting, and incident management tools such as Datadog, Grafana, PagerDuty, or similar platforms.
- Strong analytical skills with the ability to interpret logs, alerts, and diagnostic data to identify root causes quickly.
- Excellent communication skills, capable of conveying technical information clearly to non-technical stakeholders, including executive leadership.
- Experience in developing and refining incident response processes, playbooks, and escalation procedures.
- Ability to work under pressure, prioritize tasks effectively, and maintain focus during high-stress situations.
- Familiarity with ITIL or other operational frameworks is a plus.
- A proactive mindset with a focus on continuous improvement, automation, and operational maturity.
Nice to have
- Background in AdTech, digital media, or related fields, providing contextual understanding of the industry's specific challenges.
- Knowledge of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and reliability engineering principles.
- Experience with automation tools, scripting, or AI-driven incident management solutions to enhance response efficiency.
- Certifications such as ITIL, DevOps, or incident management-specific credentials are advantageous.
- Familiarity with compliance and regulatory standards relevant to incident management and data security.
- Exposure to security incident management and understanding of ITAR (International Traffic in Arms Regulations) compliance.
Skills & tools
- Expertise in distributed systems architecture and cloud infrastructure, particularly AWS or GCP.
- Hands-on experience with monitoring and incident management tools such as Datadog, Grafana, PagerDuty, and similar platforms.
- Strong log analysis and diagnostic skills, with familiarity with tools like Splunk, ELK stack, or equivalent.
- Ability to develop and improve automation scripts and workflows to reduce manual intervention during incidents.
- Excellent organizational skills to manage multiple incidents simultaneously and maintain detailed documentation.
- Effective communication skills to coordinate with technical teams and executive stakeholders.
- Leadership qualities to guide teams through complex incident scenarios and foster a culture of operational excellence.
Practical notes
The salary for this position will be determined based on various non-discriminatory factors, including the candidate's qualifications, experience, skills, and location, as well as internal equity considerations at DoubleVerify. The estimated salary range for this role is between $131,000 and $260,000. In addition to base salary, the role is eligible for bonuses, equity, and comprehensive benefits packages.
Candidates are encouraged to apply even if they do not meet every listed qualification, as diverse experiences and backgrounds are valued. The company emphasizes a commitment to equal opportunity employment and encourages applicants from all backgrounds to consider this role.
This position requires on-site presence at the New York City headquarters; no remote or hybrid arrangements are specified. The role involves working in a fast-paced environment where quick thinking, decisive action, and collaboration are essential.
DoubleVerify values continuous learning and operational excellence, and the successful candidate will be expected to stay current with industry best practices, emerging incident management tools, and evolving technologies to enhance the company's incident response capabilities.
The company provides a collaborative environment where leadership supports professional development and innovation. The role offers an opportunity to impact the reliability and stability of a leading digital media measurement platform, working with a talented team dedicated to operational excellence and customer success.