Incident Response Manager

at Stripe
📍 Ireland
EUR 88,800-133,200 per year
MIDDLE
✅ Remote

Tech Stack

Communication @ 8 SQL @ 3 Security @ 6 Splunk @ 3

Details

The Incident Ops team is a global 24/7 team responsible for driving incident response and management from detection to resolution. The team works with Reliability Engineering and across the Technology Organization to maintain Stripe's reliability. Incident Response Managers drive incidents to resolution by coordinating cross-functional resources during service outages, critical bugs, security attacks, and other events that significantly impact users. The team also manages external communications and ensures that users and senior management are informed about disruptions.

Responsibilities

  • Act as an on-call Incident Commander, driving and managing incident resolution with urgency, cross-functional collaboration, and accuracy across Engineering, Product, Policy, Risk, Public Relations, Legal, and Executive teams.
  • Lead user-facing incidents across reliability, technical, security, and data privacy domains.
  • Apply a user-first approach to determine impact, provide accurate situation reports, facilitate communications bridges, and ensure timely external communications.
  • Proactively update internal stakeholders and make data-informed decisions while partnering with Engineering, Sales, Support, and other cross-functional teams.
  • Contribute to root cause analysis, conduct post-mortems, identify remediations, and ensure problem management tasks meet service-level agreements and user expectations.
  • Improve incident handling processes, incident management metrics, and tooling based on incident trends and data in collaboration with engineering, product, and operations teams.
  • Contribute to processes, projects, forums, and groups that positively impact and grow team culture.

Requirements

  • 5+ years of demonstrable major incident experience in organizations running mission-critical applications or always-on SaaS environments.
  • Ability to lead multiple incidents concurrently, influence responders with authority and reasoning skills, resolve ambiguous problems, and drive issues to root cause.
  • Intermediate understanding of application development, application architectures, and applications deployed in cloud environments.
  • Good understanding of infrastructure, including physical, virtual, and container-based compute platforms.
  • Quantitative and analytical skills in data manipulation using SQL, Splunk, or other tools.
  • Excellent task management skills, attention to detail, and the ability to remain composed, methodical, and responsive in high-pressure environments.
  • Exceptional written and verbal English communication skills, including the ability to translate complex technical issues for internal and external stakeholders.

Preferred Qualifications

  • Domain expertise in technical, privacy, security, or crisis incidents, with a strong desire to continuously learn about products, technical issues, and systems.
  • Ability to review complex technical details regarding ongoing issues and convey key information to senior stakeholders to support real-time decision-making.
  • Experience with broad user-facing communications, such as status pages and tweets, and targeted communications, such as direct emails and support ticket responses.
  • Familiarity with operating or managing distributed architectures and correlating system behaviors based on known interdependencies.
  • Demonstrated understanding of full-stack development and support.

More jobs at Stripe

Similar jobs