Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 8
SQL @ 3
Security @ 6
Splunk @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Incident Ops team is a global 24/7 team responsible for driving incident response and management from detection to resolution. The team works with Reliability Engineering and across the Technology Organization to maintain Stripe's reliability. Incident Response Managers drive incidents to resolution by coordinating cross-functional resources during service outages, critical bugs, security attacks, and other events that significantly impact users. The team also manages external communications and ensures that users and senior management are informed about disruptions.
Responsibilities
- Act as an on-call Incident Commander, driving and managing incident resolution with urgency, cross-functional collaboration, and accuracy across Engineering, Product, Policy, Risk, Public Relations, Legal, and Executive teams.
- Lead user-facing incidents across reliability, technical, security, and data privacy domains.
- Apply a user-first approach to determine impact, provide accurate situation reports, facilitate communications bridges, and ensure timely external communications.
- Proactively update internal stakeholders and make data-informed decisions while partnering with Engineering, Sales, Support, and other cross-functional teams.
- Contribute to root cause analysis, conduct post-mortems, identify remediations, and ensure problem management tasks meet service-level agreements and user expectations.
- Improve incident handling processes, incident management metrics, and tooling based on incident trends and data in collaboration with engineering, product, and operations teams.
- Contribute to processes, projects, forums, and groups that positively impact and grow team culture.
Requirements
- 5+ years of demonstrable major incident experience in organizations running mission-critical applications or always-on SaaS environments.
- Ability to lead multiple incidents concurrently, influence responders with authority and reasoning skills, resolve ambiguous problems, and drive issues to root cause.
- Intermediate understanding of application development, application architectures, and applications deployed in cloud environments.
- Good understanding of infrastructure, including physical, virtual, and container-based compute platforms.
- Quantitative and analytical skills in data manipulation using SQL, Splunk, or other tools.
- Excellent task management skills, attention to detail, and the ability to remain composed, methodical, and responsive in high-pressure environments.
- Exceptional written and verbal English communication skills, including the ability to translate complex technical issues for internal and external stakeholders.
Preferred Qualifications
- Domain expertise in technical, privacy, security, or crisis incidents, with a strong desire to continuously learn about products, technical issues, and systems.
- Ability to review complex technical details regarding ongoing issues and convey key information to senior stakeholders to support real-time decision-making.
- Experience with broad user-facing communications, such as status pages and tweets, and targeted communications, such as direct emails and support ticket responses.
- Familiarity with operating or managing distributed architectures and correlating system behaviors based on known interdependencies.
- Demonstrated understanding of full-stack development and support.
More jobs at Stripe
Full Stack Engineer, Enterprise Engineering
Stripe · Toronto, Canada
CAD 135,200-202,800 per year
Security GRC Analyst/Program Manager, Bridge
Stripe · United States
USD 190,400-285,600 per year
Program Manager, Deal Operations
Stripe · United States
USD 134,200-201,200 per year
Specialist Solutions Architect, Money Management
Stripe · London, United Kingdom
GBP 120,500-180,700 per year
Product Support Specialist - Bridge
Stripe · United States
USD 148,300-222,500 per year
Similar jobs
Incident Response Manager - Security
Stripe · Ireland, Dublin, Ireland
EUR 89,100-133,700 per year
Solution Designer
ING · Milan, Italy
EUR 50,000 per year
Senior Systems Engineer - Windows/SQL Services
Bloomberg · New York City, United States
USD 130,000-225,000 per year
Security Engineer, ACE
Stripe · Dublin, Ireland
EUR 85,000-127,600 per year
Staff Security Engineer, Abuse Control
Stripe · London, United Kingdom, Dublin, Ireland
EUR 132,000-198,000 per year
Staff Security Engineer, Abuse Control
Stripe · London, United Kingdom, Dublin, Ireland
EUR 132,000-198,000 per year
Security Engineer - Proactive Threat
Stripe · Dublin, Ireland
EUR 112,200-168,200 per year
Forward Deployed Security Engineer
Stripe · Dublin, Ireland
EUR 132,000-198,000 per year