Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 4
Datadog @ 6
Jira @ 6
Leadership @ 6
Observability @ 4
Reporting @ 4
SRE @ 8
Salesforce @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI’s User Operations team supports customers adopting AI, resolves complex issues, provides technical guidance, and helps customers maximize the value of OpenAI products. The team works closely with Sales, Technical Success, Product, Engineering, and other groups to deliver an exceptional customer experience at scale.
This is a hands-on player-coach role responsible for building and operating OpenAI’s Incidents & Escalations function within User Operations. The role involves setting the operating model, participating in active incidents and urgent escalations, coordinating with on-call teams, driving ownership, managing communications, and ensuring incidents progress through resolution and post-incident closure.
The position coordinates cross-functional responders across Engineering, Infrastructure, Support Delivery, Product, and Go-To-Market. Responsibilities include maintaining timelines, clarifying ownership, escalating when necessary, managing internal and external communications, and providing status page updates when required. The role also owns escalation tracking, triage, mitigation, resolution, retrospectives, root-cause identification, action-item follow-through, trend analysis, and process improvements.
This role is based in San Francisco and follows a hybrid work model requiring three days in the office per week.
Responsibilities
- Participate in an on-call rotation and serve as the active incident lead during live incidents and urgent escalations.
- Own alert intake and triage across support, safety, customer, and service-impacting issues.
- Assess severity, scope, and impact and initiate the appropriate response path.
- Page and coordinate Engineering, Infrastructure, Support Delivery, Product, Legal, Policy, Go-To-Market, and other teams as needed.
- Lead incident response calls, manage timelines, clarify roles, and keep responders focused and unblocked.
- Set internal guidelines for incident communications and own internal updates, executive briefings, customer-facing updates, and external status page updates where required.
- Maintain situational awareness across customer-facing incidents and parallel workstreams.
- Create and operate processes for monitoring, processing, mitigating, and resolving critical escalations, including formal closure and handoff.
- Identify root causes, lead retrospectives, and coordinate corrective actions with accountable teams.
- Track corrective actions to closure and focus follow-through on the best possible customer outcome.
- Identify recurring operational issues, escalation patterns, and product or process gaps.
- Partner with Engineering, Infrastructure, Product, and Support leaders to reduce repeat issues and improve readiness.
- Improve incident response processes, severity frameworks, playbooks, tooling, reporting, and automation.
- Build a durable operating model for incidents and escalations as OpenAI scales globally.
Requirements
- 10+ years of experience in incident management, technical support, escalation management, SRE, technical program management, or production operations.
- 5+ years of hands-on experience working in production, on-call, or high-urgency operational environments.
- 5+ years of leadership experience, ideally in a Support, Engineering, or similar environment.
- Experience acting as Incident Commander and owning coordination, decision-making, communication, and accountability during live incidents.
- Direct experience with customer-impacting incidents, executive escalations, safety-sensitive escalations, or high-severity technical issues.
- Ability to communicate clearly under pressure with engineers, support teams, executives, customer-facing teams, and external stakeholders.
- Hands-on experience with incident communications, including internal updates, executive briefings, customer updates, and status pages.
- Experience with incident management, paging, and alerting tools such as incident.io, PagerDuty, Datadog, Jira, Salesforce, Zendesk, or similar systems.
- Understanding of monitoring and observability sufficient to reason about alerts, system health, customer impact, and incident scope.
- Ability to lead post-incident retrospectives that produce clear root causes, corrective actions, and durable improvements.
- Ability to drive action items to closure and hold teams accountable without creating unnecessary process drag.
- Strong organization and composure in ambiguous or high-pressure situations.
- Ability to balance hands-on incident execution with longer-term systems building.
- Interest in using AI and automation to improve triage, routing, summarization, reporting, knowledge management, and incident follow-through.
Benefits
- Base salary range of $234,000–$315,000 per year, plus equity.
- Medical, dental, and vision insurance with employer contributions to Health Savings Accounts.
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.