Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Bash @ 4
ChatGPT
Codex
JavaScript @ 4
LLM @ 4
PowerShell @ 4
Python @ 4
QA @ 4
SQL @ 4
Security @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Critical Harm Operations sits within User Safety & Risk Operations and builds enforcement systems for Frontier Risk and Material Harm that are accurate, fast, defensible, and built to scale. The Cyber vertical turns policy into reviewer standards, calibrated judgment, quality systems, escalation paths, and automation guardrails.
The role combines hands-on cyber judgment with systems-level operating design. The successful candidate will resolve complex dual-use questions, evolve standard operating procedures, improve reviewer and vendor capabilities, and build practical tools and automations. This is a senior individual contributor role focused on durable improvements to the operating model and the reviewers who run it.
Responsibilities
- Drive the Cyber Operations operating model across domain priorities, SOPs, escalation paths, quality health, vendor capability, roadmap inputs, and trusted access strategies.
- Serve as the senior cyber expert for complex or high-risk decisions across ChatGPT, API, Codex, agents, and emerging product surfaces.
- Translate policy ambiguity, quality misses, appeals, and reviewer disagreement into decision rules, calibration examples, training, and tooling requirements.
- Build operating systems and quality loops, including golden sets, holdouts, double-labeling, adjudication, error taxonomies, reviewer calibration, and automation evaluations.
- Raise FTE and BPO capability through onboarding, certification, coaching, recurring calibration, and vendor-performance partnership.
- Use quality, appeals, SLA, backlog, and disagreement signals to diagnose root causes and prioritize high-leverage fixes.
- Build hands-on solutions, including SQL analyses, scripts, dashboards, LLM evaluation workflows, evidence enrichment, routing logic, and lightweight automations, to improve decision quality and reduce manual effort.
- Partner with Policy, Integrity, Safety Systems, Security, Legal, Product, Engineering, and Investigations teams to operationalize changes and drive launch readiness.
Requirements
- 8+ years of hands-on cybersecurity experience in offensive security, threat intelligence, incident response, security research, red teaming, application security, DFIR, malware analysis, or a related field.
- Deep understanding of attacker tradecraft, vulnerability exploitation, credential abuse, malware, persistence, evasion, exfiltration, cloud or identity abuse, and ambiguous dual-use activity.
- Experience building or improving high-stakes operations, reviewer programs, QA systems, escalation workflows, or vendor/BPO programs.
- Ability to turn complex cyber and policy judgment into reviewer-usable SOPs, decision trees, training, and concise written recommendations.
- Experience using SQL, Python, C/C++, JavaScript, PowerShell, Bash, APIs, LLM tooling, or automation to solve operational problems.
- Understanding of human-in-the-loop automation, evaluations, monitoring, holdouts, and fallback paths for sensitive workflows.
- Ability to operate independently in ambiguity, communicate clearly across technical and non-technical audiences, and exercise sound judgment and discretion when handling sensitive material.
- Experience with trust and safety, platform abuse, cyber misuse of AI systems, or LLM safety.
Nice to Have
- Experience with golden sets, classifier or prompt evaluations, and reviewer-quality programs.
- Experience enabling global vendor reviewer operations.
Benefits
- Base salary of $252,000–$280,000 per year, plus equity.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily meals in offices and meal delivery credits as eligible.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
More jobs at OpenAI
Software Engineer, API Multimodal
OpenAI · San Francisco, United States
USD 293,000-385,000 per year
Program Manager, Critical Harm Operations
OpenAI · San Francisco, United States
USD 252,000-280,000 per year
Industrial Compute
OpenAI · United States
USD 150,000-300,000 per year
Software Engineer, API Agents
OpenAI · San Francisco, United States
USD 293,000-385,000 per year
Software Engineer, Agent Productivity
OpenAI · San Francisco, United States
USD 293,000-342,000 per year
Similar jobs
Support Engineer
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 210,000-250,000 per year
Product Support Specialist
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 131,000-210,000 per year
AI Success Engineer - AMER
OpenAI · San Francisco, United States, New York City, United States, United States
USD 234,000-260,000 per year
Solutions Engineer, Pre-Sales
OpenAI · San Francisco, United States, New York City, United States
USD 156,600-245,000 per year
Staff Forward Deployed Engineer, AI
SentinelOne · United States
USD 156,000-215,000 per year
Senior Staff Forward Deployed Engineer, AI
SentinelOne · United States
USD 184,000-253,000 per year
Senior AI Engineer
Grafana Labs · United States
USD 154,400-185,300 per year
Solutions Engineer, Pre-Sales
OpenAI · San Francisco, United States
USD 174,000-245,000 per year