Researcher, Agent Safety, Oversight and System Mitigations

at OpenAI
USD 380,000-500,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

AI @ 3 Agentic Systems Codex Security @ 6

Details

The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. The team focuses on reducing the probability of severe unintended outcomes while preserving agents' ability to act effectively and autonomously.

The work spans training methods, environments and data; evaluations and production metrics; and oversight and system mitigation mechanisms that reduce harmful actions while preserving useful agent autonomy.

This role focuses on oversight and system-level mitigations that enable increasingly capable agents to operate safely and autonomously in real environments. The role involves building oversight systems for current internal and external use, as well as studying longer-term questions about supervising, constraining, and correcting increasingly capable agentic systems.

Responsibilities

  • Design, build, and evaluate system-level controls for agent actions, including agent-based review.
  • Plan how controls fit into broader systems involving sandboxing, process isolation, and permission boundaries.
  • Work closely with a Codex harness engineering team to productionize AI controls.
  • Red-team end-to-end agentic systems to assess whether controls prevent data exfiltration, unsafe tool use, and other harmful outcomes.
  • Improve the safety–productivity tradeoff by measuring and reducing missed harmful actions, unnecessary blocks, approval burden, and latency.

Requirements

  • Strong systems or security instincts, with the ability to reason concretely about isolation boundaries, permissions, attack surfaces, and failure modes in complex systems.
  • Ability to turn ambiguous safety questions into concrete threat models, reproducible experiments, and practical mitigations.
  • Willingness to revise approaches based on evidence from deployment.
  • Ability to build robust experimental infrastructure and design evaluations that distinguish promising mitigations from brittle ones.
  • Deep interest in frontier AI alignment, safety, and control.
  • A background in AI control or security is welcome but not required.

Work Arrangement

  • Based in San Francisco, California.
  • Hybrid work model with three days in the office per week.
  • Relocation assistance is offered to new employees.

Benefits

  • Medical, dental, and vision insurance for employees and their families, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental leave, medical leave, and caregiver leave.
  • Paid time off, paid company holidays, office closures, and paid sick or safe time as required by law.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily meals in offices and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.

More jobs at OpenAI

Similar jobs