Model Policy Manager, Multimodal Safety

at OpenAI
USD 266,000-335,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

AI @ 6 ChatGPT

Details

Safety Systems manages the lifecycle of safety efforts for OpenAI’s frontier models, including system-level safeguards, model training, evaluation, and red-teaming. The Model Policy team designs policies that define safe and reliable model behavior in real-world environments.

Responsibilities

  • Design and maintain model policies for audio, image, video, and omni-modal behavior.
  • Translate theories of harm and threat models into behavioral safety policies, evaluation criteria, grading guidance, and safeguards.
  • Identify and analyze safety regressions and failure patterns to identify gaps in existing policies and inform policy iteration.
  • Develop policy artifacts supporting model training, evaluation, and deployment, including behavior instructions, human-data campaigns, golden sets, and evaluations.
  • Partner with AI researchers, domain experts, and product teams to operationalize policy into measurable model behavior.
  • Work with multimodal AI models and capabilities such as GPT-Live and ChatGPT Images.

Requirements

  • Strong judgment about the real-world risks of advanced multimodal AI systems.
  • Experience turning ambiguous safety questions into clear, data-driven policies, behavioral boundaries, and measurable evaluation criteria.
  • Ability to treat policy as an end-to-end, measurable system by testing intended model behavior and diagnosing gaps across policy, data, graders, and safeguards.
  • Strong technical judgment when designing policies around model behavior that can realistically be trained, measured, and supervised at scale.
  • Strong technical fluency and experience using AI tools to accelerate policy development, evaluate model behavior, analyze failure patterns, and turn findings into actionable improvements.
  • Hands-on experience with model data and evaluation results, including inspecting examples, analyzing failure patterns, assessing data quality, and distinguishing policy failures from grader, model, or system failures.
  • Ability to work in fast-paced, collaborative research environments where priorities shift as models, evidence, and risks change.
  • A pragmatic, evidence-driven approach to reducing risk while preserving beneficial uses of AI.
  • Experience driving consensus and action in ambiguous spaces.

Benefits

  • Base salary range of $266,000–$335,000 per year, plus equity.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent care, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

This role is based in San Francisco, California, and follows a hybrid work model requiring three days in the office per week. OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.

More jobs at OpenAI

Similar jobs