Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Communication @ 3
Security @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Safety Systems team works to build and deploy safe AGI. The Model Policy team investigates emerging model failures, defines safe and reliable model behavior, and develops the data, evaluations, monitoring, and safeguards needed to improve and validate frontier AI systems.
In this role, you will address real-world risks arising from model misalignment as models become more autonomous and operate over longer horizons. You will investigate misaligned behavior across extended trajectories and translate findings into behavioral policies, evaluations, monitoring, and safeguards.
Responsibilities
- Identify vulnerabilities that emerge as models interact with tools, data, and external systems, and translate them into model- and system-level safeguards.
- Develop threat models and empirical frameworks for understanding harmful outcomes from misaligned behavior.
- Identify the underlying behaviors and system conditions that drive harmful outcomes.
- Turn findings into policy frameworks, evaluation criteria, online measurement, and safeguards.
- Develop human data campaigns and gold sets to support measurement and evaluation of emerging behaviors and risks.
- Partner with research, engineering, security, and product teams to shape model and system safety while balancing safety, utility, and business risks.
- Inform deployment decisions, system cards, safeguards reports, and OpenAI's broader approach to agentic safety.
- Build monitoring approaches that detect regressions and emerging risks after deployment.
Requirements
- Strong background in AI agent safety, privacy, security, cybersecurity, or an adjacent field, with an adversarial mindset for investigating real-world harmful outcomes.
- Demonstrated interest in AI alignment and a strong understanding of the technical drivers of misaligned model behavior.
- Technical fluency sufficient to work directly with evaluation and training data, understand what the data shows, and identify limitations, patterns, and opportunities for deeper investigation.
- Hands-on experience with model data and evaluation results, including inspecting examples, analyzing failure patterns, assessing data quality, and distinguishing policy failures from grader, model, or system failures.
- Ability to use empirical evidence to develop and refine safety policies and safeguards.
- Ability to translate complex or ambiguous alignment risks into precise behavioral expectations and measurable evaluation criteria.
- Ability to work effectively across research, engineering, security, product, and policy teams.
- Clear communication skills when discussing complex and uncertain technical risks.
- Comfort working in fast-paced, collaborative research environments where priorities shift as models, evidence, and risks change.
Workplace And Location
The role is based in the San Francisco office. OpenAI uses a hybrid model with three days in the office per week and optional work from home on Thursdays and Fridays. Relocation support is available to eligible employees.
Compensation And Benefits
The base salary range is $207,000–$335,000 USD, plus equity. Total compensation may also include performance-related bonuses and benefits such as medical, dental, and vision insurance; health savings and flexible spending accounts; a 401(k) with employer match; paid parental, medical, and caregiver leave; paid time off; paid holidays; mental health and wellness support; life and disability coverage; a learning and development stipend; daily office meals; meal delivery credits; relocation support; charitable donation matching; and wellness stipends.