Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 1
Compliance
Machine Learning @ 1
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Models are becoming increasingly capable, moving from tools that assist humans to agents that can plan, execute, and adapt in the real world. The Preparedness team addresses frontier AI risks through measurement, mitigation, and coordination, including monitoring evolving capabilities, maintaining alignment safeguards, and establishing preparedness targets.
This role supports preparations for accelerated AI development that may culminate in recursive self-improvement. The work focuses on anticipating future misalignment risks and designing mitigations for loss-of-control scenarios. Responsibilities span pre-deployment risk assessment, control measures, recursive-self-improvement-relevant training interventions, production safety systems, institutional practices, and external-facing communications.
Focus Areas
- Scalable oversight: Establish monitoring and oversight practices for model misbehavior that remain effective as model capabilities increase.
- Automated auditing: Develop automated approaches to identify severe model misalignments, analyze production traffic, and elicit tail risks before deployment.
- Rigorous monitorability: Test and red-team measurements of model misbehavior related to loss of control, including reward hacking, sandbagging, and scheming, as well as potential losses of Chain-of-Thought monitorability.
- Model behavior science: Design experiments and evaluations to understand problematic misalignment and whether safety-relevant capabilities lag behind dangerous capabilities. This may include training model organisms of misbehavior and developing training interventions.
- Coordination and verification: Prototype technical mechanisms for verifying compliance with potential future AI safety agreements.
- AI R&D risk measurement: Track progress toward automating technical staff to inform near-term investments in alignment and security.
- RSI safety cases: Identify and address blind spots in recursive self-improvement safety mitigations.
The team alternates between rigorous, hypothesis-driven research and turning insights into interventions or control systems that affect production models, with occasional support from engineering teams.
Responsibilities
- Consider future problems OpenAI might face and determine how to prepare for them.
- Turn open-ended objectives, such as preparing for future misalignment threats, into concrete and prioritized research directions.
- Execute quickly by building scrappy prototypes and iteratively improving them into established safety-pipeline components.
- Secure buy-in from other OpenAI staff and communicate work clearly.
- Collaborate with or manage other staff as needed to rapidly scale work on these problems.
Requirements
- Exceptional technical execution ability.
- Strong strategic and research judgment, including the ability to prioritize effectively in domains with weak feedback loops.
- Passion for mitigating risks associated with recursive self-improvement.
- Motivation to do the work that most positively impacts the future of AI development.
- Experience in one or more relevant domains, such as machine learning research, AI alignment, or AI verification, is a bonus.
Benefits
- Base salary of $295,000–$445,000 per year, plus equity.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax Flexible Spending Accounts and commuter benefits.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, and office closures.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer committed to reasonable accommodations and compliance with applicable employment laws.