Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About the Team
The Frontier Assurance team brings independent scrutiny into OpenAI’s safety decisions and helps the public understand and assess its safety work. The team leads third-party assessments and safeguard testing for flagship launches, pilots assurance mechanisms such as embedded auditing, runs the misalignment disclosure process, and incorporates independent expert input into critical safety decisions.
About the Role
This role builds programs that bring independent expertise into frontier AI safety decisions and make the evidence behind those decisions understandable to the public. Responsibilities include leading external research partnerships and third-party assessments, coordinating public safety documentation, and developing approaches to independent scrutiny and transparency.
The role involves collaboration across research, engineering, product, policy, and communications. The successful candidate will help ensure external findings inform concrete decisions and that public explanations accurately reflect the evidence, limitations, and remaining uncertainty.
This position is based in San Francisco, California, and follows a hybrid work model with three days in the office per week.
Responsibilities
- Design and run third-party assessment programs for frontier models and safeguards, including independent evaluations, adversarial testing, and embedded auditing.
- Work with researchers and external partners to define assessment questions, scope, access, timelines, and deliverables.
- Build and manage strategic research partnerships with third-party evaluators, academic labs, and independent experts.
- Enable rigorous research that challenges internal assumptions.
- Bring external findings to safety decision-makers and translate them into actionable recommendations for safeguards, deployment, product, and policy.
- Track follow-up actions and ensure partners understand how their input was considered.
- Partner with research, product, policy, and communications teams to explain the evidence behind OpenAI’s safety approach, including what was tested, what was learned, and where limitations and uncertainty remain.
- Translate complex technical results into accurate, accessible communication for expert and public audiences.
- Lead public transparency programs, including system cards, summaries of third-party assessments, and updates on significant findings and follow-up actions.
- Develop approaches to sharing methods, results, and limitations while protecting privacy, security, and sensitive information.
- Create channels for researchers and civil society to ask questions, provide feedback, and inform future assessments and transparency efforts.
- Own program goals, milestones, dependencies, and risks across assessment and transparency work.
- Keep internal and external stakeholders informed and resolve obstacles to execution.
Requirements
- Understanding of AI evaluations and measurement, with the ability to engage technical teams on evaluation design and results.
- Ability to interpret evaluation findings and communicate what they do and do not establish, including methodological limitations, uncertainty, and disagreement.
- Experience building research partnerships and managing external stakeholders, particularly academic researchers and independent experts.
- Ability to synthesize technical and social science research into clear executive summaries and accessible explanations for public audiences.
- Experience delivering technical publications or public reports involving multiple contributors.
- Sound judgment about what to share, with whom, and when.
- Experience working cross-functionally across product, research, and engineering teams and coordinating complex programs with policy and communications partners.
- Understanding of and interest in frontier AI safety and policy.
Compensation and Benefits
- Base salary: $239,000–$328,000 per year.
- Equity, performance-related bonuses for eligible employees, and benefits including medical, dental, and vision insurance; flexible spending accounts; a 401(k) plan with employer match; paid parental, medical, and caregiver leave; paid time off; paid holidays; mental health and wellness support; life and disability coverage; a learning and development stipend; office meals; and relocation support for eligible employees.
- OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.