Researcher, Alignment CoT Monitorability

at OpenAI
USD 250,000-445,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

AI Debugging @ 6 LLM Machine Learning @ 6 Reinforcement Learning

Details

The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. The team measures monitorability, studies training mechanisms that affect it, explores methods to improve it, and works on monitoring, auditing, and alignment.

The role focuses on designing and running empirical experiments to improve understanding of model monitorability. The researcher will investigate how training interventions across the model-development pipeline influence whether reasoning remains legible, build evaluations that make these questions measurable, and translate findings into practical oversight and training recommendations. The role may also involve developing monitoring models or methods and applying them to OpenAI's largest training runs.

This position is based in San Francisco, California, and follows a hybrid work model requiring three days in the office per week.

Responsibilities

  • Design and run empirical studies of chain-of-thought monitorability across frontier reasoning models and training settings.
  • Build evaluations that measure whether monitors can reliably predict properties of interest, including high-stakes forms of misbehavior.
  • Investigate how pre-training, synthetic data, mid-training, post-training, reinforcement learning, and other interventions improve or degrade monitorability.
  • Analyze model behavior and turn monitoring observations into hypotheses, experiments, and recommendations.
  • Translate research findings into practical monitoring and oversight approaches that can inform real training runs.
  • Collaborate with researchers and engineers across model training, alignment evaluations, monitoring, and frontier-risk work.
  • Produce externally publishable research when results advance the broader science of alignment.

Requirements

  • Strong hands-on experience training, evaluating, or debugging large machine learning models, especially large language models.
  • Depth in alignment, interpretability, model behavior, empirical machine learning, or adjacent research.
  • Strong interest in alignment, model behavior, or interpretability.
  • Interest in chain-of-thought monitorability, monitoring methods, and scalable oversight.
  • Ability to turn ambiguous research questions into measurable experiments and follow the evidence when results are subtle or noisy.
  • Ability to move between research ideation and engineering execution.
  • Curiosity about multiple approaches to understanding model behavior and openness to different methodological lenses.
  • Ability to operate independently while collaborating closely across research and engineering teams.
  • Direct chain-of-thought interpretability experience is welcome but not required.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. The company develops AI systems and seeks to deploy them safely through its products.

OpenAI is an equal opportunity employer and does not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Background checks are administered in accordance with applicable law. Reasonable accommodations are available to applicants with disabilities.

Benefits

  • Base salary of $250,000–$445,000 per year, plus equity.
  • Medical, dental, and vision insurance for employees and their families, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent care, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, paid company holidays, office closures, and sick or safe time as required by law.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily meals in offices and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

More jobs at OpenAI

Similar jobs