Agent Post-Training, Frontier Evals and Environments Research

at OpenAI
USD 295,000-445,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

API ChatGPT Codex LLM Machine Learning @ 6 Reinforcement Learning @ 3 Statistics @ 6

Details

The Agent Post-Training team creates frontier agents for Codex, ChatGPT, the API, and other products. The team trains models that can operate computers, collaborate with people and other agents, and support coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste.

The team builds data, environments, graders, training methods, and feedback loops that shape agent capabilities and carries them through major training runs into products used by real people.

As a researcher working on Frontier Evals and Environments, you will help build model environments to drive progress toward safe AGI/ASI. Your work will guide research programs for major training runs. Example open-sourced evaluations include GDPval, SWE-bench Verified, MLE-bench, PaperBench, and SWE-Lancer.

You will work with researchers, engineers, product teams, infrastructure teams, and safety and alignment partners to determine what should go into major model runs, measure results, and ship improvements into products.

Responsibilities

  • Create ambitious reinforcement learning environments to push models to their limits and measure frontier model capabilities, skills, and behaviors.
  • Develop methodologies for automatically exploring model behavior.
  • Investigate the science of measurement, including the scalability, reliability, and variance of evaluation methodologies.
  • Help steer training for the largest training runs.
  • Design scalable systems and processes for continuous evaluation.
  • Build self-improvement loops to automate model understanding.

Requirements

  • Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, with the ability to learn quickly across unfamiliar areas.
  • Hands-on experience with large language models, reinforcement learning, RLHF/RLAIF, post-training, evaluations, graders, synthetic data, model training, coding agents, tool-using agents, or production machine learning systems.
  • Ability to work on open-ended problems involving unclear paths and noisy signals, combining research judgment with engineering execution.
  • Interest in product impact and model behavior, including usefulness, reliability, honesty, taste, and ease of use.
  • Ability to turn a vague behavioral problem into a concrete experiment by defining a hypothesis, building a pipeline, running a model, analyzing results, and deciding what to do next.
  • Ability to collaborate across research, product, infrastructure, data, evaluations, and safety functions and communicate clearly with each group.
  • Willingness to build robust systems and processes when needed.
  • Interest in training and shipping models that make agents useful for developers, enterprises, researchers, and everyday users.

Benefits

  • Equity and performance-related bonuses for eligible employees.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax Flexible Spending Accounts, dependent care, and commuter benefits.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional taxable fringe benefits may be provided.

OpenAI is an equal opportunity employer committed to reasonable accommodations and fair employment practices.

More jobs at OpenAI

Similar jobs