Research Engineer, Frontier Evals & Environments

at OpenAI
USD 205,000-380,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

AI API ChatGPT Codex LLM Machine Learning @ 6 Reinforcement Learning @ 3 Statistics @ 6

Details

The Agent Post-Training team creates the frontier agents OpenAI ships in Codex, ChatGPT, the API, and other frontier products. The team builds data, environments, graders, training methods, and feedback loops for capabilities including coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste.

As a researcher working on Frontier Evals & Environments, you will help build model environments to drive progress toward safe AGI/ASI and guide research programs for major training runs. You will collaborate with researchers, engineers, product teams, infrastructure teams, and safety and alignment partners to determine what should go into model runs, measure results, and ship improvements into products.

Prior open-source evaluations associated with this work include GDPval, SWE-bench Verified, MLE-bench, PaperBench, and SWE-Lancer.

Responsibilities

  • Create ambitious reinforcement learning environments to push models to their limits and measure frontier model capabilities, skills, and behaviors.
  • Develop methodologies for automatically exploring model behavior.
  • Analyze the science of measurement, including the scalability, reliability, and variance of evaluation methodologies.
  • Help steer training for the largest model training runs.
  • Design scalable systems and processes to support continuous evaluation.
  • Build self-improvement loops to automate model understanding.
  • Move from behavioral problems to concrete experiments by defining hypotheses, building pipelines, running models, analyzing results, and deciding next steps.
  • Work across research, product, infrastructure, data, evaluations, and safety boundaries.

Requirements

  • Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field.
  • Hands-on experience with LLMs, reinforcement learning, RLHF/RLAIF, post-training, evaluations, graders, synthetic data, model training, coding agents, tool-using agents, or production machine learning systems.
  • Ability to work on open-ended problems involving uncertain paths and noisy signals, combining research judgment with engineering execution.
  • Interest in product impact and model behavior, including making agents useful, reliable, honest, tasteful, and easy to work with.
  • Ability to communicate clearly and collaborate with research, product, infrastructure, data, evaluation, and safety teams.
  • Willingness to build robust systems and processes as needed.
  • Interest in training and shipping models for developers, enterprises, researchers, and everyday users.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. The company is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.

Benefits

  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent care, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, paid company holidays, office closures, and paid sick or safe time as required by law.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Equity, performance-related bonuses for eligible employees, and additional benefits may be provided.

More jobs at OpenAI

Similar jobs