Agent Post-Training, Computer Use Research

at OpenAI
USD 295,000-445,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

API ChatGPT Codex Compliance Data Pipelines LLM Machine Learning @ 6 Observability Reinforcement Learning @ 3 Statistics @ 6

Details

The Agent Post-Training team creates frontier agents for Codex, ChatGPT, the API, and other products. The team develops training data, environments, graders, training methods, and feedback loops for coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and other agent capabilities.

Responsibilities

  • Design and run experiments to improve agentic model behavior for complex computer use across desktops and browsers.
  • Own end-to-end improvements to the post-training stack, including reinforcement learning, data pipelines, graders, reward signals, evaluations, diagnostics, and model-behavior analysis.
  • Build evaluations and environments that expose model failures and convert those failures into training data, product fixes, or research directions.
  • Partner with Codex and ChatGPT product teams to translate user and product signals into model improvements.
  • Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and evaluation loops.
  • Help determine which integrations, capabilities, and fixes are ready for major model runs.
  • Improve large-scale training and launch systems, including experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
  • Work on cross-functional projects involving model training, product infrastructure, and production agent harnesses, including multi-agent systems and production-like training environments.
  • Debug difficult failures in shipped or near-shipped models and turn qualitative behavior into concrete hypotheses, experiments, and fixes.
  • Collaborate with researchers, engineers, product teams, infrastructure teams, and safety and alignment partners.

Requirements

  • Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field.
  • Hands-on experience with LLMs, reinforcement learning, RLHF/RLAIF, post-training, evaluations, graders, synthetic data, model training, coding agents, tool-using agents, or production machine learning systems.
  • Ability to develop experiments from behavioral problems by defining hypotheses, building pipelines, running models, analyzing results, and determining next steps.
  • Ability to work across research, product, infrastructure, data, evaluation, and safety boundaries and communicate clearly with each group.
  • Interest in open-ended problems requiring research judgment and engineering execution.
  • Interest in training and shipping models that make agents useful for developers, enterprises, researchers, and everyday users.

Benefits

  • Base pay of $295,000–$445,000 per year, plus equity and performance-related bonuses for eligible employees.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax flexible spending and commuter accounts.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

OpenAI is an equal opportunity employer committed to reasonable accommodations and compliance with applicable employment laws.

More jobs at OpenAI

Similar jobs