Agent Post-Training, API & Power Users

at OpenAI
USD 295,000-445,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

API @ 3 ChatGPT Codex LLM Machine Learning @ 6 Observability Statistics @ 6

Details

The Agent Post-Training team creates frontier agents for Codex, ChatGPT, the API, and other products. The team develops training data, environments, graders, training methods, and feedback loops for capabilities including coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste.

As a member of the API and power-users team, you will improve the capabilities, reliability, and product fit of OpenAI's agentic models for power users and API developers. The work may include designing evaluations from real developer workflows, building training environments around production-like tool use, converting qualitative model failures into training data or post-training interventions, and driving behavior improvements from discovery through post-training, integration, and launch.

Responsibilities

  • Design and run experiments that improve model behavior in API and power-user workflows, including function calling, tool use, coding, planning, long-horizon execution, factuality, instruction following, error recovery, and calibrated reasoning.
  • Build evaluations, graders, and environments from real developer and power-user workflows, then turn observed failures into training data, model-behavior hypotheses, and shipped improvements.
  • Partner with API and power users to identify high-leverage behavior gaps and convert product signals into post-training interventions.
  • Improve model behavior when models are composed into systems, including reliable tool use, respect for developer intent, partial-failure handling, clarification, and coherence across multi-step tasks.
  • Own end-to-end model behavior projects, from qualitative failure analysis through data generation, training experiments, evaluation design, integration into major runs, and launch readiness.
  • Develop feedback loops using power-user traces, API usage patterns, and production-like environments to discover agentic model failures and capability gaps.
  • Help determine which agentic capabilities, behavioral fixes, and partner-team integrations are ready for major model runs.
  • Debug failures in shipped or near-shipped models by analyzing traces, evaluations, training data, model outputs, and product context.
  • Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and evaluation loops.
  • Improve large-scale training and launch systems in areas such as experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
  • Work on cross-functional projects involving model training, product infrastructure, and the production agent harness, including multi-agent systems and training against production-like environments.

Requirements

  • Strong technical fundamentals in machine learning, software engineering, systems, statistics, or applied research.
  • Ability to learn quickly across unfamiliar parts of the technology stack.
  • Hands-on experience with LLMs, post-training, RL, RLHF, RLAIF, evaluations, graders, synthetic data, coding agents, tool-using agents, API products, or production machine learning systems.
  • Strong understanding of model behavior and the ability to form concrete hypotheses from transcripts, traces, evaluation failures, or API interactions.
  • Comfort working on ambiguous capability problems with noisy signals and qualitative failures.
  • Strong interest in developer and expert-user experience, particularly models embedded in real workflows, API products, and agent harnesses.
  • Ability to work across research, product, infrastructure, data, evaluations, and safety boundaries and communicate clearly with each group.
  • Willingness to build reliable systems and processes as needed.

Benefits

  • Equity and performance-related bonuses for eligible employees.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent care, parking, and transit expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional taxable fringe benefits may include charitable donation matching and wellness stipends.

More jobs at OpenAI

Similar jobs