Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API @ 3
ChatGPT @ 3
Codex @ 3
Data Pipelines
LLM
Machine Learning @ 6
Observability
Reinforcement Learning @ 3
Statistics @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Agent Post-Training team creates the frontier agents OpenAI ships in Codex, ChatGPT, the API, and other products. The team develops training data, environments, graders, training methods, evaluations, and feedback loops for agents capable of coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and other capabilities.
The role focuses on improving the capabilities, reliability, and product fit of OpenAI's agentic models. Responsibilities may span research, engineering, data, evaluations, infrastructure, product integration, and launch. The role may be based in San Francisco or London, subject to team needs and location approval.
Responsibilities
- Design and run experiments that improve agentic model behavior across coding, tool use, function calling, computer use, multi-agent collaboration, long-horizon tasks, factuality, instruction following, and calibrated reasoning.
- Own end-to-end improvements to the post-training stack, including reinforcement learning, data pipelines, graders, reward signals, evaluations, diagnostics, and model-behavior analysis.
- Build evaluations and environments that expose model failures, then turn those failures into training data, product fixes, or new research directions.
- Partner with Codex, API/platform, and ChatGPT product teams to understand user needs and translate product signals into model improvements.
- Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and evaluation loops.
- Help determine which integrations, capabilities, and fixes are ready for inclusion in major model runs.
- Improve large-scale training and launch systems, including experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
- Work on cross-functional projects involving model training, product infrastructure, and production agent harnesses, including multi-agent systems and production-like training environments.
- Debug difficult failures in shipped or near-shipped models and turn qualitative behavior into hypotheses, experiments, and fixes.
Requirements
- Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field.
- Hands-on experience with large language models, reinforcement learning, RLHF/RLAIF, post-training, evaluations, graders, synthetic data, model training, coding agents, tool-using agents, or production machine learning systems.
- Ability to work on open-ended problems involving ambiguous goals, noisy signals, research judgment, and engineering execution.
- Interest in product impact and model behavior, including making agents useful, reliable, honest, tasteful, and easy to work with.
- Ability to translate behavioral problems into concrete experiments by defining hypotheses, building pipelines, running models, analyzing results, and determining next steps.
- Ability to collaborate across research, product, infrastructure, data, evaluations, and safety teams.
- Willingness to build reliable systems and processes when needed.
Benefits
- Equity and performance-related bonuses for eligible employees.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
More jobs at OpenAI
Technical Program Manager, Enterprise
OpenAI · San Francisco, United States
USD 257,000-445,000 per year
Product Design Leadership, Growth
OpenAI · San Francisco, United States
USD 347,000-405,000 per year
Head of Technical Success, Government
OpenAI · Washington, United States
USD 374,000-450,000 per year
Software Engineer, Plugin Developer Platform
OpenAI · San Francisco, United States
USD 185,000-490,000 per year
Red Team Specialist - Cyber
OpenAI · San Francisco, United States, Seattle, United States
USD 198,000-320,000 per year
Similar jobs
Agent Post-Training, Context Research
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Agent Post-Training, Computer Use Research
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Agent Post-Training, Connectors Research
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Agent Post-Training, Artifacts Research
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Research Engineer, Codex
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Agent Post-Training, Frontier Evals and Environments Research
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Agent Post-Training, Personality
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Agent Post-Training, API & Power Users
OpenAI · San Francisco, United States
USD 295,000-445,000 per year