Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API
ChatGPT @ 3
Codex @ 3
Data Pipelines
LLM
Machine Learning @ 6
Observability
Reinforcement Learning @ 3
Statistics @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Agent Post-Training team creates frontier agents for Codex, ChatGPT, the API, and other products. The team develops training data, environments, graders, training methods, and feedback loops for capabilities including coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste.
As a Context Researcher on Agent Post-Training, you will scale compute spent on context and work in the frontier training stack to enable new model-training paradigms through the Codex Chronicle product interface. You will collaborate with researchers, engineers, product teams, infrastructure teams, and safety and alignment partners to determine what should be included in major model runs, measure results, and ship improvements to products used by real people.
Responsibilities
- Design and run experiments that improve scaling of compute on context.
- Own end-to-end improvements to the post-training stack, including reinforcement learning, data pipelines, graders, reward signals, evaluations, diagnostics, and model-behavior analysis.
- Build evaluations and environments that expose model failures, then convert those failures into training data, product fixes, or new research directions.
- Partner with Codex and ChatGPT product teams to understand user needs and translate product signals into model improvements.
- Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and evaluation loops that shape downstream agent behavior.
- Help determine which integrations, capabilities, and fixes are ready for inclusion in major model runs.
- Improve large-scale training and launch systems, including experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
- Take on cross-functional projects involving model training, product infrastructure, and production agent harnesses, such as multi-agent systems or training directly against production-like environments.
- Debug difficult failures in shipped or near-shipped models and turn qualitative behavior into concrete hypotheses, experiments, and fixes.
Requirements
- Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field.
- Hands-on experience with LLMs, reinforcement learning, RLHF/RLAIF, post-training, evaluations, graders, synthetic data, model training, coding agents, tool-using agents, or production machine learning systems.
- Ability to work on open-ended problems involving unclear paths and noisy signals, combining research judgment with engineering execution.
- Interest in product impact and model behavior, including what makes an agent useful, reliable, honest, tasteful, and easy to work with.
- Ability to turn a vague behavioral problem into a concrete experiment by defining a hypothesis, building a pipeline, running a model, analyzing results, and determining next steps.
- Ability to work across research, product, infrastructure, data, evaluations, and safety boundaries and communicate clearly with each group.
- Willingness to build reliable systems and processes when needed.
- Interest in training and shipping models that make agents useful for developers, enterprises, researchers, and everyday users.
Benefits
- Equity and performance-related bonuses for eligible employees.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily meals in offices and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional taxable fringe benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer committed to reasonable accommodations for applicants with disabilities. Background checks will be administered in accordance with applicable law.