Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
ChatGPT
Codex
Data Pipelines
Experimentation
LLM
Machine Learning @ 6
Observability
Reinforcement Learning @ 3
Statistics @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Codex Research team creates frontier agents for Codex, ChatGPT, the API, and other products. The team works on coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and model behavior.
The role focuses on improving the capabilities, reliability, and product fit of OpenAI's agentic models. Responsibilities may span research, engineering, data, evaluations, infrastructure, and product integration, including taking capabilities from initial experimentation through integration and launch.
Responsibilities
- Design and run experiments to improve agentic model behavior across coding, tool use, function calling, computer use, multi-agent collaboration, long-horizon tasks, factuality, instruction following, and calibrated reasoning.
- Own end-to-end improvements to the post-training stack, including reinforcement learning, data pipelines, graders, reward signals, evaluations, diagnostics, and model-behavior analysis.
- Build evaluations and environments that expose model failures, then turn those failures into training data, product fixes, or research directions.
- Partner with Codex, API and platform, ChatGPT, and general-agent product teams to translate product signals into model improvements.
- Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and evaluation loops.
- Help determine which integrations, capabilities, and fixes are ready for major model runs.
- Improve large-scale training and launch systems, including experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
- Work on cross-functional projects involving model training, product infrastructure, production agent harnesses, multi-agent systems, and production-like environments.
- Debug difficult failures in shipped or near-shipped models and turn qualitative behavior into hypotheses, experiments, and fixes.
Requirements
- Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field.
- Hands-on experience with LLMs, reinforcement learning, RLHF/RLAIF, post-training, evaluations, graders, synthetic data, model training, coding agents, tool-using agents, or production machine learning systems.
- Ability to work on open-ended problems involving ambiguous paths and noisy signals, combining research judgment with engineering execution.
- Interest in product impact and model behavior, including what makes an agent useful, reliable, honest, tasteful, and easy to work with.
- Ability to turn a vague behavioral problem into a concrete experiment by defining a hypothesis, building a pipeline, running a model, analyzing results, and determining next steps.
- Ability to work across research, product, infrastructure, data, evaluations, and safety boundaries and communicate clearly with each group.
- Willingness to build robust systems and processes when needed.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.