Applied AI Engineer, Codex Core Agent
📍 New York City, United States
📍 San Francisco, United States
📍 Seattle, United States
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Codex
Experimentation
LLM @ 3
Machine Learning @ 3
Python @ 2
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Codex Core Agent team builds the kernel of Codex, focusing on improving the agent, accelerating research, and deploying improvements to production. The team works across production performance, including tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; shared infrastructure; and feedback loops that improve models and agent behavior through real-world usage.
This role focuses on improving agent performance on real software engineering tasks and closing the gap between research capability and real-world usefulness. You will work with research, infrastructure, and product teams to make agents powerful, useful, steerable, and reliable in practice, while delivering measurable improvements in solve rate, usefulness, and economic value.
Responsibilities
- Design and iterate on agent behaviors across real-world coding tasks and long-horizon workflows.
- Work with research teams to develop and run evaluations measuring agent performance, regressions, failure modes, and edge cases.
- Improve performance through prompting, tool-use strategies, context construction, and model-facing experimentation.
- Analyze production failures and systematically improve robustness and reliability.
- Build feedback loops and data systems that bring better real-task data into evaluation and research.
- Work with product teams to shape user-facing agent experiences and the interfaces the agent depends on.
- Help define success criteria for agents completing complex tasks end to end.
Requirements
- Experience building or shipping machine learning or LLM-powered products.
- Strong Python skills and familiarity with modern machine learning tooling.
- Experience with model evaluation, fine-tuning, or prompt design.
- A systems- and user-outcome-oriented approach rather than a focus solely on model metrics.
- Ability to debug messy, real-world failures and turn them into improvements.
- Interest in turning research and model potential into systems that work reliably for users.
Bonus Qualifications
- Experience with agent frameworks or tool-using LLM systems.
- Research experience with code generation models or developer tooling.
- Experience working with large, messy datasets or production logs.
Benefits
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health and dependent care expenses, as well as commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, and paid office closures.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.