AI Systems Engineer, Codex Agents

at OpenAI
USD 230,000-385,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

AI @ 6 API @ 6 Codex Debugging Distributed Systems @ 3 Experimentation GPU @ 3 LLM @ 3 Machine Learning Observability Profiling @ 3 Python @ 3 Rust @ 6 Security @ 3

Details

The Codex Core Agents team builds the agent harness that turns model capability into real-world action. The team owns the systems around the model, including prompting and interpreting model outputs, safely executing actions in real environments, and feeding production experience back into better models and agent behavior.

The team works across harnesses, model interaction, inference, sandboxed execution, orchestration, evaluations, production reliability, and the performance envelope around tokens, latency, cost, capacity, and quality. The harness is open source and increasingly part of how models are trained and evaluated.

The role focuses on building AI systems that make Codex agents dependable in production. The ideal candidate is an agent-systems builder who can work across low-level systems and ML workflows and debug Codex behavior end to end across the harness, model behavior, inference and runtime stack, GPU fleet, and product surface.

Responsibilities

  • Design and build the core agent harness and execution loop that enables Codex agents to interpret model outputs, use tools, execute code, and complete long-horizon tasks safely.
  • Build sandboxing, isolation, orchestration, state, and workflow infrastructure for agents operating in real development environments.
  • Develop evaluation, experimentation, and debugging systems that distinguish harness issues, model behavior, inference and runtime issues, and product failures.
  • Run ablations across prompts, model-facing interfaces, context construction, tool-use strategies, and harness behavior to improve solve rate, reliability, latency, and cost.
  • Improve observability, profiling, and diagnostics across the agent stack, including backend systems, inference, GPUs, and fleet capacity.
  • Work closely with research to make the harness trainable, measurable, and useful for improving frontier agentic models.
  • Build shared primitives that make Codex faster, safer, more reliable, and easier for other teams and open-source users to build on.
  • Work with research, infrastructure, and product to design agent harness capabilities, run experiments, build frameworks for assessing production agent performance, and turn production failures into durable improvements.

Requirements

  • Experience building or operating production systems in distributed systems, infrastructure, developer tooling, sandboxing, virtualization, cloud platforms, or ML systems.
  • Ability to work across Rust systems code, Python configuration layers, APIs, agent orchestration, evaluations, logs and traces, inference behavior, runtime constraints, and user outcomes.
  • Hands-on experience with LLM applications, coding agents, evaluations, model deployment, inference, compiler or runtime performance, or developer platforms.
  • Strong focus on reliability, safety, performance, debuggability, and clean abstractions.
  • Ability to debug from evidence and move quickly from ambiguous production failures to practical, durable fixes.
  • Interest in working close to research while shipping changes to production.
  • Ability to write meaningful code, demonstrate strong ownership, and lead scoped or multi-team AI systems work.

Bonus Qualifications

  • Deep Rust, systems, sandboxing, isolation, or low-level platform experience.
  • Experience with coding agents, agent harnesses, tool-using LLM systems, model evaluations, or post-training feedback loops.
  • Background in compilers, kernels, runtimes, inference optimization, GPU systems, benchmarking, profiling, or performance engineering.
  • Experience building production infrastructure used by many engineers or users under demanding reliability and security constraints.
  • Open-source infrastructure or developer-platform experience with a strong focus on APIs and usability.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI is an equal opportunity employer and does not discriminate on the basis of legally protected characteristics. Background checks are administered in accordance with applicable law. Reasonable accommodations are available to applicants with disabilities.

More jobs at OpenAI

Similar jobs