Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 3
CI/CD @ 6
Data Pipelines @ 1
GPU @ 3
Kubernetes @ 3
Machine Learning
Observability @ 6
Python @ 3
Reinforcement Learning @ 1
Robotics
Rust @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve people's lives.
The Simulation Infrastructure Engineer will turn simulation systems into reliable, automated, production-quality pipelines that power model training, evaluation, and hardware-in-the-loop validation. This role owns the automation, orchestration, and tool integration that apply simulation to concrete robotics tasks, including CI/CD for SIL/HIL, presubmit checks, automatic model evaluation, metric computation and reporting, and runtime infrastructure for running simulations at scale. The role collaborates closely with Sim Realism, Sim Environments, Research, and Ops to make simulation an integrated, reproducible, and measurable part of ML and robotics workflows.
This role is based in San Francisco, California, and requires in-person work four days per week.
Responsibilities
- Build and maintain presubmit checks and continuous integration and deployment pipelines for simulation code, environments, and tasks so simulation artifacts are testable, versioned, and reproducible.
- Implement end-to-end automation to run model evaluation in simulation (SIL) and orchestrate hardware-in-the-loop (HIL) runs.
- Compute realism and task metrics, generate dashboards and alerts, and ensure evaluation is repeatable and auditable.
- Create robust APIs and connectors so research, training, and data-collection systems can schedule, seed, and evaluate batches of simulations.
- Support reinforcement learning rollouts, imitation-data collection, and presubmit model checks.
- Build scheduling, batching, and orchestration for very large numbers of concurrent rollouts, targeting tens of thousands of rollouts and large RL workloads.
- Solve engine-level scaling challenges, including parallelization and batching multiple runs per engine, and optimize cloud/GPU runtime reliability.
- Produce metrics and tooling for measuring simulation health, throughput, fidelity regressions, and cost.
- Create presubmit and canary tests that identify simulation regressions early.
- Implement artifact versioning, environment immutability through images and asset versions, experiment provenance, and policies for resource quotas and cost control across the simulation farm.
- Work closely with Sim Environments, Sim Realism, Research, and Ops to ensure simulation improvements directly translate into better model evaluation and training results.
Requirements
- Deep software engineering and infrastructure experience, including building CI/CD at scale, authoring reliable pipelines, and shipping production services that coordinate many moving parts.
- Experience with distributed compute and cloud GPU workloads, including scheduling, batching, GPU orchestration, and throughput and cost optimization.
- Experience building or maintaining HIL/SIL workflows or other simulation-to-hardware integrations, with an understanding of the operational challenges of bridging software and hardware testbeds.
- Strong automation, observability, and metrics skills, including defining meaningful KPIs and surfacing regressions early.
- Ability to design APIs and developer tooling that enable research and software engineering teams to submit jobs, reproduce experiments, and interpret results.
- Experience with Python, C++, or Rust; container orchestration with Kubernetes; distributed task queues; and CI systems.
- Experience with reinforcement learning tooling, task generators, or large-scale data pipelines is a bonus.
- Ability to collaborate across teams to turn experimental simulation work into dependable production tooling.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI is an equal opportunity employer and does not discriminate on the basis of legally protected characteristics. Background checks are administered in accordance with applicable law. Reasonable accommodations are available to applicants with disabilities.
Benefits
- Medical, dental, and vision insurance for employees and their families, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and paid sick or safe time as required by applicable law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional taxable fringe benefits may include charitable donation matching and wellness stipends.
Base pay may vary depending on market location, job-related knowledge, skills, and experience. Total compensation also includes equity and performance-related bonuses for eligible employees.