Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Distributed Systems @ 4
GPU
Machine Learning @ 4
Networking
Robotics
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The OpenAI Robotics team focuses on unlocking general-purpose robotics and advancing AGI-level intelligence in dynamic, real-world settings. The team works across the model stack, integrating hardware and software across a broad range of robotic form factors.
As a Senior Software Engineer, ML Systems & Training Infrastructure, you will be a hands-on engineering force multiplier for the robotics team. You will help maintain the training framework and surrounding infrastructure, review and improve code, debug failures across machine learning systems and infrastructure, and unblock researchers and engineers when training workflows encounter issues.
The role is based in San Francisco, California, and requires working in the office five days per week.
Responsibilities
- Review, improve, and clean up code across training frameworks and adjacent infrastructure.
- Identify risky or low-quality changes before they are merged and raise the code quality bar without slowing the team down.
- Debug issues across machine learning training systems, GPUs, clusters, networking, and related infrastructure.
- Help researchers and engineers resolve broken training jobs, unreliable workflows, and brittle internal tooling.
- Improve the reliability, maintainability, and usability of the robotics team's training framework.
- Move quickly on practical engineering problems that directly affect team velocity.
Requirements
- Strong software engineering fundamentals and excellent code review judgment.
- Experience with machine learning systems, training frameworks, GPUs, distributed systems, infrastructure, or similarly complex technical environments.
- Ability to read and debug unfamiliar codebases quickly and identify root causes.
- Ability to ship high-quality code with strong velocity and pragmatic judgment.
- A low-ego, responsive approach focused on helping researchers and engineers move faster.
- Preference for highly effective hands-on individual contributor work over broad, process-heavy initiatives.
- Experience reviewing messy, fast-moving, or AI-generated codebases.
Compensation and Benefits
- Base salary: $295,000–$380,000 USD per year.
- Equity, performance-related bonuses for eligible employees, and benefits including medical, dental, and vision insurance; flexible spending accounts; a 401(k) with employer match; paid parental, medical, and caregiver leave; paid time off and holidays; mental health and wellness support; life and disability coverage; a learning and development stipend; office meals; and relocation support for eligible employees.
OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.