Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
API @ 6
AWS @ 1
Algorithms @ 7
Codex
Data Pipelines @ 4
Data Structures @ 7
Deep Learning @ 4
Distributed Systems @ 4
Kubernetes @ 1
LLM
Machine Learning @ 4
PyTorch @ 4
Python @ 7
Reinforcement Learning @ 3
Rust @ 7
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI’s API Multicloud team extends the API platform into strategic cloud environments, starting with AWS. The team enables API technologies in AWS-native environments in partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. Its work includes AWS-hosted Codex, model customization and post-training as a service, and stateful runtime environments for agentic workloads.
The role focuses on building and improving AI systems that help strategic partners adapt OpenAI models to cloud-native use cases. Responsibilities span post-training workflows, evaluation, data pipelines, model behavior, and API and infrastructure integration. The role works across partner needs and core ML systems to diagnose training and evaluation issues and improve the underlying platform.
Responsibilities
- Partner with strategic customers and internal teams to define target model behaviors, diagnose failure modes, and translate real-world needs into training, evaluation, and system requirements.
- Build and scale production ML systems for model customization, post-training, and fine-tuning-as-a-service workflows.
- Investigate whether training and customization workflows produce the intended outcomes, and identify improvements to data, evaluation, training, or infrastructure.
- Partner with backend and infrastructure engineers to integrate ML capabilities into AWS-native API environments.
- Propose and implement improvements to post-training systems, tooling, APIs, and developer workflows based on partner deployments.
- Work with Research and Applied teams to bring model improvements, training workflows, and evaluation best practices into production.
- Help design systems that allow strategic partners and enterprise customers to safely customize OpenAI models.
- Debug and improve systems spanning model behavior, training data, APIs, distributed infrastructure, and customer-facing product surfaces.
- Operate with high ownership in a 0-to-1 environment where requirements are ambiguous, systems evolve quickly, and reliability is important.
Requirements
- Master’s or PhD in Computer Science, Machine Learning, or a related field, or equivalent practical experience.
- 7+ years of professional engineering experience in relevant ML, infrastructure, or product-driven engineering roles.
- Strong ML engineering experience building, training, fine-tuning, evaluating, or deploying production AI systems.
- Hands-on experience with deep learning, transformer models, and frameworks such as PyTorch or TensorFlow.
- Familiarity with training and fine-tuning large language models, including supervised fine-tuning, distillation, preference optimization, reinforcement learning, or other post-training techniques.
- Strong software engineering fundamentals, including data structures, algorithms, systems design, and high-quality production code in Python, Rust, or similar languages.
- Experience with model customization, evaluation systems, data pipelines, distributed systems, cloud infrastructure, or production ML platform tradeoffs.
- Ability to work across model behavior, APIs, and infrastructure while collaborating with Research, Safety, product engineering, infrastructure, and external technical partners.
- Comfort working through ambiguity, owning problems end-to-end, and learning whatever is needed to complete the work.
- Bonus experience with AWS, Kubernetes, agents, tool use, runtime environments, AI developer platforms, or speech models.
Benefits
- Base salary of $295,000–$445,000 per year, plus equity.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax Flexible Spending Accounts and commuter benefits.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and paid sick or safe time as required by law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and is committed to providing reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.