Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 7
Debugging @ 4
Leadership @ 6
Machine Learning
Networking @ 4
Observability @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for advanced AI workloads. The Silicon & Systems team is seeking a Technical Lead to lead deployment and operations for custom silicon and associated systems, bringing platforms from lab validation into production data center environments and ensuring operational readiness and reliability at scale.
Responsibilities
- Lead a team responsible for deployment and operations of custom silicon and systems in data center environments.
- Own the path from hardware bring-up and validation through production deployment, operational readiness, and sustained fleet support.
- Partner with silicon, systems, software, infrastructure, networking, data center, supply chain, and external partner teams.
- Define deployment processes, operational playbooks, technical readiness criteria, escalation paths, and reliability practices.
- Drive execution across lab bring-up, rack and system integration, data center deployment, fleet monitoring, debugging, and issue resolution.
- Contribute hands-on through architecture reviews, deployment planning, failure analysis, operational debugging, and system-level decision-making.
- Identify and address gaps in tooling, observability, automation, validation coverage, and operational processes.
- Establish metrics for deployment readiness, reliability, performance, maintainability, and operational health.
- Build an engineering culture grounded in ownership, technical rigor, operational excellence, and high-velocity execution.
- Ensure custom hardware platforms can be deployed and operated reliably, repeatably, and safely at scale.
- Contribute to the architecture and design of future machine learning systems.
Requirements
- 8+ years of engineering experience in hardware systems, infrastructure, data center deployment, production operations, systems engineering, silicon bring-up, or related technical domains.
- Strong technical depth in one or more of hardware deployment, data center operations, rack-scale systems, silicon bring-up, systems validation, fleet operations, reliability engineering, infrastructure automation, or hardware/software integration.
- Experience bringing complex hardware systems from development or validation into production environments.
- Experience working closely with silicon, systems, software, infrastructure, networking, or data center teams.
- Experience with deployment planning, operational readiness, incident response, debugging, and root-cause analysis for production systems.
- Experience building tooling, automation, observability, or operational processes that improve deployment quality and fleet reliability.
- Demonstrated ability to hire, develop, and lead senior technical talent.
- Ability to move between people leadership, technical strategy, and hands-on operational problem-solving.
- Strong written and verbal communication skills in high-urgency, cross-functional technical environments.
- Experience working in fast-moving environments.
- Ability to mentor engineers while remaining deeply engaged in technical execution and to lead through ambiguity.
- Candidates may need to meet certain legal status requirements under U.S. export control laws and regulations.
Benefits
- Equity, performance-related bonuses for eligible employees, and a range of employee benefits.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for healthcare, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
More jobs at OpenAI
Analytics Engineer, GTM
OpenAI · New York City, United States, San Francisco, United States
USD 220,000-335,000 per year
Cyber Operations Strategist, Critical Harm Operations
OpenAI · Toronto, Canada
CAD 140,000-188,000 per year
Product Manager, Multimodal Safety
OpenAI · San Francisco, United States
USD 293,000-325,000 per year
Forward Deployed Engineer (FDE), Legal - New York City
OpenAI · New York City, United States, San Francisco, United States
USD 162,000-280,000 per year
Product Manager, Learning
OpenAI · San Francisco, United States
USD 293,000-385,000 per year
Similar jobs
Principal Software Engineer – Infrastructure
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Senior MLOps Engineer - DSX Enablement
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Senior Technical Program Manager, AI Infrastructure and Capacity Operations
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Engineering Manager, Agentic GenAI Platform
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year