Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
API @ 4
AWS @ 4
Azure @ 4
BGP @ 7
Change Management @ 4
Cumulus Linux @ 4
GPU @ 4
Git @ 4
Grafana @ 4
HPC @ 6
Linux @ 4
Networking @ 7
Observability
Prometheus @ 4
Python @ 4
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of large-scale AI infrastructure networks. The team operates production AI networks across Industrial Compute’s data centers and works with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads.
The role combines hands-on production operations with automation, observability, and incident response across a global AI network. The successful candidate will operate high-availability data center, cloud, AI, or HPC networks and work across physical-layer troubleshooting, routing and fabric behavior, change execution, and root-cause analysis.
Responsibilities
- Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute’s data centers.
- Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR).
- Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks.
- Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact.
- Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance.
- Support new AI cluster deployments, data center expansions, and infrastructure migrations in partnership with deployment and engineering teams.
- Partner with cloud service providers, colocation providers, Smart Hands teams, and hardware vendors to maintain production infrastructure.
- Perform root-cause analysis (RCA) for production incidents and drive permanent corrective actions that eliminate recurring issues.
- Build and maintain monitoring, telemetry, dashboards, and alerting to improve network observability and proactive issue detection.
- Develop and improve operational runbooks, playbooks, troubleshooting documentation, and standard operating procedures.
- Automate repetitive operational tasks using Python and infrastructure automation frameworks to reduce toil and improve efficiency.
- Continuously identify opportunities to improve service reliability, scalability, operational maturity, and engineering efficiency.
Requirements
- Bachelor’s degree in Computer Science, Network Engineering, or a related discipline, or equivalent practical experience.
- 5+ years of experience operating large-scale data center, cloud, AI, or HPC network infrastructure.
- Experience supporting production network environments with high-availability requirements.
- Hands-on experience with one or more of Cisco NX-OS, Arista EOS, NVIDIA Spectrum / Cumulus Linux, or Juniper JunOS.
- Strong knowledge of Layer 2 and Layer 3 networking, BGP, OSPF, ECMP, MLAG, LACP, VRFs, and VLANs.
- Experience troubleshooting physical infrastructure, including fiber optics, transceivers, DAC/AOC cables, and high-speed Ethernet links.
- Experience performing software upgrades, hardware maintenance, and production change management.
- Excellent analytical and troubleshooting skills, with the ability to communicate technical risk clearly across teams.
Preferred Skills
- Experience operating AI or High-Performance Computing (HPC) network environments.
- Experience with NVIDIA AI networking technologies and GPU infrastructure.
- Experience supporting RoCE v2 or RDMA-based Ethernet fabrics, including Priority Flow Control (PFC), Explicit Congestion Notification (ECN), Data Center Quantized Congestion Notification (DCQCN), Quality of Service (QoS), and lossless Ethernet networking.
- Experience supporting 100G, 200G, 400G, and 800G Ethernet networks.
- Experience with GPU platforms including NVIDIA HGX, DGX, GB200, or equivalent AI infrastructure.
- Experience supporting distributed storage environments such as VAST, DDN, or similar technologies.
- Experience working with cloud service providers such as AWS, Azure, or Google Cloud, and with third-party colocation providers.
- Experience with network monitoring and telemetry technologies, including Prometheus, Grafana, gNMI, streaming telemetry, SNMP, or similar tools.
- Experience developing automation using Python, Git, REST APIs, Terraform, or similar automation frameworks.
Work Environment and On-Call
- Participate in a 24x7 on-call rotation supporting mission-critical AI infrastructure.
- Support time-sensitive production incidents, maintenance windows, capacity expansions, and network changes with a focus on service availability and minimal customer impact.
- Travel up to 30% to data center locations for new turnups and acceptance activities, as needed.
Benefits
- Base salary range of $157,000–$221,000, plus equity.
- Medical, dental, and vision insurance for employees and their families, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and paid sick or safe time as applicable.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily meals in offices and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional taxable fringe benefits may be provided, including charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and is committed to providing reasonable accommodations to applicants with disabilities. Background checks will be administered in accordance with applicable law.