Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 3
Ansible @ 3
Azure @ 3
Change Management
GPU
Git @ 3
IaC
MySQL @ 3
Networking @ 6
Observability @ 3
PostgreSQL @ 3
Python @ 3
RDBMS
Reporting @ 3
SQL @ 3
Security @ 3
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
OpenAI's Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for large-scale data centers supporting AI research, products, and infrastructure partners.
The role focuses on designing, validating, and scaling resilient, secure, and scalable controls and operational technology (OT) network architectures for high-density AI data centers. The engineer will work across controls systems, OT infrastructure, telemetry, commissioning, deployment, and operations while partnering with mechanical, electrical, IT/networking, security, and external delivery teams.
Responsibilities
- Define controls, automation, and OT network requirements for AI data center campuses.
- Develop reference architectures, engineering standards, and reusable design templates.
- Review and develop basis-of-design and functional design documents, including OT network diagrams, IP/VLAN schemes, telemetry architectures, data flow diagrams, and commissioning requirements.
- Design OT and infrastructure network architectures covering physical and logical topology, IP addressing, subnetting, VLANs, routing, switching, redundancy, segmentation, firewall policy coordination, out-of-band management, monitoring, and remote access.
- Define day-two network operations requirements, including change management, configuration backups, golden configurations, monitoring thresholds, firmware lifecycle, rollback plans, and post-change validation.
- Partner with electrical, mechanical, IT/networking, security, and operations teams to align OT network systems with GPU deployments, campus-wide telemetry, and failure-domain isolation requirements.
- Define integration patterns and protocol requirements for BACnet/IP, BACnet MSTP, Modbus TCP/RTU, OPC UA, IEC-61850 MMS/GOOSE, MQTT, SNMP, syslog, NTP/PTP, IRIG-B, and vendor-specific interfaces.
- Lead technical evaluations of controls integrators, network equipment suppliers, design consultants, contractors, and commissioning agents.
- Review network equipment submittals, configurations, firmware assumptions, certifications, test reports, and quality documentation.
- Support factory witnessed testing, site acceptance testing, network readiness checks, failover testing, and integrated systems testing.
- Troubleshoot packet loss, latency, duplicate IPs, routing errors, firewall drops, protocol incompatibilities, time synchronization drift, and intermittent device communication failures.
Requirements
- 8+ years of relevant experience in controls engineering, industrial automation, OT networking, mission-critical facilities, or similar critical infrastructure environments.
- Strong expertise in resilient OT network architecture, implementation, troubleshooting, and lifecycle support.
- Experience with OT/IT boundary design, secure enterprise integration, firewall policy design, redundant topologies, out-of-band management, and monitoring.
- Hands-on Layer 3 OT network design experience, including IP addressing, subnetting, routing, VRFs, ACLs, inter-VLAN traffic control, and network segmentation.
- Hands-on Layer 2 security and switching experience, including MACsec, port security, loop prevention, and switch-level access control.
- Experience designing resilient OT network topologies using PRP, HSR, Cisco REP, RSTP/MSTP, and ring or star topologies.
- Experience designing resilient infrastructure network architectures using HSRP/VRRP, spine-leaf topologies, redundant uplinks, and failure-domain isolation.
- Experience with Cisco and Juniper switches and routers, Palo Alto firewalls, Rockwell Automation Stratix switches, Siemens Ruggedcom, or comparable industrial networking platforms.
- Experience with network management and observability platforms such as Cisco Catalyst Center/DNA Center, Palo Alto Panorama, Juniper Mist, industrial NMS tools, packet brokers, and OT monitoring platforms.
- Experience with industrial Ethernet, VPN tunneling, IPsec connectivity, and secure remote access.
- Experience with virtualized OT or controls server environments such as VMware vSAN, Microsoft Azure Stack HCI, Hyper-V, or comparable platforms.
- Experience with BACnet/IP, BACnet MSTP, Modbus TCP/RTU, OPC UA, IEC-61850 MMS/GOOSE, MQTT, SNMP, syslog, NTP/PTP, IRIG-B, and vendor-specific interfaces.
- Experience producing technical design documentation, commissioning plans, and acceptance test procedures.
- Experience with factory witnessed testing, site acceptance testing, failover testing, telemetry validation, protocol compatibility testing, and root-cause analysis.
- Ability to use logs, packet captures, and field observations to make technical decisions and communicate risk clearly.
- Bachelor's degree in Electrical Engineering, Computer Engineering, Network Engineering, Systems Engineering, or a related discipline.
Preferred Skills
- Master's degree in Electrical Engineering, Computer Engineering, Network Engineering, Systems Engineering, or a related discipline.
- Experience leading multi-campus OT network integration, commissioning, and operations across cross-functional teams, contractors, vendors, and delivery partners.
- Networking certifications such as Cisco CCNA/CCNP, Palo Alto PCNSA/PCNSE, Juniper JNCIA/JNCIS, or similar credentials.
- Cybersecurity certifications such as CISSP, GICSP, ISA/IEC 62443, CompTIA Security+, or similar credentials.
- Experience with network automation, Git-based configuration management, and Infrastructure as Code using Ansible, Terraform, Python, or similar tools.
- Experience with scripting, APIs, and automation workflows for OT network operations.
- Experience using AI agents or MCP-connected tools for telemetry analysis and troubleshooting.
- Experience with PostgreSQL, SQL Server, MySQL, or similar relational database systems used for OT telemetry, historian integrations, troubleshooting, and reporting.
Work Environment And Travel
- Periodic travel is required to data center campuses, vendors, labs, construction sites, commissioning activities, and controls/network cutovers.
- The engineer should be comfortable working in office, lab, construction, and live data center environments, including PPE, lockout/tagout, cyber hygiene, and change-control requirements.
- Work may include time-sensitive support during commissioning, startup, vendor testing, cutovers, network changes, telemetry issues, automation failures, and operational events.
Benefits
- Base salary of $257,000–$327,000 per year, plus equity.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax FSA, dependent care FSA, and commuter accounts.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
More jobs at OpenAI
GRC Program Manager, Assurance Engineering & Control Systems
OpenAI · San Francisco, United States
USD 216,000-252,000 per year
Android Systems Engineer, Consumer Devices
OpenAI · San Francisco, United States
USD 216,000-342,000 per year
Senior Staff Software Engineer, Identity
OpenAI · Mountain View, United States, San Francisco, United States
USD 345,000-405,000 per year
Analytics Engineer, GTM
OpenAI · San Francisco, United States, New York City, United States
USD 220,000-335,000 per year
Product Designer, Payments
OpenAI · San Francisco, United States
USD 245,000-310,000 per year
Similar jobs
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Network Operations Engineer, AI Networking
OpenAI · San Francisco, United States
USD 157,000-221,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year