Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Ansible
Azure @ 4
Chef @ 4
IaC
Observability @ 6
SRE
Security @ 4
Terraform
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Infrastructure Engineering function sits within IT and is responsible for reliably building, deploying, and operating critical on-premises and hybrid environments that power internal services and critical R&D environments.
This is an early, high-leverage technical role focused on applying Site Reliability Engineering discipline to environments where uptime, safety, recoverability, and security are non-negotiable. The role replaces bespoke, one-off infrastructure with standardized infrastructure-as-code building blocks that improve reliability and operational leverage as the company scales.
The role involves designing, building, and operating reliable, secure, and scalable infrastructure supporting identity, access, endpoint, and shared platform services. The successful candidate will be a senior technical owner for infrastructure and identity systems end to end, including architecture, implementation, policy enforcement, upgrades, recovery, and day-to-day operations.
The role is based at the San Francisco headquarters and requires in-office presence.
Responsibilities
- Design, build, and operate reliable infrastructure across on-premises, hybrid, shared, and product-adjacent environments.
- Establish standardized infrastructure patterns that replace bespoke implementations with repeatable, auditable, secure-by-default systems.
- Own the lifecycle of critical infrastructure platforms, including provisioning, deployment, upgrades, patching, recovery, and long-term reliability.
- Build infrastructure as code and configuration management using tools such as Terraform, Chef, and Ansible.
- Mature identity-adjacent and policy-enforced infrastructure, including Microsoft Entra and Azure management patterns.
- Build observability, alerting, and incident response mechanisms that improve availability, recoverability, and operational confidence.
- Automate high-toil and high-risk workflows with guardrails, progressive rollout patterns, and safe rollback paths.
- Translate incidents, design reviews, and operational learnings into durable fixes, reusable patterns, and stronger technical standards.
Requirements
- 10+ years of hands-on experience operating and architecting mission-critical infrastructure in high-reliability environments.
- Experience running security infrastructure.
- Experience serving as the senior technical owner for the design and maturation of complex on-premises, hybrid, or cloud-integrated systems, with durable architectural patterns used by multiple teams.
- Ability to apply Site Reliability Engineering principles at scale using observability, automation, and incident learnings to materially reduce risk and operational toil.
- Comfort operating in ambiguous environments, making sound architectural decisions under pressure while staying close to technical detail.
- Ability to influence cross-functional partners across security, identity, network, and platform teams through architecture, implementation, operational data, and clear technical writing.
- Experience operating infrastructure for R&D or specialized labs, manufacturing, or other safety-critical environments where uptime and recoverability are essential.
- Experience with fleet, endpoint, or virtual desktop platforms such as FleetDM, Chef, or Azure Virtual Desktop.
- Experience partnering closely with identity or security engineering teams on hardened, policy-enforced infrastructure at scale.
Benefits
- Base salary of $293,000–$385,000 per year, plus equity.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and sick or safe time as required by applicable law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.