Software Engineer, Fleet Management

at OpenAI
USD 230,000-490,000 per year
MIDDLE
✅ Hybrid
✅ Relocation

Tech Stack

AI @ 3 Algorithms CI/CD @ 3 ChatGPT Chef @ 3 GPU Kubernetes @ 3 LLM Linux @ 3 Networking Terraform @ 3

Details

The Fleet team at OpenAI supports the computing environment that powers cutting-edge research and product development. The team oversees large-scale systems spanning data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. This work enables OpenAI's models to operate at scale, supporting internal research and products such as ChatGPT, with a focus on safety, reliability, and responsible AI deployment.

About the Role

The Software Engineer, Operating Systems & Orchestration will build systems to manage hardware, configurations, vendors, and the people interacting with infrastructure. The role involves designing and developing solutions that integrate individual nodes and servers into unified clusters, helping advance AI research by streamlining the overall research user experience.

This role is based in San Francisco, California. OpenAI uses a hybrid work model requiring three days in the office per week and offers relocation assistance to new employees.

Responsibilities

  • Design and build systems to manage cloud and bare-metal fleets at scale.
  • Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms.
  • Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows.
  • Automate infrastructure processes to reduce repetitive toil and improve system reliability.
  • Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack.
  • Continuously improve tools, automation, processes, and documentation to enhance operational efficiency.

Requirements

  • Strong software engineering skills with experience in large-scale infrastructure environments.
  • Broad knowledge of cluster-level systems, such as Kubernetes, CI/CD pipelines, Terraform, and cloud providers.
  • Deep expertise in server-level systems, including systems, containerization, Chef, Linux kernels, firmware management, and host routing.
  • Passion for optimizing the performance and reliability of large compute fleets.
  • Ability to thrive in dynamic environments and solve complex infrastructure challenges.
  • Commitment to automation, efficiency, and continuous improvement.

Benefits

  • Base salary range of $230,000–$490,000 per year.
  • Equity, performance-related bonuses for eligible employees, and additional benefits.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax FSA, dependent care FSA, commuter, parking, and transit accounts.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.

More jobs at OpenAI

Similar jobs