Performance Modeling Lead

at OpenAI
USD 293,000-385,000 per year
SENIOR
✅ Hybrid
✅ Relocation

Tech Stack

AI @ 7 Communication @ 6 Distributed Systems @ 4 GPU InfiniBand @ 3 Machine Learning @ 4 Mentoring @ 4 NVLink @ 3 Networking @ 4

Details

OpenAI’s Hardware organization develops system and infrastructure solutions for advanced AI workloads, working across research, software, and external hardware partners. The team focuses on understanding and optimizing performance across the full system stack through rigorous quantitative analysis of real-world workloads.

The Performance Modeling Lead will build and lead a small, high-impact team responsible for answering forward-looking architectural questions across AI infrastructure systems. The role involves developing modeling frameworks and methodologies to evaluate system-level tradeoffs and guide reference architectures, vendor designs, and long-term infrastructure strategy. This position is based in San Francisco, California, and follows a hybrid work model requiring three days in the office per week.

Responsibilities

  • Build and own a performance modeling framework and toolchain to evaluate AI systems across multiple levels of abstraction.
  • Analyze and quantify architectural tradeoffs across compute, memory, networking, storage, and system topology.
  • Develop performance models to guide decisions regarding:
    • Scale-up versus scale-out architectures.
    • Interconnect and network design.
    • Memory hierarchy and system balance.
  • Translate modeling outputs into clear recommendations for internal teams and external hardware vendors.
  • Influence reference designs and vendor roadmaps through data-driven insights.
  • Partner closely with machine learning, systems, and hardware teams to understand workload characteristics and requirements.
  • Lead and grow a small team of two to three engineers, setting technical direction and maintaining high standards for modeling rigor.
  • Continuously improve modeling fidelity by validating models against real system behavior and measurements.

Requirements

  • Experience owning or building performance modeling frameworks used to drive real system design decisions.
  • Deep knowledge of AI/ML workloads, including training and/or inference at scale.
  • Understanding of system-level tradeoffs across compute, memory, and networking in large-scale distributed systems.
  • Comfort working across abstraction layers, from workload behavior to hardware implementation.
  • Experience using analytical or simulation-based modeling to inform architectural decisions.
  • Ability to operate in ambiguous problem spaces and turn open-ended questions into structured analysis.
  • Clear communication skills and the ability to influence internal teams and external partners.

Preferred Skills

  • Experience working with hardware vendors, including ODM/JDM, silicon, and networking vendors.
  • Background in data center infrastructure or hyperscale systems.
  • Familiarity with accelerators such as GPUs and ASICs, and interconnects such as NVLink, InfiniBand, and Ethernet.
  • Experience influencing hardware roadmaps or reference architectures.
  • Prior experience leading or mentoring engineers.

Benefits

  • Base salary of $293,000–$385,000 per year, plus equity.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent care, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, and paid sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily meals in offices and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

More jobs at OpenAI

Similar jobs