Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 5
Data Analysis @ 3
Data Science @ 5
Distributed Systems
Experimentation @ 6
GPU
Machine Learning
Mathematics @ 3
Python @ 6
Reinforcement Learning
SQL @ 6
Statistics @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About the Role
OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI models.
We’re looking for a Data Scientist to partner closely with Capacity Systems Engineering, Infrastructure, Product, and Research to optimize inference capacity across our global GPU fleet. This role combines statistical modeling, large-scale data analysis, forecasting, and systems thinking to drive critical decisions around infrastructure investments, performance-efficiency trade-offs, and customer experience.
You’ll transform complex operational data into actionable insights that directly influence how OpenAI allocates and scales one of the world’s largest AI compute environments.
Responsibilities
- Build statistical and machine learning models to profile and improve GPU utilization, latency, throughput, and overall fleet efficiency.
- Develop forecasting models for inference demand across products, regions, and model families.
- Analyze production workloads to identify latency bottlenecks and capacity constraints, highlighting optimization opportunities.
- Partner with Capacity Systems Engineering to inform infrastructure planning and long-term GPU investment strategies.
- Design experiments and simulations to evaluate scheduling policies, serving strategies, and infrastructure tradeoffs.
- Build dashboards and operational metrics that enable leadership to make data-driven capacity decisions.
- Collaborate with Product, Research, Finance, and Infrastructure teams to align compute planning with business growth and model roadmaps.
- Communicate technical findings clearly to both engineering teams and executive leadership.
Requirements
- MS or PhD in Statistics, Computer Science, Operations Research, Applied Mathematics, Economics, or related quantitative discipline (or equivalent industry experience).
- 5+ years of experience working in the infrastructure data science space.
- Strong expertise in Python and SQL.
- Experience building forecasting, optimization, or predictive models.
- Strong understanding of experimentation, statistical inference, and causal analysis.
- Experience communicating analytical insights to executive stakeholders.
Preferred Skills
- Capacity planning
- Distributed systems
- AI infrastructure
- Datacenter design and buildout
- Queueing theory
- Time-series forecasting
- Operations research
- Supply-demand modeling
- Reinforcement learning for resource allocation
- Cost optimization
Benefits
- Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings Accounts
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit)
- 401(k) retirement plan with employer match
- Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks)
- Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees
- 13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time (1 hour per 30 hours worked, or more, as required by applicable state or local law)
- Mental health and wellness support
- Employer-paid basic life and disability coverage
- Annual learning and development stipend to fuel your professional growth
- Daily meals in our offices, and meal delivery credits as eligible
- Relocation support for eligible employees
- Additional taxable fringe benefits, such as charitable donation matching and wellness stipends, may also be provided.
More details about our benefits are available to candidates during the hiring process.
This role is at-will and OpenAI reserves the right to modify base pay and other compensation components at any time based on individual performance, team or company results, or market conditions.