Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Airflow @ 4
Dagster @ 4
Data Engineering @ 6
Data Pipelines @ 6
ELT @ 4
ETL @ 4
Observability @ 4
Python @ 6
Reporting @ 3
SQL @ 7
Snowflake @ 4
Spark @ 4
dbt @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Scaling Analytics team serves as the data backbone for OpenAI's infrastructure organization, enabling leaders and operators to make informed decisions across infrastructure deployment, hardware operations, supply chain, capacity planning, and site execution.
As OpenAI's Industrial Compute expands across global data center campuses, the team develops data models, pipelines, metrics, and reporting systems that transform fragmented operational data into actionable insights.
The Data Engineer will help build and scale the analytical foundations that support OpenAI's infrastructure organization. The role partners with Hardware Operations, Capacity Planning, Supply Chain, Infrastructure Delivery, Finance, and Engineering teams to create reliable data products for critical operational and strategic decisions.
Responsibilities
- Design, build, and maintain scalable data pipelines supporting infrastructure deployment, operations, capacity planning, and supply chain functions.
- Develop trusted datasets and reporting systems providing visibility into hardware inventory, deployment status, site readiness, capacity utilization, and operational performance.
- Partner with cross-functional stakeholders to define metrics, establish data standards, and improve decision-making across infrastructure organizations.
- Create scalable data models that enable consistent reporting and analytics across multiple data sources and operational systems.
- Improve data quality, lineage, observability, and governance practices across critical infrastructure datasets.
- Support executive reporting, operational reviews, forecasting exercises, and strategic planning initiatives through reliable analytical foundations.
- Collaborate with engineering teams to integrate new data sources and operational telemetry into existing analytics ecosystems.
- Build solutions that reduce manual reporting efforts and improve the speed and accuracy of infrastructure decision-making.
- Document systems, processes, and analytical frameworks to improve long-term maintainability and organizational resilience.
Requirements
- 5+ years of experience building and maintaining production data pipelines and analytical systems.
- Strong proficiency in SQL and experience designing scalable data models.
- Proficiency in Python or another programming language commonly used for data engineering.
- Experience with modern data warehouses such as Snowflake, BigQuery, or Redshift.
- Experience with orchestration frameworks such as Airflow or Dagster.
- Experience designing reliable ETL/ELT workflows with a focus on maintainability, performance, and operational excellence.
- Experience partnering with cross-functional stakeholders to translate business requirements into technical solutions.
- Experience implementing data quality checks, monitoring, and observability practices in production environments.
Preferred Skills
- Experience supporting infrastructure, hardware operations, supply chain, manufacturing, logistics, or capacity planning organizations.
- Familiarity with large-scale operational telemetry and business-critical reporting environments.
- Experience with distributed processing frameworks such as Spark.
- Experience with transformation frameworks such as dbt.
- Experience developing executive reporting and operational review metrics.
- Experience operating in fast-paced, ambiguous environments with evolving priorities.
- Interest in building analytical foundations that support large-scale AI infrastructure deployments.
Benefits
- Equity, performance-related bonuses for eligible employees, and comprehensive benefits.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health and dependent care expenses, parking, and transit.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, office closures, and paid sick or safe time as required by applicable law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily meals in offices and meal delivery credits as eligible.
- Relocation support for eligible employees.