Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Airflow @ 6
ChatGPT
Compliance
Dagster @ 6
Data Engineering @ 7
Data Pipelines
Data Science @ 4
ETL @ 6
Experimentation
Flink @ 4
Hadoop @ 4
Java @ 6
Marketing @ 4
Python @ 6
Scala @ 6
Security
Spark @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Statsig team at OpenAI builds and operates the experimentation platform that powers product development, measurement, and decision-making across the company. The team works with product, engineering, and infrastructure teams to ensure experiments are trustworthy, statistically rigorous, and scalable for frontier AI products.
The role focuses on building data pipelines and core tables that power analyses, safety systems, business decisions, product growth, and the prevention of bad actors. The Data Engineer will also collaborate with researchers working on ChatGPT and model training.
Responsibilities
- Design, build, and manage data pipelines, ensuring user event data is integrated into the data warehouse.
- Develop canonical datasets to track product metrics such as user growth, engagement, and revenue.
- Collaborate with Infrastructure, Data Science, Product, Marketing, Finance, and Research teams to understand data needs and provide solutions.
- Implement robust and fault-tolerant systems for data ingestion and processing.
- Participate in data architecture and engineering decisions.
- Ensure the security, integrity, and compliance of data according to industry and company standards.
Requirements
- 3+ years of experience as a data engineer and 8+ years of software engineering experience, including data engineering.
- Proficiency in at least one programming language commonly used in data engineering, such as Python, Scala, or Java.
- Experience with distributed processing technologies and frameworks such as Hadoop and Flink.
- Experience with distributed storage systems such as HDFS and S3.
- Expertise with ETL schedulers such as Airflow, Dagster, Prefect, or similar frameworks.
- Solid understanding of Spark, including the ability to write, debug, and optimize Spark code.
Compensation and Benefits
- Base compensation range: $293,000–$325,000 USD.
- Equity, performance-related bonuses for eligible employees, and benefits including medical, dental, and vision insurance.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, company holidays, and office closures.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.
The role uses a hybrid work model and values in-person collaboration for technical design, iteration, and cross-functional partnership. OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.