Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Airflow @ 3
Compliance
Dagster @ 3
Data Engineering @ 6
Data Pipelines
Data Science @ 3
Databricks @ 3
ETL @ 3
Fivetran @ 3
Flink @ 3
Hadoop @ 3
Java @ 5
Python @ 5
Scala @ 5
Security
Snowflake @ 3
Spark @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About the Team
At OpenAI, we’re building the connective tissue between our mission and our people. People Innovation Labs is a fast-moving engineering team embedded in the People organization, focused on rethinking how we find and retain the best talent and empower everyone to do their best work. From recruiting to culture, we’re designing systems that give our People Team a significant edge by infusing OpenAI’s models and first-principles thinking into every aspect of our work. Our projects range from greenfield 0-1 products like OpenHouse (our internal knowledge hub) to AI-powered automations and scalable recruiting tools. We’re defining the future of work at OpenAI, creating a blueprint for how AI can supercharge productivity, culture, and innovation.
About the Role
We’re seeking a Data Engineer to build data-intensive systems that will power People Innovation Labs’ internal products and enable the People Analytics function to do their best work. These data pipelines are crucial for our build-out of people products backed by business systems of record and for ongoing people data analytics.
One example of an employee-facing product you’ll help us build is OpenHouse, which serves as a culture and communication hub and an organization-wide front door into all other aspects of People Innovation Labs’ work. OpenHouse and other products in our portfolio are built by full stack product engineers who are deeply curious about culture, recruiting and people development, and want to know everything from the business strategy and metrics down through the code that gets us there. In this role, you will work with People Innovation Labs leadership and software engineers and the People Analytics team to build the data systems that enable this work.
Responsibilities
- Design, build and manage people data pipelines, ensuring all data is seamlessly integrated into our Databricks warehouse.
- Develop canonical datasets to track key people metrics and People Innovation Labs product metrics.
- Work collaboratively with various teams, including, Data Platform, Data Science, People Analytics, and Compensation and Equity to understand their data needs and provide solutions.
- Implement robust and fault-tolerant systems for data ingestion and processing.
- Participate in data architecture and engineering decisions, bringing your strong experience and knowledge to bear as the primary data engineering expert on the team.
- Ensure the security, integrity, and compliance of data according to industry and company standards.
Requirements
- Have 3+ years of experience as a data engineer and 8+ years of any software engineering experience (including data engineering).
- Proficiency in at least one programming language commonly used within Data Engineering, such as Python, Scala, or Java.
- Experience with data warehousing technologies such as Databricks and Snowflake, and expertise with ETL schedulers such as Fivetran, Airflow, Dagster, Prefect, or similar.
- Experience with distributed processing technologies and frameworks, such as Spark, Hadoop, Flink and distributed storage systems (e.g., HDFS, S3).
Benefits
- Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings Accounts
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit)
- 401(k) retirement plan with employer match
- Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks)
- Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees
- 13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time (1 hour per 30 hours worked, or more, as required by applicable state or local law)
- Mental health and wellness support
- Employer-paid basic life and disability coverage
- Annual learning and development stipend to fuel your professional growth
- Daily meals in our offices, and meal delivery credits as eligible
- Relocation support for eligible employees
- Additional taxable fringe benefits, such as charitable donation matching and wellness stipends, may also be provided.
More details about our benefits are available to candidates during the hiring process.