Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Data Pipelines
Deep Learning @ 3
Machine Learning @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Data Understanding team creates high-quality datasets and quantized representations for OpenAI. Its work includes synthesizing multimodal data, building vector-quantized representations, and processing, filtering, deduplicating, performing quality control, and tokenizing data for large-model training runs.
The role focuses on advancing how OpenAI prepares, curates, synthesizes, and understands multimodal data at scale. The work includes research and production problems involving images, audio, and video; multimodal content synthesis and supervision; noisy data pipelines; quality filters; model-based data preparation; and measuring whether dataset changes improve model performance.
Responsibilities
- Choose and pursue important multimodal data problems.
- Own and drive a research agenda from problem selection through long-running work and measurable impact.
- Develop new or improved machine learning ideas.
- Synthesize multimodal content, including images, audio, and video, and their supervisions.
- Improve noisy data pipelines and build better quality filters.
- Use models to automate data preparation.
- Measure whether changes in datasets improve model performance.
- Collaborate using OpenAI's empirical research approach.
Requirements
- A strong track record of new or improved machine learning ideas demonstrated through publications, projects, or applied research.
- Ability to own and drive a research agenda.
- Excitement about empirical and collaborative research.
Preferred Qualifications
- Experience with multimodal learning, audio, vision, video, synthetic data, or data-centric machine learning.
- Thoughtfulness about AI's impact, including privacy, provenance, and data quality.
- Experience building high-performance deep learning or large-scale data processing systems.
Benefits
- Medical, dental, and vision insurance for employees and their families, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and paid sick or safe time as required by applicable law.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Potential additional taxable fringe benefits, including charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer. Background checks and reasonable accommodations are handled in accordance with applicable law.