Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms
Communication @ 6
Compliance
Distributed Systems @ 6
Kubernetes @ 5
Software Development @ 5
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Data Acquisition team within the Foundations organization at OpenAI is responsible for data collection supporting model training operations. The team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams.
Responsibilities
- Own and lead engineering projects involving data acquisition, including web crawling, data ingestion, and search.
- Collaborate with Data Processing, Architecture, and Scaling teams to ensure smooth data flow and system operability.
- Work closely with the legal team on compliance and data privacy matters.
- Develop and deploy highly scalable distributed systems capable of handling petabytes of data.
- Architect and implement algorithms for data indexing and search capabilities.
- Build and maintain backend services for data storage, including key-value databases and synchronization.
- Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks.
- Conduct and analyze data experiments to provide insights into system performance.
Requirements
- BS, MS, or PhD in Computer Science or a related field.
- 4+ years of industry experience in software development.
- Experience with large web crawlers is a plus.
- Strong expertise in large stateful distributed systems and data processing.
- Proficiency in Kubernetes and Infrastructure-as-Code concepts.
- Willingness and enthusiasm for trying new approaches and technologies.
- Ability to handle multiple tasks and adapt to changing priorities.
- Strong written and verbal communication skills.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. OpenAI develops AI systems and seeks to safely deploy them through its products.
OpenAI is an equal opportunity employer and does not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristics.
Background checks for applicants will be administered in accordance with applicable law. Qualified applicants with arrest or conviction records will be considered consistent with applicable laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act for US-based candidates.
OpenAI is committed to providing reasonable accommodations to applicants with disabilities. Additional information is available through OpenAI's employment and applicant privacy policies.