Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CloudFormation @ 3
Communication @ 6
Datadog @ 3
Design Patterns
GPU
Grafana @ 3
IaC @ 3
Kubernetes @ 3
Microservices @ 3
Observability @ 3
Prometheus @ 3
Security @ 3
Splunk @ 3
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Join the engineering teams that bring OpenAI’s ideas safely to the world.
The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. The team seeks to learn from deployment and distribute the benefits of AI while ensuring that this powerful tool is used responsibly and safely. Safety is more important than unfettered growth.
About the Role
As OpenAI continues to grow, it is looking for experienced, problem-solving engineers to ensure its systems scale. The role focuses on rapidly iterating on performant and reliable products while delivering technology safely to millions of users around the world.
You will help ensure the reliability, scalability, and performance of rapidly evolving infrastructure. You will work closely with software engineers, product managers, data scientists, researchers, and designers to build and maintain resilient systems capable of handling a growing user base and workload.
Responsibilities
- Design and implement solutions to ensure infrastructure scalability as demands increase.
- Build and maintain load, chaos, and synthetic testing software used by development teams to improve system reliability.
- Build and maintain automation tools to streamline repetitive tasks and improve reliability.
- Build and maintain platforms for CPU/storage, GPU, and network lifecycle management to drive efficiency, accountability, and dynamic resource optimization.
- Implement fault-tolerant and resilient design patterns to minimize service disruptions.
- Develop and maintain service level objectives (SLOs) and service level indicators (SLIs) to measure and ensure system reliability.
- Partner with researchers, engineers, product managers, and designers to bring new features and research capabilities to users.
- Participate in an on-call rotation to respond to critical incidents and help ensure 24/7 system availability.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent work experience.
- Proven experience as a software engineer focused on reliability or in a similar role at a fast-paced, rapidly scaling company.
- Strong proficiency in cloud infrastructure.
- Proficiency in programming languages.
- Experience with containerization technologies and container orchestration platforms such as Kubernetes.
- Knowledge of Infrastructure as Code (IaC) tools such as Terraform or CloudFormation.
- Excellent problem-solving and troubleshooting skills.
- Strong communication and collaboration skills.
- Experience with observability tools such as Datadog, Prometheus, Grafana, and Splunk.
- Experience with microservices architecture and service mesh technologies.
- Knowledge of security best practices in cloud environments.
- Experience accelerating engineering reliability through effective tooling and systems.
- Ability to own problems end-to-end, address performance bottlenecks, and collaborate cross-functionally on reliability and scalability.
This role is exclusively based in OpenAI’s San Francisco headquarters.
Benefits
- Medical, dental, and vision insurance for employees and their families, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental leave, medical leave, and caregiver leave.
- Paid time off and paid company holidays.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional taxable fringe benefits may include charitable donation matching and wellness stipends.
OpenAI is an equal opportunity employer and is committed to providing reasonable accommodations to applicants with disabilities.