Backend Software Engineer (Evals)

at OpenAI
USD 230,000-385,000 per year
SENIOR
✅ On-site
✅ Relocation

Tech Stack

AI @ 4 API @ 4 Data Science @ 4 Distributed Systems @ 4 FastAPI @ 6 LLM Machine Learning @ 4 PostgreSQL Python @ 6

Details

The Support Automation team applies AI models to real-world challenges, developing automation products that improve work across OpenAI. The team combines rapid prototyping with a focus on long-term quality, reliability, and reusable solutions.

The role focuses on designing and building evaluation infrastructure to measure the quality of OpenAI’s support automation. You will build robust backend systems and services that support how knowledge is created, accessed, and applied across OpenAI, working closely with Data Science and Research partners to design evaluations at scale.

Responsibilities

  • Design reliable, reproducible, and extensible evaluation pipelines.
  • Build infrastructure for continuous evaluation monitoring, including regression and drift monitoring and robust golden datasets.
  • Develop feedback loops that improve support automation.
  • Design, build, and maintain backend services and APIs for intelligent automation and knowledge systems.
  • Integrate and structure data across internal platforms, transforming it into formats optimized for downstream systems and AI workflows.
  • Collaborate with data, research, and engineering teams to integrate OpenAI models into high-leverage workflows.
  • Own the full development lifecycle of new backend systems and internal platform capabilities.
  • Build scalable and maintainable systems while rapidly iterating on new ideas.

Requirements

  • 4+ years of backend engineering experience at product-driven companies, excluding internships.
  • Proficiency with backend technologies; the technology stack includes Python, FastAPI, and Postgres.
  • Experience designing and scaling distributed systems, APIs, or data processing pipelines.
  • Experience building AI agents or applications, including designing evaluations and improving performance through prompting or scaffolding.
  • Familiarity with evaluation methods for large language models and patterns such as multi-agent workflows, tool use, and long context.
  • Experience creating production evaluations and/or measuring the performance of machine learning or large language models at scale.
  • A pragmatic mindset and comfort shipping iteratively while building toward a long-term vision.

Benefits

  • Equity, performance-related bonuses for eligible employees, and comprehensive benefits.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent care, parking, and transit expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

OpenAI is an equal opportunity employer. Background checks and reasonable accommodations are handled in accordance with applicable law.

More jobs at OpenAI

Similar jobs