Researcher, Alignment Interpretability

at OpenAI
USD 295,000-500,000 per year
MIDDLE
✅ On-site
✅ Relocation

Tech Stack

AI @ 3 Deep Learning @ 6 Machine Learning @ 3 Python @ 5

Details

The Interpretability team studies internal representations of deep learning models to understand model behavior and engineer models with more understandable representations. The team is particularly interested in applying this understanding to ensure the alignment of powerful AI systems and works collaboratively in a curiosity-driven environment.

The researcher will develop and carry out a research plan in mechanistic interpretability, collaborating with a highly motivated team to help ensure future models remain safe as they grow in capability.

Responsibilities

  • Develop and publish research on techniques for understanding representations of deep networks.
  • Engineer infrastructure for studying model internals at scale.
  • Collaborate across teams on projects that OpenAI is uniquely suited to pursue.
  • Guide research directions toward demonstrable usefulness and/or long-term scalability.

Requirements

  • Enthusiasm for OpenAI's mission of ensuring AGI benefits all of humanity and alignment with OpenAI's charter.
  • Interest in long-term AI safety and alignment, with deep consideration of technical paths to safe AGI.
  • Experience in AI safety and alignment, mechanistic interpretability, or closely related disciplines.
  • A Ph.D. or research experience in computer science, machine learning, or a related field.
  • Ability to thrive in environments involving large-scale AI systems and make use of resources in this area.
  • 2+ years of research engineering experience.
  • Proficiency in Python or similar programming languages.
  • Deep curiosity.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. The company develops AI systems and seeks to deploy them safely.

OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. Background checks are administered in accordance with applicable law.

Benefits

  • Base salary range of $295,000–$500,000, plus equity.
  • Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
  • Pre-tax accounts for health, dependent care, and commuter expenses.
  • 401(k) retirement plan with employer match.
  • Paid parental, medical, and caregiver leave.
  • Paid time off, company holidays, office closures, and paid sick or safe time as required by law.
  • Mental health and wellness support.
  • Employer-paid basic life and disability coverage.
  • Annual learning and development stipend.
  • Daily office meals and eligible meal delivery credits.
  • Relocation support for eligible employees.
  • Additional benefits may include charitable donation matching and wellness stipends.

More jobs at OpenAI

Similar jobs