Member of Technical Staff - Post-Training and RL

USD 180,000-600,000 per year
MIDDLE SENIOR
✅ On-site

Tech Stack

AI @ 6 Communication @ 6 Prioritization @ 6 Reinforcement Learning @ 3

Details

Work on critical post-training and reinforcement learning challenges, including reward modeling, preference optimization using RLHF and DPO, and reinforcement learning to improve reasoning, truthfulness, and real-world capabilities. The first project will be clarified before an offer.

The team operates with a flat organizational structure and values engineering excellence, initiative, strong prioritization, hands-on contribution, and concise, accurate communication.

Responsibilities

  • Develop solutions for post-training and reinforcement learning challenges.
  • Work on reward modeling, preference optimization, RLHF, DPO, and reinforcement learning.
  • Improve AI model reasoning, truthfulness, alignment, and real-world capabilities.

Requirements

  • Strong interest in developing truth-seeking AI systems.
  • Deep enthusiasm for building useful models through post-training and reinforcement learning techniques.
  • Extensive use of AI models and interest in advancing reinforcement learning and alignment methods.
  • Previous experience in post-training, RLHF, or training models used by millions of people is a plus, but relevant experience is not required.
  • Pride in work and ability to thrive in a meritocratic environment.
  • Strong communication, prioritization, and problem-solving skills.

Benefits

  • Equity.
  • Comprehensive medical, vision, and dental coverage.
  • 401(k) retirement plan.
  • Short- and long-term disability insurance.
  • Life insurance.
  • Various discounts and perks.

More jobs at SpaceXAI

Similar jobs