Member of Technical Staff - RL Training Framework

USD 180,000-440,000 per year
MIDDLE
✅ On-site

Tech Stack

AI Debugging @ 3 Distributed Systems @ 3 JAX @ 5 LLM @ 3 Observability Python @ 5 Reinforcement Learning @ 6 Rust @ 5

Details

The RL infrastructure team is looking for an engineer to help develop the reinforcement learning training framework. The team develops AI systems and operates with a hands-on, engineering-focused culture.

Responsibilities

  • Design and implement the systems supporting all reinforcement learning workloads, from small-scale ablations to production training runs.
  • Profile, debug, and optimize end-to-end training performance.
  • Improve the scalability and observability of the reinforcement learning stack.

Requirements

  • Experience building, debugging, and optimizing the efficiency of large-scale distributed systems.
  • Ability to work in unfamiliar areas and solve problems at all levels of the stack.
  • Proficiency in one or more of Python, JAX, Rust, and C++.
  • Experience with large-scale LLM training infrastructure is preferred.
  • Strong knowledge of reinforcement learning techniques is preferred.
  • Experience with reinforcement learning numerics is preferred.

Benefits

  • Base salary of $180,000–$440,000 USD.
  • Equity.
  • Comprehensive medical, vision, and dental coverage.
  • Access to a 401(k) retirement plan.
  • Short- and long-term disability insurance.
  • Life insurance.
  • Various other discounts and perks.

More jobs at SpaceXAI

Similar jobs