Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Debugging @ 3
Distributed Systems @ 3
JAX @ 5
LLM @ 3
Observability
Python @ 5
Reinforcement Learning @ 6
Rust @ 5
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The RL infrastructure team is looking for an engineer to help develop the reinforcement learning training framework. The team develops AI systems and operates with a hands-on, engineering-focused culture.
Responsibilities
- Design and implement the systems supporting all reinforcement learning workloads, from small-scale ablations to production training runs.
- Profile, debug, and optimize end-to-end training performance.
- Improve the scalability and observability of the reinforcement learning stack.
Requirements
- Experience building, debugging, and optimizing the efficiency of large-scale distributed systems.
- Ability to work in unfamiliar areas and solve problems at all levels of the stack.
- Proficiency in one or more of Python, JAX, Rust, and C++.
- Experience with large-scale LLM training infrastructure is preferred.
- Strong knowledge of reinforcement learning techniques is preferred.
- Experience with reinforcement learning numerics is preferred.
Benefits
- Base salary of $180,000–$440,000 USD.
- Equity.
- Comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance.
- Various other discounts and perks.
More jobs at SpaceXAI
Human Data - Business Operations Analyst
SpaceXAI · Palo Alto, United States
USD 122,000-180,000 per year
Program Manager, Harmful Activity
SpaceXAI · New York City, United States, Palo Alto, United States, Bastrop, United States
USD 110,000-145,000 per year
Human Data Manager
SpaceXAI · Palo Alto, United States
USD 100,000-186,000 per year
Analytics Engineer - X
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Research Analyst
SpaceXAI · United States, New York City, United States
USD 129,600-158,400 per year
Similar jobs
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · London, United Kingdom
GBP 260,000-630,000 per year
Software Engineer, Model Runtime
OpenAI · San Francisco, United States
USD 266,000-445,000 per year
Senior AI Infrastructure Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
AI Systems Engineer, Codex Agents
OpenAI · San Francisco, United States
USD 230,000-385,000 per year
Member of Technical Staff - Multimodal Understanding
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year