Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Communication @ 6
Prioritization @ 6
Reinforcement Learning @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Work on critical post-training and reinforcement learning challenges, including reward modeling, preference optimization using RLHF and DPO, and reinforcement learning to improve reasoning, truthfulness, and real-world capabilities. The first project will be clarified before an offer.
The team operates with a flat organizational structure and values engineering excellence, initiative, strong prioritization, hands-on contribution, and concise, accurate communication.
Responsibilities
- Develop solutions for post-training and reinforcement learning challenges.
- Work on reward modeling, preference optimization, RLHF, DPO, and reinforcement learning.
- Improve AI model reasoning, truthfulness, alignment, and real-world capabilities.
Requirements
- Strong interest in developing truth-seeking AI systems.
- Deep enthusiasm for building useful models through post-training and reinforcement learning techniques.
- Extensive use of AI models and interest in advancing reinforcement learning and alignment methods.
- Previous experience in post-training, RLHF, or training models used by millions of people is a plus, but relevant experience is not required.
- Pride in work and ability to thrive in a meritocratic environment.
- Strong communication, prioritization, and problem-solving skills.
Benefits
- Equity.
- Comprehensive medical, vision, and dental coverage.
- 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance.
- Various discounts and perks.
More jobs at SpaceXAI
Team Lead, Human Data Operations - Vision, Image & Video
SpaceXAI · World, Dubai, United Arab Emirates, Spain, France, Australia, United Kingdom, Singapore, Singapore, Palo Alto, United States, Indonesia, Canada, Netherlands, Ireland, India, United States, Switzerland, Philippines, Japan, Germany, South Korea, Asia, Philippines
USD 104,000-156,000 per year
Expert Team Lead, SWE
SpaceXAI · World, Dubai, United Arab Emirates, Spain, France, Australia, United Kingdom, Singapore, Singapore, Palo Alto, United States, Indonesia, Canada, Netherlands, Ireland, India, United States, Switzerland, Philippines, Japan, Germany, South Korea, Asia, Philippines
USD 104,000-170,400 per year
Summer 2027 Software Engineering Internship/Co-op
SpaceXAI · Palo Alto, United States
USD 30-40 per hour
Spring 2027 Software Engineering Internship/Co-op
SpaceXAI · Palo Alto, United States
USD 30-40 per hour
Member of Technical Staff - Evaluation Infrastructure
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Similar jobs
Member of Technical Staff - Imagine Model
SpaceXAI · Palo Alto, United States, Seattle, United States
USD 180,000-440,000 per year
Member of Technical Staff - Mid-Training
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Incident Manager - Detection & Response
Anthropic · Washington, United States, New York City, United States, San Francisco, United States, Seattle, United States
USD 290,000-365,000 per year
Software Engineer, Staff: Applied AI, Science & Engineering
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Staff Research Engineer, Multi-Agent Scaling
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 500,000-850,000 per year
Applied AI Architect, International Policy
Anthropic · Washington, United States, New York City, United States, San Francisco, United States
USD 240,000-315,000 per year
Manager, Product Program Management
Anthropic · San Francisco, United States, New York City, United States
USD 365,000 per year
Incident Response Manager - Privacy
Anthropic · San Francisco, United States, New York City, United States
USD 290,000-365,000 per year