Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Agentic Systems @ 2
Algorithms @ 3
Data Pipelines @ 6
Distributed Systems @ 3
GPU @ 3
JAX @ 6
Kubernetes @ 3
LLM
Machine Learning @ 3
PyTorch @ 6
Python @ 6
Reinforcement Learning @ 6
Rust @ 5
Spark @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Join the multimodal team to advance superhuman multimodal intelligence across image, video, audio, and text. The role spans data curation and acquisition, tokenizer training, large-scale pre-training, post-training and alignment, infrastructure and scaling, evaluation, tooling and demos, and end-to-end product experiences.
Collaborate with pre-training, post-training, reasoning, data, applied, and product teams to deliver capabilities in multimodal reasoning, world modeling, tool use, agentic behaviors, and interactive human-AI collaboration.
Responsibilities
- Design, build, and optimize large-scale distributed systems for multimodal pre-training, post-training, inference, data processing, and tokenization at web and petabyte scale.
- Develop high-throughput pipelines for data acquisition, preprocessing, filtering, generation, decoding, loading, crawling, visualization, and management of images, videos, audio, and text.
- Advance multimodal capabilities including spatial-temporal compression, cross-modal alignment, world modeling, reasoning, emergent abilities, audio/image/video understanding and generation, real-time video processing, and noisy data handling.
- Drive data quality studies involving human and synthetic curation, filtering techniques, analysis, and scalable pipelines supporting trillion-parameter models.
- Create evaluation frameworks, internal benchmarks, reward models, and metrics that capture real-world usage, failure modes, interactive dynamics, and human-AI synergy.
- Innovate on algorithms, modeling approaches, hardware/software/algorithm co-design, and scaling paradigms for state-of-the-art performance.
- Build research tooling, user-friendly interfaces, prototypes, demos, full-stack applications, and enable rapid iteration based on feedback.
- Work across the stack from pre-training through supervised fine-tuning, reinforcement learning, and post-training to enable reasoning, tool calling, agentic behaviors, orchestration, and seamless real-time interactions.
Requirements
- Hands-on experience with multimodal pre-training, post-training, or fine-tuning involving vision, audio, video, or cross-modal systems.
- Expert-level proficiency in Python, with strong experience in at least one of JAX, PyTorch, or XLA.
- Proven track record building or optimizing large-scale distributed machine learning systems, including training or inference optimization, GPU utilization, multi-GPU/TPU setups, or hardware co-design.
- Deep experience designing and running data pipelines at scale, including curation, filtering, generation, and quality studies for noisy, real-world multimodal data.
- Strong fundamentals in evaluation design, benchmarks, reward modeling, or reinforcement learning techniques, particularly for interactive or agentic behaviors.
- Proactive self-starter who thrives in high-intensity environments and is passionate about advancing multimodal AI.
- Willingness to own end-to-end initiatives and deliver breakthrough user experiences.
Preferred Skills and Experience
- Experience leading major improvements in model capabilities through better data, modeling, algorithms, or scaling.
- Familiarity with state-of-the-art multimodal large language models, scaling laws, tokenizers, compression techniques, reasoning, or agentic systems.
- Proficiency in Rust and/or C++ for performance-critical components.
- Hands-on work with large-scale orchestration tools such as Spark, Ray, or Kubernetes.
- Background building full-stack tooling, performant interfaces, real-time research demos or applications, or end-to-end products.
- Passion for end-to-end user experience in interactive, real-time multimodal AI systems.
Benefits
- Equity.
- Comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance.
- Various discounts and perks.
- Equal opportunity employment.
More jobs at SpaceXAI
Human Data Manager
SpaceXAI · Palo Alto, United States
USD 100,000-186,000 per year
Analytics Engineer - X
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Research Analyst
SpaceXAI · United States, New York City, United States
USD 129,600-158,400 per year
AI Tutor - Yoruba
SpaceXAI · World, United States
USD 35-45 per hour
AI Tutor - Slovak
SpaceXAI · World, United States
USD 35-45 per hour
Similar jobs
Machine Learning Engineer, API Multicloud
OpenAI · San Francisco, United States
USD 295,000-445,000 per year
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · London, United Kingdom
GBP 260,000-630,000 per year
Member of Technical Staff - Imagine Model
SpaceXAI · Palo Alto, United States, Seattle, United States
USD 180,000-440,000 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Research Engineer, Discovery
Anthropic · San Francisco, United States
USD 350,000-850,000 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year