Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Agentic Systems @ 3
Algorithms
Communication @ 6
Debugging @ 3
JAX @ 5
Machine Learning @ 3
Prioritization @ 6
PyTorch @ 5
Python @ 5
Reinforcement Learning @ 3
Rust @ 5
Spark @ 5
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
SpaceXAI is seeking a multimodal engineer to develop AI experiences beyond text, with a focus on high-fidelity understanding and generation across image and video modalities. The role also incorporates audio where it enhances visual content, such as synchronized audio for video. Responsibilities span data curation, modeling, training, inference serving, and product integration across pretraining and post-training phases.
Responsibilities
- Create and drive engineering agendas to advance multimodal capabilities, including image and video generation, editing, understanding, controllable and long-horizon synthesis, agentic planning, reinforcement learning training, and world simulation.
- Integrate audio into visual experiences to create richer video content.
- Improve data quality through annotation, filtering, augmentation, synthetic generation, captioning, and in-depth data studies, particularly for visual and audio data.
- Design evaluation frameworks, metrics, benchmarks, evaluations, and reward models for image, video, and audio quality and coherence.
- Implement efficient algorithms for state-of-the-art model performance, including real-time inference, distillation, and scalable serving for visual content.
- Develop scalable data collection and processing pipelines for multimodal datasets, primarily focused on image and video.
- Collaborate cross-functionally to integrate AI solutions into production and iterate based on user feedback.
Requirements
- Track record of leading studies that significantly improve neural network capabilities and performance through improved data or modeling.
- Experience with data-driven experiment design, systematic analysis, and iterative model debugging.
- Experience developing or working with large-scale distributed machine learning systems.
- Ability to deliver optimal end-to-end user experiences.
- Hands-on contribution, initiative, excellence, strong work ethic, prioritization skills, and excellent communication.
Preferred Skills and Experience
- Experience with supervised fine-tuning (SFT), reinforcement learning, evaluations, human or synthetic data collection, or agentic systems.
- Proficiency in Python, JAX/XLA, PyTorch, Rust/C++, Spark, Ray, and related large-scale frameworks.
- Domain expertise in multimodal applications such as graphics engines, rendering techniques, image and video understanding and generation, world models, real-time simulation, or controllable and long-horizon visual content creation.
- Audio or speech processing and music or audio generation experience is a plus where it supports video.
- Experience with agentic reinforcement learning training, controllable or long-horizon generation, or multimodal agents that reason and act across modalities, especially in visual domains.
Compensation and Benefits
- Base salary: $180,000–$440,000 USD per year.
- Equity.
- Comprehensive medical, vision, and dental coverage.
- 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance.
- Various discounts and perks.
SpaceXAI is an equal opportunity employer.
More jobs at SpaceXAI
Human Data Manager
SpaceXAI · Palo Alto, United States
USD 100,000-186,000 per year
Analytics Engineer - X
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Research Analyst
SpaceXAI · United States, New York City, United States
USD 129,600-158,400 per year
AI Tutor - Yoruba
SpaceXAI · World, United States
USD 35-45 per hour
AI Tutor - Slovak
SpaceXAI · World, United States
USD 35-45 per hour
Similar jobs
Member of Technical Staff - Multimodal Understanding
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Research Scientist, Robotics Research - PhD New College Grad 2026
Nvidia · Seattle, United States
USD 168,000-264,500 per year
Senior Robotics Research Scientist
Nvidia · Seattle, United States
USD 192,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · London, United Kingdom
GBP 260,000-630,000 per year
Senior Deep Learning Algorithm Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Applied Research Scientist – AI Native Numerical Methods
Nvidia · United States
USD 192,000-356,500 per year