Member of Technical Staff - Imagine Model

USD 180,000-440,000 per year
MIDDLE
✅ On-site

Tech Stack

AI @ 3 Agentic Systems @ 3 Algorithms Communication @ 6 Debugging @ 3 JAX @ 5 Machine Learning @ 3 Prioritization @ 6 PyTorch @ 5 Python @ 5 Reinforcement Learning @ 3 Rust @ 5 Spark @ 5

Details

SpaceXAI is seeking a multimodal engineer to develop AI experiences beyond text, with a focus on high-fidelity understanding and generation across image and video modalities. The role also incorporates audio where it enhances visual content, such as synchronized audio for video. Responsibilities span data curation, modeling, training, inference serving, and product integration across pretraining and post-training phases.

Responsibilities

  • Create and drive engineering agendas to advance multimodal capabilities, including image and video generation, editing, understanding, controllable and long-horizon synthesis, agentic planning, reinforcement learning training, and world simulation.
  • Integrate audio into visual experiences to create richer video content.
  • Improve data quality through annotation, filtering, augmentation, synthetic generation, captioning, and in-depth data studies, particularly for visual and audio data.
  • Design evaluation frameworks, metrics, benchmarks, evaluations, and reward models for image, video, and audio quality and coherence.
  • Implement efficient algorithms for state-of-the-art model performance, including real-time inference, distillation, and scalable serving for visual content.
  • Develop scalable data collection and processing pipelines for multimodal datasets, primarily focused on image and video.
  • Collaborate cross-functionally to integrate AI solutions into production and iterate based on user feedback.

Requirements

  • Track record of leading studies that significantly improve neural network capabilities and performance through improved data or modeling.
  • Experience with data-driven experiment design, systematic analysis, and iterative model debugging.
  • Experience developing or working with large-scale distributed machine learning systems.
  • Ability to deliver optimal end-to-end user experiences.
  • Hands-on contribution, initiative, excellence, strong work ethic, prioritization skills, and excellent communication.

Preferred Skills and Experience

  • Experience with supervised fine-tuning (SFT), reinforcement learning, evaluations, human or synthetic data collection, or agentic systems.
  • Proficiency in Python, JAX/XLA, PyTorch, Rust/C++, Spark, Ray, and related large-scale frameworks.
  • Domain expertise in multimodal applications such as graphics engines, rendering techniques, image and video understanding and generation, world models, real-time simulation, or controllable and long-horizon visual content creation.
  • Audio or speech processing and music or audio generation experience is a plus where it supports video.
  • Experience with agentic reinforcement learning training, controllable or long-horizon generation, or multimodal agents that reason and act across modalities, especially in visual domains.

Compensation and Benefits

  • Base salary: $180,000–$440,000 USD per year.
  • Equity.
  • Comprehensive medical, vision, and dental coverage.
  • 401(k) retirement plan.
  • Short- and long-term disability insurance.
  • Life insurance.
  • Various discounts and perks.

SpaceXAI is an equal opportunity employer.

More jobs at SpaceXAI

Similar jobs