ML Infrastructure Engineer

USD 180,000-440,000 per year
MIDDLE
✅ On-site

Tech Stack

Ansible @ 3 CUDA @ 3 Data Pipelines Deep Learning @ 3 Distributed Systems @ 3 Experimentation GPU @ 3 JAX @ 2 Linux @ 3 Machine Learning @ 3 Mentoring Networking @ 3 Puppet @ 3 PyTorch @ 2 Python @ 6 Rust @ 6 Slurm @ 3

Details

About the Role

As an ML Infrastructure Engineer, you will play a pivotal role in building and optimizing the reliable, high-performance ML platform that powers SpaceXAI’s most demanding training and inference workloads.

Responsibilities

  • Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses
  • Developing data pipelines and integrating large-scale data, training, and inference systems
  • Collaborating with ML teams to productionize models and ensure seamless integration across the stack
  • Ensuring scalability, reliability, and efficiency of large-scale machine learning systems
  • Working across the full stack to solve complex problems independently
  • Mentoring junior engineers and contributing to the growth of the team

Requirements

  • Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or other quantitative discipline; or equivalent work experience
  • 2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications
  • 2+ years experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists
  • Strong proficiency with Python and experience with compiled languages such as C++ or Rust

Preferred Skills and Experience

  • Deep familiarity with modern ML frameworks such as JAX or PyTorch
  • Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking
  • Comfortable with Linux systems and orchestration tools
  • Experience with job schedulers (e.g., Slurm), configuration management (Puppet/Ansible), or related infrastructure tooling

Compensation and Benefits

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

More jobs at SpaceXAI

Similar jobs