Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Ansible @ 3
CUDA @ 3
Data Pipelines
Deep Learning @ 3
Distributed Systems @ 3
Experimentation
GPU @ 3
JAX @ 2
Linux @ 3
Machine Learning @ 3
Mentoring
Networking @ 3
Puppet @ 3
PyTorch @ 2
Python @ 6
Rust @ 6
Slurm @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About the Role
As an ML Infrastructure Engineer, you will play a pivotal role in building and optimizing the reliable, high-performance ML platform that powers SpaceXAI’s most demanding training and inference workloads.
Responsibilities
- Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses
- Developing data pipelines and integrating large-scale data, training, and inference systems
- Collaborating with ML teams to productionize models and ensure seamless integration across the stack
- Ensuring scalability, reliability, and efficiency of large-scale machine learning systems
- Working across the full stack to solve complex problems independently
- Mentoring junior engineers and contributing to the growth of the team
Requirements
- Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or other quantitative discipline; or equivalent work experience
- 2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications
- 2+ years experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists
- Strong proficiency with Python and experience with compiled languages such as C++ or Rust
Preferred Skills and Experience
- Deep familiarity with modern ML frameworks such as JAX or PyTorch
- Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking
- Comfortable with Linux systems and orchestration tools
- Experience with job schedulers (e.g., Slurm), configuration management (Puppet/Ansible), or related infrastructure tooling
Compensation and Benefits
$180,000 - $440,000 USD
Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
More jobs at SpaceXAI
Manager, Human Data Operations
SpaceXAI · Dubai, United Arab Emirates, Indonesia, Ireland, India, United States, South Korea, United Kingdom, Japan, Philippines, Singapore, Singapore
USD 108,000-162,000 per year
Software Engineer - Platform Infrastructure (Rust, C++)
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
AI Tutor - Sinhala
SpaceXAI · World
USD 35-45 per hour
Software Engineer - Network Software and Services
SpaceXAI · Dublin, Ireland
EUR 80,000-150,000 per year
AI Tutor - Kazakh
SpaceXAI · World
USD 35-45 per hour
Similar jobs
Customer Engineer
Nebius · United States
USD 179,500-224,300 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member Of Technical Staff (Ai Inference Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States, New York City, United States
USD 220,000-485,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Software Engineer, Ai Networking
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Solutions Architect
Nebius · United States, Canada
USD 250,000-320,000 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year