Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Ansible @ 3
CUDA @ 3
Communication @ 6
Data Pipelines
Deep Learning @ 3
Distributed Systems @ 3
Experimentation
GPU @ 3
JAX @ 2
Linux @ 3
Machine Learning @ 3
Networking @ 3
Prioritization @ 6
Puppet @ 3
PyTorch @ 2
Python @ 6
Rust @ 6
Slurm @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
SpaceXAI is seeking an ML Infrastructure Engineer to build and optimize the reliable, high-performance machine learning platform that powers recommendations on X. The role involves working across the stack in a highly collaborative, hands-on engineering environment.
Responsibilities
- Design, build, and scale GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on machine learning hypotheses.
- Develop data pipelines and integrate large-scale data, training, and inference systems.
- Collaborate with machine learning teams to productionize models and ensure seamless integration across the stack.
- Ensure the scalability, reliability, and efficiency of large-scale machine learning systems.
- Work across the full stack to solve complex problems independently.
- Mentor junior engineers and contribute to the growth of the team.
Requirements
- Bachelor's, Master's, postgraduate, or PhD degree in computer science, machine learning, or another quantitative discipline, or equivalent work experience.
- At least 2 years of industry experience working with high-traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications.
- At least 2 years of experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists.
- Strong proficiency with Python and experience with compiled languages such as C++ or Rust.
- Deep familiarity with modern ML frameworks such as JAX or PyTorch.
- Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking.
- Comfort working with Linux systems and orchestration tools.
- Experience with job schedulers such as Slurm, configuration management tools such as Puppet or Ansible, or related infrastructure tooling.
- Strong communication skills, work ethic, prioritization skills, curiosity, and a willingness to contribute directly to the company's mission.
Benefits
- Equity.
- Comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance.
- Various discounts and perks.
SpaceXAI is an equal opportunity employer.
More jobs at SpaceXAI
Human Data Manager
SpaceXAI · Palo Alto, United States
USD 100,000-186,000 per year
Analytics Engineer - X
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Research Analyst
SpaceXAI · United States, New York City, United States
USD 129,600-158,400 per year
AI Tutor - Yoruba
SpaceXAI · World, United States
USD 35-45 per hour
AI Tutor - Slovak
SpaceXAI · World, United States
USD 35-45 per hour
Similar jobs
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Key Customers Solutions Architect
Nebius · United States, Canada
USD 208,800-261,000 per year
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Nvidia · Warsaw, Poland
PLN 221,200-507,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year