Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
CUDA
Communication @ 6
GPU @ 6
Linux
Networking
Prioritization @ 6
Profiling @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
SpaceXAI is building one of the world’s largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, this role owns both the raw GPU supercomputer and the platform layer that runs on top of it. The work spans low-level GPU kernel optimization, Linux kernel internals, large-scale orchestration, and virtualization to make training and inference faster, more reliable, and more scalable.
Responsibilities
- Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads.
- Develop and tune low-level CUDA kernels, including GeMM and Attention, using CUTLASS, Tensor Cores, and Nsight for maximum performance.
- Profile, debug, and eliminate bottlenecks across the GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation.
- Collaborate closely with AI research teams to deliver production-grade performance and scalability.
Requirements
- Deep low-level systems programming experience with C, C++, PTX, and SASS.
- Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale.
- Hands-on experience with GPU kernel optimization, including CUTLASS, custom kernels, and Nsight profiling.
- A track record of building or running high-performance infrastructure for AI workloads, including training or inference platforms.
- Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios.
- Strong communication skills, work ethic, prioritization skills, curiosity, and a hands-on approach.
Benefits
- Base salary of $180,000–$440,000 USD.
- Equity.
- Comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance.
- Various discounts and perks.
- Equal opportunity employment.
More jobs at SpaceXAI
Human Data - Business Operations Analyst
SpaceXAI · Palo Alto, United States
USD 122,000-180,000 per year
Program Manager, Harmful Activity
SpaceXAI · New York City, United States, Palo Alto, United States, Bastrop, United States
USD 110,000-145,000 per year
Human Data Manager
SpaceXAI · Palo Alto, United States
USD 100,000-186,000 per year
Analytics Engineer - X
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Research Analyst
SpaceXAI · United States, New York City, United States
USD 129,600-158,400 per year
Similar jobs
Senior Developer Technology Engineer - Agentic SoC Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - AI Performance And Efficiency Tools
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Senior Developer Technology Engineer, High-Performance Databases
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
System Software Engineer – Data Center Compute Diagnostics
Nvidia · Durham, United States
USD 152,000-241,500 per year
Senior System Software Engineer – Data Center Compute Diagnostics
Nvidia · Durham, United States
USD 224,000-356,500 per year
Senior Developer Technology Engineer - Edge Agentic AI
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
ML Infrastructure Engineer
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year