Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
CUDA
Communication @ 6
GPU @ 6
Linux
Networking
Prioritization @ 6
Profiling @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
SpaceXAI is building one of the world’s largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, this role owns both the raw GPU supercomputer and the platform layer that runs on top of it. The work spans low-level GPU kernel optimization, Linux kernel internals, large-scale orchestration, and virtualization to make training and inference faster, more reliable, and more scalable.
Responsibilities
- Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads.
- Develop and tune low-level CUDA kernels, including GeMM and Attention, using CUTLASS, Tensor Cores, and Nsight for maximum performance.
- Profile, debug, and eliminate bottlenecks across the GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation.
- Collaborate closely with AI research teams to deliver production-grade performance and scalability.
Requirements
- Deep low-level systems programming experience with C, C++, PTX, and SASS.
- Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale.
- Hands-on experience with GPU kernel optimization, including CUTLASS, custom kernels, and Nsight profiling.
- A track record of building or running high-performance infrastructure for AI workloads, including training or inference platforms.
- Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios.
- Strong communication skills, work ethic, prioritization skills, curiosity, and a hands-on approach.
Benefits
- Base salary of $180,000–$440,000 USD.
- Equity.
- Comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance.
- Various discounts and perks.
- Equal opportunity employment.
More jobs at SpaceXAI
Team Lead, Human Data Operations - Post-Training
SpaceXAI · United Kingdom, Indonesia, Ireland, India, Japan, South Korea, Philippines, United States, Dubai, United Arab Emirates, Singapore, Singapore
USD 104,000-156,000 per year
Expert Team Lead, Human Data Operations
SpaceXAI · United Kingdom, Indonesia, Ireland, India, Japan, South Korea, Philippines, United States, Dubai, United Arab Emirates, Singapore, Singapore
USD 104,000-170,400 per year
AI Tutor - Design Specialist
SpaceXAI · World
USD 35-75 per hour
Supervisor, IT Support
SpaceXAI · Memphis, United States
USD 91,800-137,700 per year
AV Administrator, Events & Experiences
SpaceXAI · New York City, United States
USD 115,000-135,000 per year
Similar jobs
Senior Developer Technology Engineer - Agentic SoC Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Solutions Architecture Intern - Summer 2027
Nvidia · Santa Clara, United States
USD 20-71 per hour
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Senior Cloud Software Engineer, Developer Tools
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Developer Technology Engineer, High-Performance Databases
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
System Software Engineer – Data Center Compute Diagnostics
Nvidia · Durham, United States
USD 152,000-241,500 per year