Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Debugging @ 6
Distributed Systems @ 3
Docker @ 3
GPU
Grafana @ 3
Helm @ 5
Kubernetes @ 5
Linux @ 3
Networking @ 3
Observability @ 3
OpenTelemetry @ 3
Performance Analysis @ 5
Profiling @ 5
Prometheus @ 3
Rust @ 3
VictoriaMetrics @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Responsibilities
- Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
- Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
- Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
- Maintain and innovate on our codebase to ensure scalability and reliability.
- Develop tools to enhance team productivity and streamline workflows.
Requirements
- Systems programming experience in C, C++, or Rust.
- Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications.
- Hands-on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production-grade operations.
Preferred Skills and Experience
- Collaborate in a fast-paced, open environment to design and foundational systems.
- Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers.
- Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives).
- Proficiency in performance analysis, profiling, and low-level optimization techniques.
- Solid understanding of computer networks and the TCP/IP stack.
- Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark).
- Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows.
- Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel.
- Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar).
Compensation and Benefits
- $180,000 - $440,000 USD base salary.
- Equity.
- Comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan.
- Short & long-term disability insurance.
- Life insurance.
- Various other discounts and perks.
More jobs at SpaceXAI
AI Tutor - Sinhala
SpaceXAI · World
USD 35-45 per hour
Software Engineer - Network Software and Services
SpaceXAI · Dublin, Ireland
EUR 80,000-150,000 per year
AI Tutor - Kazakh
SpaceXAI · World
USD 35-45 per hour
AI Tutor - Mongolian
SpaceXAI · World
USD 35-45 per hour
AI Tutor - Tajik
SpaceXAI · World
USD 35-45 per hour
Similar jobs
Senior Systems Engineer, Storage - DGX Cloud
Nvidia · United States
USD 208,000-414,000 per year
Senior Software Engineer, GoLang - DSX MaxQ
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, Ai Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Perplexity AI · New York City, United States, Ireland, London, United Kingdom, San Francisco, United States
USD 250,000-485,000 per year