Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Debugging @ 6
Distributed Systems @ 3
Docker @ 3
GPU
Grafana @ 3
Helm @ 3
Kubernetes @ 3
Linux @ 3
Networking @ 3
Observability @ 3
OpenTelemetry @ 3
Performance Analysis @ 5
Profiling @ 5
Prometheus @ 3
Rust @ 3
VictoriaMetrics @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The team builds AI systems and large-scale infrastructure for supercomputing clusters. Employees are expected to be hands-on, contribute directly to the mission, communicate clearly, and work effectively in a flat organizational structure.
Responsibilities
- Design, build, and implement large-scale distributed systems powering one of the world's largest supercomputing clusters.
- Profile, debug, and optimize performance across GPUs, the Linux kernel, networking, and filesystems.
- Collaborate on hardware, software, and algorithm co-design for AI training.
- Maintain and improve the codebase for scalability and reliability.
- Develop tools to improve team productivity and streamline workflows.
Requirements
- Systems programming experience in C, C++, or Rust.
- Strong computer systems fundamentals, including how computers execute code from transistors through high-level applications.
- Hands-on expertise with Kubernetes, including cluster architecture, pod lifecycle, CNI networking, CSI storage, service mesh, and production-grade operations.
- Strong debugging skills across the full stack, from the kernel and operating system through container orchestration layers.
- Knowledge of operating systems internals, including process scheduling, memory management, filesystems, and synchronization primitives.
- Proficiency in performance analysis, profiling, and low-level optimization techniques.
- Understanding of computer networks and the TCP/IP stack.
- Experience with Linux kernel concepts or systems-level debugging tools such as perf, gdb, strace, and Wireshark.
- Experience deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows.
- Understanding of containerization technologies including Docker, containerd, and CRI-O, and their interaction with the Linux kernel.
- Experience with observability and monitoring in distributed systems using Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar tools.
Benefits
- Equity.
- Medical, vision, and dental coverage.
- 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance.
- Various discounts and perks.
More jobs at SpaceXAI
Research Analyst
SpaceXAI · United States, New York City, United States
USD 129,600-158,400 per year
AI Tutor - Yoruba
SpaceXAI · World, United States
USD 35-45 per hour
AI Tutor - Slovak
SpaceXAI · World, United States
USD 35-45 per hour
AI Tutor - Serbian
SpaceXAI · World
USD 35-45 per hour
AI Tutor - Sorani Kurdish
SpaceXAI · World
USD 35-45 per hour
Similar jobs
Senior Software Engineer, AIOps and Observability
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Software Engineer, Golang - DSX MaxQ
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Senior HPC Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Engineering Manager, Agentic GenAI Platform
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Systems Engineer, Storage - DGX Cloud
Nvidia · United States
USD 208,000-414,000 per year