Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Communication @ 6
Debugging @ 5
Linux @ 6
Networking @ 3
Observability
Prioritization @ 6
Security @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The team builds AI systems and operates with a focus on engineering excellence, hands-on contribution, initiative, communication, and strong prioritization.
Responsibilities
- Build and maintain a lean, high-reliability, Linux-based operating system supporting a supercomputer network fabric.
- Write or rewrite high-performance device drivers to maximize the capabilities of network and compute hardware.
- Manage the deployment, operation, and production debugging of the platform.
- Resolve and root-cause anomalies to continuously improve system reliability.
- Develop and improve observability and configuration-management tools.
- Collaborate with hardware teams and external partners to design next-generation supercomputer hardware and bring up the operating system and software stack on it.
Requirements
- Hands-on systems programming experience in C or C++.
- Strong operating-systems fundamentals, including scheduling, memory management, and I/O.
- Understanding of computer networking, including host operating-system stacks and network devices such as switches and routers.
- Excellent written and verbal communication skills.
Preferred Skills and Experience
- Deep knowledge of the Linux kernel and its networking stack; experience with another production-grade kernel is also transferable.
- Proficiency with debugging tools across kernel, network, and userspace, including ftrace, perf, Wireshark/tcpdump, eBPF, and gdb.
- Experience bringing up new hardware, including flashing bare-metal software on a microcontroller with JTAG or bootstrapping a Linux kernel and root filesystem on new hardware.
- A track record of developing and optimizing device and peripheral drivers, ideally for high-speed interfaces or systems involving DMA and cache coherency.
- Understanding of peripheral hardware interface standards such as Ethernet, PCIe, I2C, and SPI, including the ability to review schematics and debug these interfaces.
- Experience with product security, operating-system and system hardening, and secure-boot schemes.
Compensation and Benefits
- Base salary: $180,000–$440,000 USD per year.
- Equity.
- Comprehensive medical, vision, and dental coverage.
- 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance.
- Various discounts and perks.
More jobs at SpaceXAI
Exceptional Designer
SpaceXAI · Spain
USD 180,000-400,000 per year
Sr. Security Engineer - GRC Fintech & Financial Services
SpaceXAI · Washington, United States, New York City, United States, Palo Alto, United States
USD 152,000-258,000 per year
Software Engineer - Evals
SpaceXAI · Palo Alto, United States
USD 175,000-275,000 per year
ML Infrastructure Engineer
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Software Engineer - Platform Infrastructure (Rust, C++)
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer - Connectivity Gateway
Bloomberg · New York City, United States
USD 160,000-240,000 per year
Distinguished Engineer, Storage – AI Cloud
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
Senior Storage Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year