Software Engineer - Platform Infrastructure (Rust, C++)

USD 180,000-440,000 per year
MIDDLE
✅ On-site

Tech Stack

AI Debugging @ 6 Distributed Systems @ 3 Docker @ 3 GPU Grafana @ 3 Helm @ 5 Kubernetes @ 5 Linux @ 3 Networking @ 3 Observability @ 3 OpenTelemetry @ 3 Performance Analysis @ 5 Profiling @ 5 Prometheus @ 3 Rust @ 3 VictoriaMetrics @ 3

Details

Responsibilities

  • Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
  • Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
  • Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
  • Maintain and innovate on our codebase to ensure scalability and reliability.
  • Develop tools to enhance team productivity and streamline workflows.

Requirements

  • Systems programming experience in C, C++, or Rust.
  • Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications.
  • Hands-on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production-grade operations.

Preferred Skills and Experience

  • Collaborate in a fast-paced, open environment to design and foundational systems.
  • Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers.
  • Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives).
  • Proficiency in performance analysis, profiling, and low-level optimization techniques.
  • Solid understanding of computer networks and the TCP/IP stack.
  • Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark).
  • Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows.
  • Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel.
  • Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar).

Compensation and Benefits

  • $180,000 - $440,000 USD base salary.
  • Equity.
  • Comprehensive medical, vision, and dental coverage.
  • Access to a 401(k) retirement plan.
  • Short & long-term disability insurance.
  • Life insurance.
  • Various other discounts and perks.

More jobs at SpaceXAI

Similar jobs