Software Engineer - Platform Infrastructure (Rust, C++)

USD 180,000-440,000 per year
MIDDLE
✅ On-site

Tech Stack

AI Debugging @ 6 Distributed Systems @ 3 Docker @ 3 GPU Grafana @ 3 Helm @ 3 Kubernetes @ 3 Linux @ 3 Networking @ 3 Observability @ 3 OpenTelemetry @ 3 Performance Analysis @ 5 Profiling @ 5 Prometheus @ 3 Rust @ 3 VictoriaMetrics @ 3

Details

The team builds AI systems and large-scale infrastructure for supercomputing clusters. Employees are expected to be hands-on, contribute directly to the mission, communicate clearly, and work effectively in a flat organizational structure.

Responsibilities

  • Design, build, and implement large-scale distributed systems powering one of the world's largest supercomputing clusters.
  • Profile, debug, and optimize performance across GPUs, the Linux kernel, networking, and filesystems.
  • Collaborate on hardware, software, and algorithm co-design for AI training.
  • Maintain and improve the codebase for scalability and reliability.
  • Develop tools to improve team productivity and streamline workflows.

Requirements

  • Systems programming experience in C, C++, or Rust.
  • Strong computer systems fundamentals, including how computers execute code from transistors through high-level applications.
  • Hands-on expertise with Kubernetes, including cluster architecture, pod lifecycle, CNI networking, CSI storage, service mesh, and production-grade operations.
  • Strong debugging skills across the full stack, from the kernel and operating system through container orchestration layers.
  • Knowledge of operating systems internals, including process scheduling, memory management, filesystems, and synchronization primitives.
  • Proficiency in performance analysis, profiling, and low-level optimization techniques.
  • Understanding of computer networks and the TCP/IP stack.
  • Experience with Linux kernel concepts or systems-level debugging tools such as perf, gdb, strace, and Wireshark.
  • Experience deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows.
  • Understanding of containerization technologies including Docker, containerd, and CRI-O, and their interaction with the Linux kernel.
  • Experience with observability and monitoring in distributed systems using Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar tools.

Benefits

  • Equity.
  • Medical, vision, and dental coverage.
  • 401(k) retirement plan.
  • Short- and long-term disability insurance.
  • Life insurance.
  • Various discounts and perks.

More jobs at SpaceXAI

Similar jobs