Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

USD 250,000-485,000 per year
MIDDLE SENIOR
✅ Remote ✅ On-site

Tech Stack

AI AWS @ 3 CUDA @ 3 Distributed Systems @ 6 GCP @ 3 GPU @ 3 Go @ 3 Grafana HPC InfiniBand @ 3 Kubernetes @ 6 LLM Machine Learning Networking @ 3 Observability Prometheus Rust @ 3 SGLang Slurm TensorRT vLLM

Details

Perplexity serves hundreds of millions of queries a month, and every one of them fans out into multiple AI inference requests running in real time. Behind that sits a large GPU fleet spread across several cloud providers. Today, inference engineers and researchers build models while also managing networking, securing capacity, and operating the underlying GPU clusters. This role will take ownership of that infrastructure and hide its complexity behind a unified, self-serve platform for running training and inference workloads.

Responsibilities

  • Build a self-serve compute platform: Design and own systems that let inference engineers and researchers launch training jobs and operate inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure.
  • Operate the GPU fleet: Own provisioning, lifecycle management, reliability, and capacity integration across providers, giving teams a consistent way to use compute regardless of where it runs.
  • Solve for GPU scarcity: Build scheduling and placement logic that finds available capacity across providers, packs it efficiently, and gets the right workload onto the right hardware under real constraints.
  • Support two different workloads: Keep long-running distributed training jobs healthy while guaranteeing the availability and latency of production inference services on the same fleet.
  • Own Kubernetes for GPU orchestration: Write operators and CRDs, and manage many clusters across providers so the platform behaves consistently everywhere it runs.
  • Make failure boring: Build fault tolerance, autoscaling, and observability that keep the fleet utilized and let workloads survive node loss, provider disruptions, and capacity shifts without human intervention.
  • Set technical direction across teams: Partner with inference and cloud infrastructure engineers to turn operational constraints into a coherent platform architecture and roadmap.

Requirements

  • Deep Kubernetes experience, including custom operators, CRDs, and multi-cluster federation.
  • Experience managing GPU clusters at scale, including NVIDIA hardware, CUDA, and networking such as InfiniBand or RoCE.
  • Experience orchestrating compute across multiple clouds, such as CoreWeave, AWS, GCP, or similar, with an understanding of provider differences.
  • Strong distributed systems fundamentals, including scheduling, resource allocation, and fault tolerance under load.
  • Ability to write infrastructure and systems-level code in Go, Rust, or C++.
  • Experience supporting both long-running training jobs and high-availability inference services.
  • Ability to own problems end-to-end and work effectively when the path forward is not fully defined.

Additional Experience

  • Inference serving stacks such as vLLM, SGLang, or TensorRT-LLM.
  • Slurm or other HPC schedulers.
  • GPU kernel development in CUDA or Triton.
  • Production experience with high-speed interconnects such as InfiniBand, RoCE, or RDMA.
  • Observability for ML workloads using Prometheus, Grafana, or Weights & Biases.

Benefits

Full-time U.S. employees receive a comprehensive benefits program including equity, health, dental, vision, retirement, fitness, commuter and dependent care accounts, and more. Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region of residence. USD salary ranges apply only to U.S.-based positions; international salaries are set based on the local market. Final offer amounts are determined by multiple factors, including experience and expertise, and may vary from the listed amounts.

More jobs at Perplexity AI

Similar jobs