Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
📍 New York City, United States
📍 San Francisco, United States
📍 Seattle, United States
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 3
CUDA @ 3
Distributed Systems @ 6
GCP @ 3
GPU @ 3
Go @ 3
Grafana
HPC
InfiniBand @ 3
Kubernetes @ 6
LLM
Machine Learning
Networking @ 3
Observability
Prometheus
Rust @ 3
SGLang
Slurm
TensorRT
vLLM
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Perplexity serves hundreds of millions of queries a month, and every one of them fans out into multiple AI inference requests running in real time. Behind that sits a large GPU fleet spread across several cloud providers. Today, inference engineers and researchers build models while also managing networking, securing capacity, and operating the underlying GPU clusters. This role will take ownership of that infrastructure and hide its complexity behind a unified, self-serve platform for running training and inference workloads.
Responsibilities
- Build a self-serve compute platform: Design and own systems that let inference engineers and researchers launch training jobs and operate inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure.
- Operate the GPU fleet: Own provisioning, lifecycle management, reliability, and capacity integration across providers, giving teams a consistent way to use compute regardless of where it runs.
- Solve for GPU scarcity: Build scheduling and placement logic that finds available capacity across providers, packs it efficiently, and gets the right workload onto the right hardware under real constraints.
- Support two different workloads: Keep long-running distributed training jobs healthy while guaranteeing the availability and latency of production inference services on the same fleet.
- Own Kubernetes for GPU orchestration: Write operators and CRDs, and manage many clusters across providers so the platform behaves consistently everywhere it runs.
- Make failure boring: Build fault tolerance, autoscaling, and observability that keep the fleet utilized and let workloads survive node loss, provider disruptions, and capacity shifts without human intervention.
- Set technical direction across teams: Partner with inference and cloud infrastructure engineers to turn operational constraints into a coherent platform architecture and roadmap.
Requirements
- Deep Kubernetes experience, including custom operators, CRDs, and multi-cluster federation.
- Experience managing GPU clusters at scale, including NVIDIA hardware, CUDA, and networking such as InfiniBand or RoCE.
- Experience orchestrating compute across multiple clouds, such as CoreWeave, AWS, GCP, or similar, with an understanding of provider differences.
- Strong distributed systems fundamentals, including scheduling, resource allocation, and fault tolerance under load.
- Ability to write infrastructure and systems-level code in Go, Rust, or C++.
- Experience supporting both long-running training jobs and high-availability inference services.
- Ability to own problems end-to-end and work effectively when the path forward is not fully defined.
Additional Experience
- Inference serving stacks such as vLLM, SGLang, or TensorRT-LLM.
- Slurm or other HPC schedulers.
- GPU kernel development in CUDA or Triton.
- Production experience with high-speed interconnects such as InfiniBand, RoCE, or RDMA.
- Observability for ML workloads using Prometheus, Grafana, or Weights & Biases.
Benefits
Full-time U.S. employees receive a comprehensive benefits program including equity, health, dental, vision, retirement, fitness, commuter and dependent care accounts, and more. Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region of residence. USD salary ranges apply only to U.S.-based positions; international salaries are set based on the local market. Final offer amounts are determined by multiple factors, including experience and expertise, and may vary from the listed amounts.