Member Of Technical Staff (Ai Inference Engineer)
USD 220,000-485,000 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API
CUDA @ 7
Debugging
Deep Learning @ 3
Distributed Systems @ 6
GPU @ 7
InfiniBand
JAX @ 3
Kubernetes
LLM @ 4
Machine Learning
NCCL
NVLink
Observability
Profiling
PyTorch @ 3
Python @ 6
Rust @ 6
TensorFlow @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL—and we need another engineer to join us.
Responsibilities
Examples of real work the team does:
- New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
- GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
- Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
- Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
- Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.
Requirements
- Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.
- You understand modern LLM architectures and are able to bring them up reliably in a production environment.
- You've built and operated production distributed systems under real load—ideally performance-critical ones.
- Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.
- You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.
- Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you.
Good if you touched any of
- ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.
- Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.
- Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.
- Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.
- Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads.
Qualifications
- 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
- Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
- Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
- Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).
More jobs at Perplexity AI
Engineering Manager (TLM, Agents)
Perplexity AI · San Francisco, United States
USD 300,000-405,000 per year
Member Of Technical Staff (Secure Intelligence Institute)
Perplexity AI · San Francisco, United States
USD 220,000-405,000 per year
Member of Technical Staff (AI Software Engineer, Agents)
Perplexity AI · San Francisco, United States
USD 220,000-405,000 per year
Member of Technical Staff (Engineering Lead, Developer Experience & Relations)
Perplexity AI · San Francisco, United States, New York City, United States
USD 220,000-405,000 per year
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Perplexity AI · New York City, United States, Ireland, London, United Kingdom, San Francisco, United States
USD 250,000-485,000 per year
Similar jobs
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, Ai Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Training Performance Engineer
OpenAI · San Francisco, United States
USD 250,000-445,000 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Systems Software Engineer, AI Stack And Performance - DGX Station
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year