Member of Technical Staff (AI Inference Engineer)

USD 220,000-485,000 per year
MIDDLE
✅ On-site

Tech Stack

API CUDA @ 6 Debugging Deep Learning @ 2 Distributed Systems @ 3 GPU @ 6 InfiniBand JAX @ 2 Kubernetes LLM @ 3 Machine Learning NCCL NVLink Profiling PyTorch @ 2 Python @ 3 Rust @ 3 TensorFlow @ 2

Details

We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL.

Responsibilities

  • Support transformer-based retrieval, text-generation, and multimodal models in inference infrastructure, including weight loading, request scheduling, KV-cache management, and API Gateway support.
  • Port in-house CUDA kernels to NVIDIA's CuTe DSL for current and future GPU platforms.
  • Develop the internal Rust-based inference server to address Python limitations and support growing traffic.
  • Profile and resolve bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
  • Build dashboards, alerts, and automated remediation to identify regressions before users do.
  • Respond to and learn from production incidents.

Requirements

  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Deep experience with GPU programming and performance work, using CUDA, Triton, CUTLASS, or similar technologies.
  • Understanding of modern LLM architectures and experience bringing them up reliably in production environments.
  • Experience building and operating production distributed systems under real load, ideally performance-critical systems.
  • Comfort working across Rust for serving runtimes, Python for model code, and CUDA/CuTe DSL for kernels.
  • Familiarity with at least one deep learning framework, such as PyTorch, JAX, or TensorFlow.
  • Understanding of GPU architectures, including memory hierarchy, warp scheduling, and tensor cores.
  • Understanding of common LLM architectures and inference optimization techniques, such as quantization, speculative decoding, and prefill-decode disaggregation.
  • Ability to own problems end-to-end and work independently in a fast-moving environment.

Preferred Experience

  • ML compilers and framework internals, including PyTorch internals, torch.compile, and custom operators.
  • Distributed GPU communication, including NCCL, NVLink, InfiniBand, RDMA libraries, and model or tensor parallelism.
  • Low-precision inference, including INT8, FP8, and FP4 quantization and mixed-precision serving.
  • Profiling and debugging tools, including Nsight Compute, Nsight Systems, CUDA-GDB, and PTX/SASS analysis.
  • Container orchestration, including Kubernetes, GPU scheduling, and autoscaling inference workloads.

Benefits

Full-time U.S. employees receive benefits including equity, health, dental, vision, retirement, fitness, commuter and dependent care accounts, and more. Full-time employees outside the U.S. receive benefits tailored to their region of residence.

More jobs at Perplexity AI

Similar jobs