Senior Software Engineer, AI Inference Systems

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ Hybrid

Tech Stack

AI AWS @ 4 Algorithms @ 4 Azure @ 4 CI/CD @ 4 CUDA @ 3 Communication @ 6 Data Structures @ 4 Debugging @ 6 Deep Learning @ 4 Distributed Systems @ 4 Docker @ 4 GCP @ 4 GPU @ 7 GitHub @ 6 Go @ 1 HPC IaC Kubernetes @ 4 LLM @ 4 LLVM @ 4 Linux @ 3 Machine Learning NCCL @ 3 Observability @ 4 Parallel Programming @ 4 Profiling @ 6 PyTorch @ 4 Python @ 7 Rust @ 1 SGLang @ 4 Slurm @ 4 vLLM @ 4

Details

We are seeking highly skilled and motivated software engineers to build AI inference systems that serve large-scale models with extreme efficiency. You will architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale workloads across multi-GPU, multi-node, and multi-cloud environments. You will collaborate across inference, compiler, scheduling, and performance teams to advance accelerated computing for AI.

Responsibilities

  • Contribute features to vLLM that support the newest models and NVIDIA GPU hardware features.
  • Profile and optimize vLLM using speculative decoding, data, tensor, expert, and pipeline parallelism, and prefill-decode disaggregation.
  • Develop, optimize, and benchmark GPU kernels using fusion, autotuning, and memory/layout optimization techniques.
  • Build and extend high-level DSLs and compiler infrastructure to improve kernel developer productivity and hardware utilization.
  • Define and build inference benchmarking methodologies and tools.
  • Contribute new benchmarks and NVIDIA submissions to the MLPerf Inference benchmarking suite.
  • Architect scheduling and orchestration for containerized, large-scale inference deployments on GPU clusters across cloud platforms.
  • Conduct and publish original research in ML Systems and integrate research ideas and prototypes into NVIDIA software products.

Requirements

  • Bachelor's degree or equivalent experience in Computer Science, Computer Engineering, or Software Engineering with 7+ years of experience; alternatively, a master's degree with 5+ years of experience; or a PhD with a thesis and top-tier publications in ML Systems, GPU architecture, or high-performance computing.
  • Strong programming skills in Python and C/C++. Experience with Go or Rust is a plus.
  • Solid knowledge of algorithms and data structures, operating systems, computer architecture, parallel programming, distributed systems, and deep learning theory.
  • Experience with performance engineering in ML frameworks such as PyTorch and inference engines such as vLLM and SGLang.
  • Familiarity with GPU programming and performance, including CUDA, memory hierarchy, streams, and NCCL.
  • Proficiency with profiling and debugging tools such as Nsight Systems and Nsight Compute.
  • Experience with containers and orchestration technologies including Docker, Kubernetes, and Slurm.
  • Familiarity with Linux namespaces and cgroups.
  • Excellent debugging, problem-solving, and communication skills, with the ability to work effectively in a fast-paced, multifunctional environment.

Preferred Qualifications

  • Experience building and optimizing LLM inference engines such as vLLM and SGLang.
  • Hands-on experience with ML compilers and DSLs such as Triton, TorchDynamo/Inductor, MLIR/LLVM, and XLA.
  • Experience with GPU libraries and features such as CUTLASS, CUDA Graph, and Tensor Cores.
  • Experience with containerization and virtualization technologies such as containerd, CRI-O, and CRIU.
  • Experience with AWS, GCP, or Azure; infrastructure as code; CI/CD; and production observability.
  • Contributions to open-source projects or publications, including GitHub pull requests, published papers, and artifacts.

Benefits

The role includes eligibility for equity and benefits. NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.

More jobs at Nvidia

Similar jobs