Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 4
Algorithms @ 4
Azure @ 4
CI/CD @ 4
CUDA @ 3
Communication @ 6
Data Structures @ 4
Debugging @ 6
Deep Learning @ 4
Distributed Systems @ 4
Docker @ 4
GCP @ 4
GPU @ 7
GitHub @ 6
Go @ 1
HPC
IaC
Kubernetes @ 4
LLM @ 4
LLVM @ 4
Linux @ 3
Machine Learning
NCCL @ 3
Observability @ 4
Parallel Programming @ 4
Profiling @ 6
PyTorch @ 4
Python @ 7
Rust @ 1
SGLang @ 4
Slurm @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are seeking highly skilled and motivated software engineers to build AI inference systems that serve large-scale models with extreme efficiency. You will architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale workloads across multi-GPU, multi-node, and multi-cloud environments. You will collaborate across inference, compiler, scheduling, and performance teams to advance accelerated computing for AI.
Responsibilities
- Contribute features to vLLM that support the newest models and NVIDIA GPU hardware features.
- Profile and optimize vLLM using speculative decoding, data, tensor, expert, and pipeline parallelism, and prefill-decode disaggregation.
- Develop, optimize, and benchmark GPU kernels using fusion, autotuning, and memory/layout optimization techniques.
- Build and extend high-level DSLs and compiler infrastructure to improve kernel developer productivity and hardware utilization.
- Define and build inference benchmarking methodologies and tools.
- Contribute new benchmarks and NVIDIA submissions to the MLPerf Inference benchmarking suite.
- Architect scheduling and orchestration for containerized, large-scale inference deployments on GPU clusters across cloud platforms.
- Conduct and publish original research in ML Systems and integrate research ideas and prototypes into NVIDIA software products.
Requirements
- Bachelor's degree or equivalent experience in Computer Science, Computer Engineering, or Software Engineering with 7+ years of experience; alternatively, a master's degree with 5+ years of experience; or a PhD with a thesis and top-tier publications in ML Systems, GPU architecture, or high-performance computing.
- Strong programming skills in Python and C/C++. Experience with Go or Rust is a plus.
- Solid knowledge of algorithms and data structures, operating systems, computer architecture, parallel programming, distributed systems, and deep learning theory.
- Experience with performance engineering in ML frameworks such as PyTorch and inference engines such as vLLM and SGLang.
- Familiarity with GPU programming and performance, including CUDA, memory hierarchy, streams, and NCCL.
- Proficiency with profiling and debugging tools such as Nsight Systems and Nsight Compute.
- Experience with containers and orchestration technologies including Docker, Kubernetes, and Slurm.
- Familiarity with Linux namespaces and cgroups.
- Excellent debugging, problem-solving, and communication skills, with the ability to work effectively in a fast-paced, multifunctional environment.
Preferred Qualifications
- Experience building and optimizing LLM inference engines such as vLLM and SGLang.
- Hands-on experience with ML compilers and DSLs such as Triton, TorchDynamo/Inductor, MLIR/LLVM, and XLA.
- Experience with GPU libraries and features such as CUTLASS, CUDA Graph, and Tensor Cores.
- Experience with containerization and virtualization technologies such as containerd, CRI-O, and CRIU.
- Experience with AWS, GCP, or Azure; infrastructure as code; CI/CD; and production observability.
- Contributions to open-source projects or publications, including GitHub pull requests, published papers, and artifacts.
Benefits
The role includes eligibility for equity and benefits. NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior AI Performance and Efficiency Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Perplexity AI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 250,000-485,000 per year
Senior Software SDET Test Development Engineer
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year