Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 4
Algorithms @ 7
Azure @ 4
CI/CD @ 4
CUDA @ 3
Communication @ 6
Data Structures @ 7
Debugging @ 6
Deep Learning @ 7
Distributed Systems @ 7
Docker @ 4
GCP @ 4
GPU @ 7
GitHub @ 6
Go @ 1
HPC
IaC
Kubernetes @ 4
LLM @ 4
LLVM @ 4
Linux @ 3
Machine Learning
NCCL @ 3
Observability @ 4
Parallel Programming @ 7
Performance Optimization
Profiling @ 6
PyTorch @ 4
Python @ 7
Rust @ 1
SGLang @ 4
Slurm @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are seeking highly skilled and motivated software engineers to build AI inference systems that serve large-scale models with extreme efficiency. You will architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale workloads across multi-GPU, multi-node, and multi-cloud environments. You will collaborate across inference, compiler, scheduling, and performance teams to advance accelerated computing for AI.
Responsibilities
- Contribute features to vLLM that support the newest models and NVIDIA GPU hardware features.
- Profile and optimize vLLM using speculative decoding, data, tensor, expert, and pipeline parallelism, and prefill-decode disaggregation.
- Develop, optimize, and benchmark GPU kernels using fusion, autotuning, and memory and layout optimization techniques.
- Build and extend high-level DSLs and compiler infrastructure to improve kernel developer productivity while approaching peak hardware utilization.
- Define and build inference benchmarking methodologies and tools.
- Contribute new benchmarks and NVIDIA submissions to the MLPerf Inference benchmarking suite.
- Architect scheduling and orchestration for containerized, large-scale inference deployments on GPU clusters across cloud environments.
- Conduct and publish original research in ML systems, survey recent publications, and integrate research ideas and prototypes into NVIDIA software products.
Requirements
- Bachelor's degree or equivalent experience in Computer Science, Computer Engineering, or Software Engineering and 7+ years of experience; alternatively, a master's degree in one of these fields with 5+ years of experience; or a PhD with a thesis and top-tier publications in ML systems, GPU architecture, or high-performance computing.
- Strong programming skills in Python and C/C++. Experience with Go or Rust is a plus.
- Strong computer science fundamentals, including algorithms and data structures, operating systems, computer architecture, parallel programming, distributed systems, and deep learning theory.
- Knowledge of and passion for performance engineering in ML frameworks such as PyTorch and model-serving systems such as vLLM and SGLang.
- Familiarity with GPU programming and performance, including CUDA, memory hierarchy, streams, and NCCL.
- Proficiency with profiling and debugging tools such as Nsight Systems and Nsight Compute.
- Experience with containers and orchestration technologies including Docker, Kubernetes, and Slurm.
- Familiarity with Linux namespaces and cgroups.
- Excellent debugging, problem-solving, and communication skills, with the ability to work effectively in a fast-paced, multifunctional environment.
Preferred Qualifications
- Experience building and optimizing LLM inference engines such as vLLM and SGLang.
- Hands-on experience with ML compilers and DSLs such as Triton, TorchDynamo/Inductor, MLIR/LLVM, and XLA.
- Experience with GPU libraries and features such as CUTLASS, CUDA Graph, and Tensor Cores.
- Experience contributing to containerization and virtualization technologies such as containerd, CRI-O, and CRIU.
- Experience with AWS, GCP, or Azure; infrastructure as code; CI/CD; and production observability.
- Contributions to open-source projects or publications, including GitHub pull requests, published papers, and artifacts.
Company Information
NVIDIA's mission is to advance AI research and development by creating technologies that enable people to harness the power of AI. The team consists of experts in AI, systems, and performance optimization, including leaders recognized with academic and industry research awards.
Compensation and Benefits
The base salary range is 170,000 CAD–220,000 CAD for Level 4 and 225,000 CAD–275,000 CAD for Level 5. Base salary is determined by location, experience, and the pay of employees in similar positions. Employees are also eligible for equity and benefits.
Applications will be accepted at least until August 18, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.