Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
CUDA @ 6
Communication @ 4
Deep Learning @ 6
GPU
HPC
JAX @ 4
LLM @ 4
MPI @ 4
Machine Learning @ 4
NCCL @ 4
Performance Analysis
PyTorch @ 4
Python @ 4
Reinforcement Learning @ 6
SGLang @ 4
System Architecture @ 4
TensorRT @ 6
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is developing technologies in artificial intelligence, high-performance computing, and visualization. The GPU serves as the foundation of many of its products and services.
This role focuses on bringing advanced CUDA features and distributed runtime technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, and JAX. The work covers multi-GPU and multi-node workloads, ranging from training at scales of up to 100,000 GPUs to inference with microsecond-level latency. You will collaborate with teams working on CUDA features and runtimes for deep learning and high-performance computing applications.
Responsibilities
- Integrate new CUDA features and runtime abstractions into AI frameworks, from proof of concept through performance analysis and production.
- Analyze AI workloads and frameworks to identify lower-level requirements and opportunities for innovation.
- Collaborate with teams working on the latest AI models.
- Drive improvements in the AI compiler-runtime interface to build high-performance multi-GPU and multi-node solutions.
- Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads.
- Influence the core CUDA roadmap to support next-generation deep learning frameworks.
- Collaborate across teams and time zones.
- Work with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design performant and programmable systems and frameworks.
- Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
- Write clean, effective, and maintainable code, transitioning prototypes into open-source releases, upstream framework integrations, internal tools, or commercial products.
Requirements
- Bachelor's, master's, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
- At least 8 years of relevant industry experience or equivalent academic experience after completing the degree.
- Development experience with deep learning frameworks such as PyTorch and JAX, and inference engines such as TRT-LLM, vLLM, and SGLang.
- Rapid prototyping and development experience with Python, C++, CUDA, or related domain-specific languages.
- Strong understanding of AI models, parallelism, and compiler technologies such as
torch.compile. - Experience benchmarking performance on AI clusters.
- Familiarity with performance profiler toolchains such as PyTorch Profiler or NVIDIA Nsight Systems.
- Understanding of high-performance computing and AI communication concepts.
- Good understanding of computer system architecture, hardware-software interactions, operating system principles, and systems software fundamentals.
- Adaptability and willingness to learn new frameworks and tools.
- Ability to work and communicate effectively across teams and time zones.
Preferred Qualifications
- Deep expertise in the performance internals and execution graphs of major deep learning, autograd, training, and inference frameworks, including PyTorch, JAX, TensorRT, vLLM, SGLang, NeMo, Megatron, and MaxText.
- Hands-on experience with CUDA, communication libraries such as NCCL, MPI, or UCX, and distributed machine learning techniques such as pipeline parallelism and tensor parallelism.
- Expertise in training, distributed inference, mixture-of-experts, reinforcement learning, or kernel authoring with CUDA, Triton, or cuTe.
- Background in deep learning compilers, including graph-level and code-generation technologies such as Triton, XLA, and
torch.compile. - Experience programming compute and communication overlap in distributed runtimes.
Compensation and Benefits
- Base salary for Level 4: USD 184,000–287,500 per year.
- Base salary for Level 5: USD 224,000–356,500 per year.
- Additional equity and benefits are provided.
- Applications will be accepted at least until July 1, 2026.
- This posting is for an existing vacancy.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Software Solutions Engineer
Nvidia · Poland
PLN 230,200-487,500 per year
Senior Offensive Security Engineer, Automotive
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Systems Operations and Administrator
Nvidia · Santa Clara, United States
USD 112,000-218,500 per year
Senior Software Engineer, Agent Simulation and Evaluation
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior AI Product Engineer
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Similar jobs
Senior Deep Learning Framework Communications Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Deep Learning Communication Architect
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, Machine Learning Inference
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
AI Inference Performance Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Principal Software Engineer, E2E Performance and Goodput — CSP Engagements
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year