Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agentic AI @ 4
Algorithms
CUDA @ 4
Communication @ 4
Deep Learning @ 4
GPU @ 3
GenAI
Generative AI @ 4
JAX @ 6
LLM
MPI @ 4
Machine Learning @ 4
NCCL @ 4
Performance Optimization @ 4
Profiling @ 4
PyTorch @ 6
Python @ 7
SGLang @ 6
TensorRT @ 6
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for an experienced and highly motivated software professional to work on pioneering initiatives and projects at the intersection of CUDA and deep learning systems. The team explores and prototypes next-generation ideas that bridge deep learning algorithms and CUDA, with a focus on modern accelerator architectures, model optimization, custom kernel development, and cluster-scale AI systems.
This research-oriented role focuses on maximizing hardware performance for emerging AI workloads, from single GPUs to supercomputer clusters. Exploratory prototypes may transition into open-source releases, upstream framework integrations, internal tools, or closed-source commercial products.
Responsibilities
- Explore, research, and prototype systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA through modeling, simulation, and silicon prototyping.
- Architect and optimize distributed computing systems that scale from a single node to massive, cluster-scale supercomputing environments.
- Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads.
- Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
- Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms.
- Improve accelerator compute utilization, memory bandwidth, cross-node network communication efficiency, and programmability.
- Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
- Write clean, effective, and maintainable code.
Requirements
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
- 8+ years of relevant industry experience or equivalent academic experience after degree achievement.
- Strong proficiency in C++ and Python.
- Solid understanding of deep learning fundamentals, particularly transformers.
- Strong understanding of distributed computing, multi-node scaling, and the performance challenges of cluster-scale execution.
- Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
- Familiarity with GPU and other deep learning accelerator architectures.
- Hands-on experience with CUDA programming, kernel optimization, and workload profiling.
- Experience profiling and optimizing generative AI models, including large language models.
- Research background in machine learning systems or adjacent fields, with experience profiling and optimizing vision models, generative AI architectures, or diffusion models.
- Track record of initiative and willingness to investigate problems across the stack.
Preferred Qualifications
- Deep expertise in the performance internals and execution graphs of deep learning training and inference frameworks such as PyTorch, JAX, TensorRT, vLLM, sgLang, NeMo, or Megatron.
- Hands-on experience with communication libraries such as NCCL, MPI, or UCX.
- Experience with distributed machine learning techniques, including pipeline, tensor, and expert parallelism.
- Knowledge of numerical methods and low-precision arithmetic, including NVFP4, MXFP4, FP8, and INT8, and their impact on deep learning accuracy and performance.
- Background in deep learning compilers and ML systems, including graph-level and code-generation tools such as Triton, XLA, and torch.compile.
- Experience with highly parallel or reinforcement-learning-style simulation environments.
- Experience designing and implementing agentic AI systems for complex systems and infrastructure problems.
Compensation and Benefits
- Base salary for Level 4: USD 184,000–287,500 per year.
- Base salary for Level 5: USD 224,000–356,500 per year.
- Eligibility for equity and benefits.
- Applications will be accepted at least until August 9, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
AI Inference Performance Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Deep Learning Frameworks CUDA Software Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Deep Learning Communication Architect
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, Machine Learning Inference
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year