Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agentic Systems @ 4
Algorithms
CUDA @ 4
Communication @ 4
Deep Learning @ 4
GPU
GenAI
Generative AI @ 4
JAX @ 6
MPI @ 4
Machine Learning @ 4
NCCL @ 4
Performance Optimization @ 4
Profiling @ 7
PyTorch @ 6
Python @ 7
Reinforcement Learning @ 3
SGLang @ 6
TensorRT @ 6
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for an experienced and highly motivated software professional to work on pioneering initiatives and projects at the intersection of CUDA and deep learning systems. The team explores advanced deep learning architectures, massive-scale distributed computing, low-level hardware optimization, model optimization, custom kernel development, and cluster-scale AI systems design.
Responsibilities
- Explore, research, and prototype novel systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA through modeling, simulation, and silicon prototyping.
- Architect and optimize distributed computing systems that scale from a single node to massive, cluster-scale supercomputing environments.
- Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads.
- Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
- Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms that improve accelerator compute utilization, memory bandwidth, cross-node network communication efficiency, and programmability.
- Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
- Write clean, effective, and maintainable code so exploratory prototypes can transition into open-source releases, upstream framework integrations, internal tools, or commercial products.
Requirements
- BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
- 8+ years of relevant industry experience or equivalent academic experience after degree achievement.
- Strong proficiency in C++ and Python.
- Solid understanding of deep learning fundamentals, with a focus on transformers.
- Strong understanding of distributed computing principles, multi-node scaling, and the performance challenges of cluster-scale execution.
- Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
- Familiarity with deep learning accelerator architectures such as GPUs, along with hands-on CUDA programming and kernel optimization experience.
- Strong analytical skills and experience using profiling tools to understand software performance on hardware.
- Experience profiling and optimizing vision models, generative AI architectures, or diffusion models.
- Background in deep learning compilers, including graph-level and code-generation technologies such as Triton, XLA, and torch.compile.
Preferred Qualifications
- Deep expertise in the performance internals and execution graphs of major deep learning autograd, training, and inference frameworks, including PyTorch, JAX, TensorRT, vLLM, sgLang, NeMo, Megatron, and MaxText.
- Hands-on experience with CUDA communication libraries such as NCCL, MPI, and UCX, and distributed machine learning techniques such as pipeline parallelism and tensor parallelism.
- Knowledge of numerical methods and low-precision arithmetic, including NVFP4, MXFP4, FP8, and INT8, and their implications for model accuracy and performance.
- Familiarity with systems requirements for reinforcement learning or highly parallel simulation environments, and/or a research background in machine learning systems or adjacent fields.
- Experience applying machine learning, especially agentic systems, to systems problems.
Compensation and Benefits
The base salary is determined based on location, experience, and the pay of employees in similar positions. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Employees are also eligible for equity and benefits.
Applications will be accepted at least until July 1, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.