Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 6
Communication @ 4
Deep Learning @ 4
GPU
HPC @ 4
JAX @ 4
LLM @ 4
Performance Analysis
PyTorch @ 4
Python @ 6
SGLang @ 4
System Architecture @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Responsibilities
- Integrate new CUDA features and Runtime abstractions in AI frameworks: from PoC to performance analysis to production
- Perform deep analysis of AI workloads and frameworks to identify requirements and opportunities to innovate in the lower layers of the stack. Collaborate hands-on with teams working on the latest AI models
- Own and drive improvements in the AI Compiler-Runtime interface to build speed-of-light multi-GPU multi-node solutions
- Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads
- Influence the roadmap of core CUDA to facilitate building next-gen DL frameworks
- Collaborate with a very dynamic team across multiple time zones
- Collaborate closely with AI researchers, HW and SW architects, kernel and compiler authors and CUDA driver experts to co-design systems and frameworks that enhance performance and programmability
- Develop exploratory tools and runtime systems to profile and accelerate new paradigms in deep learning
- Write clean, effective, and maintainable code, ensuring exploratory prototypes can smoothly transition into open-source releases, upstream framework integrations, internal tools, or closed-source commercial products
Requirements
- BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience)
- 8+ years of relevant industry experience or equivalent academic experience after completed degree
- Development experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang
- Rapid prototyping and development with Python, C++, CUDA or related DSLs
- Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile)
- Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems)
- Understanding of HPC/AI communication concepts
- Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals)
- Adaptability and passion to learn new frameworks and tools
- Flexibility to work and communicate effectively across different teams and timezones
Benefits
- Eligible for equity and benefits
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Deep Learning Framework Communications Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Ai Networking
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, CUTLASS Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Deep Learning Communication Architect
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Deep Learning Software Engineer, TensorRT Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
AI Inference Performance Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year