Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Algorithms
CUDA @ 4
Deep Learning @ 8
GPU @ 7
GenAI
Generative AI
LLM @ 4
OpenCL @ 4
Profiling
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking outstanding engineers to help shape the future of large language model inference. The team focuses on improving the algorithmic performance and efficiency of systems that represent LLMs, developing new inference algorithms and protocols, improving existing models, and integrating these improvements into NVIDIA's solutions for large-scale, sophisticated tasks.
Responsibilities
- Explore and incorporate contemporary research on generative AI, agents, and inference systems into the NVIDIA LLM software stack.
- Conduct in-depth analysis, profiling, and optimization of agentic LLM workloads to reduce request latency and increase request throughput while maintaining workflow fidelity.
- Design and implement scalable systems to accelerate agentic workflows and efficiently handle sophisticated datacenter-scale use cases.
- Advise future iterations of NVIDIA software, hardware, and systems by collaborating with diverse teams at NVIDIA and external partners.
- Formalize strategic requirements presented by workload use cases.
Requirements
- Bachelor's, master's, or doctoral degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent experience.
- 15+ years of experience in deep learning and deep learning systems design.
- Proficiency in Python and C++ programming.
- Strong understanding of computer architecture and GPU/parallel datacenter computing fundamentals.
- Demonstrated interest in analyzing, modeling, and tuning application performance.
Preferred Qualifications
- Experience building large-scale LLM inference systems, especially systems involving compound AI.
- Experience with processor and system-level performance modeling.
- GPU programming experience with CUDA or OpenCL.
Benefits
The role includes eligibility for equity and benefits. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.
Applications will be accepted at least until July 26, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Senior Deep Learning Software Engineer, Inference
Nvidia · United States
USD 152,000-287,500 per year
AI Inference Performance Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Deep Learning Software Engineer, Inference
Nvidia · Netherlands
PLN 221,200-507,000 per year
Deep Learning Compiler Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year