Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 7
CUDA @ 4
Communication @ 7
Data Structures @ 7
Deep Learning @ 7
Distributed Systems @ 4
GPU @ 4
GenAI
Generative AI
LLM @ 4
Machine Learning @ 4
Mentoring @ 1
NLP
PyTorch @ 7
TensorFlow @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Research Engineer passionate about generative AI inference. The team develops optimized inferencing technologies to support generative AI workloads and contributes across the machine learning lifecycle, including conceptualization, applied research, optimized inference engineering, and deployment. The role involves collaboration with research teams, engineers, and the open-source community.
Responsibilities
- Design and evaluate routing policies for large language model traffic to make effective use of mixture-of-model systems.
- Build and run agentic benchmarks, such as Terminal-Bench, to measure algorithm quality and convert results into calibration data and routing profiles.
- Contribute to open-source repositories through design documentation, code reviews, documentation, and community contributions.
- Collaborate with engineering teams across NVIDIA to ensure software integrates seamlessly across the NVIDIA accelerated serving stack.
Requirements
- Bachelor's or master's degree in Computer Science or equivalent experience.
- 8+ years of industry experience with deep learning frameworks such as PyTorch or TensorFlow.
- Experience designing or running LLM evaluations or benchmarks, ideally agentic benchmarks, and drawing statistically sound conclusions.
- Understanding of modern techniques in machine learning, deep neural networks, natural language processing, or speech recognition.
- An empirical research mindset, including forming hypotheses about new algorithms, running calibrations, and iterating on results.
- Strong communication and interpersonal skills, with the ability to work in a dynamic and distributed team.
- Strong computer science fundamentals, including algorithms and data structures, computational complexity, parallel and distributed computing, and system software.
- A desire to continuously grow and learn new things.
- Mentoring experience with junior engineers and interns is a significant advantage.
Preferred Qualifications
- Experience architecting or developing large-scale distributed systems for deep learning.
- Experience creating agentic benchmarks and publishing research.
- Knowledge of CPU and/or GPU architecture.
- GPU programming experience with CUDA.
Compensation and Benefits
- Base salary range for Level 4: USD 192,000–304,750 per year.
- Base salary range for Level 5: USD 224,000–356,500 per year.
- Eligibility for equity and benefits.
- The base salary is determined by location, experience, and compensation for employees in similar positions.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Senior Software Engineer - Local AI
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Machine Learning Engineer, Accelerated Apache Spark
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, Deep Learning Inference – TensorRT
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Senior AI Systems and Algorithms Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year