Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 7
CUDA @ 4
Communication @ 7
Data Structures @ 7
Deep Learning @ 7
Distributed Systems @ 4
GPU @ 4
GenAI
Generative AI
LLM @ 4
Machine Learning @ 4
Mentoring @ 1
NLP
PyTorch @ 7
TensorFlow @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Research Engineer passionate about generative AI inference. The team develops optimized inferencing technologies to support generative AI workloads and contributes across the machine learning lifecycle, including conceptualization, applied research, optimized inference engineering, and deployment. The role involves collaboration with research teams, engineers, and the open-source community.
Responsibilities
- Design and evaluate routing policies for large language model traffic to make effective use of mixture-of-model systems.
- Build and run agentic benchmarks, such as Terminal-Bench, to measure algorithm quality and convert results into calibration data and routing profiles.
- Contribute to open-source repositories through design documentation, code reviews, documentation, and community contributions.
- Collaborate with engineering teams across NVIDIA to ensure software integrates seamlessly across the NVIDIA accelerated serving stack.
Requirements
- Bachelor's or master's degree in Computer Science or equivalent experience.
- 8+ years of industry experience with deep learning frameworks such as PyTorch or TensorFlow.
- Experience designing or running LLM evaluations or benchmarks, ideally agentic benchmarks, and drawing statistically sound conclusions.
- Understanding of modern techniques in machine learning, deep neural networks, natural language processing, or speech recognition.
- An empirical research mindset, including forming hypotheses about new algorithms, running calibrations, and iterating on results.
- Strong communication and interpersonal skills, with the ability to work in a dynamic and distributed team.
- Strong computer science fundamentals, including algorithms and data structures, computational complexity, parallel and distributed computing, and system software.
- A desire to continuously grow and learn new things.
- Mentoring experience with junior engineers and interns is a significant advantage.
Preferred Qualifications
- Experience architecting or developing large-scale distributed systems for deep learning.
- Experience creating agentic benchmarks and publishing research.
- Knowledge of CPU and/or GPU architecture.
- GPU programming experience with CUDA.
Compensation and Benefits
- Base salary range for Level 4: USD 192,000–304,750 per year.
- Base salary range for Level 5: USD 224,000–356,500 per year.
- Eligibility for equity and benefits.
- The base salary is determined by location, experience, and compensation for employees in similar positions.
More jobs at Nvidia
User Interface - User Experience Designer
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior Application Engineer, HPC and AI for Physics
Nvidia · United States
USD 140,000-270,200 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior System Software Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Principal Machine Learning Engineer, Accelerated Apache Spark
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Learning Software Engineer, Inference and Model Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year