Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
GPU
HPC
LLM @ 7
Leadership @ 6
Performance Analysis @ 7
Performance Optimization @ 6
Technical Leadership @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's accelerated computing platform is enabling generational improvements in large language models, while the scale and complexity of these models create new challenges in computational efficiency. This role will drive a unified strategy for making LLMs more efficient from research through deployment by combining model innovation, systems expertise, and hardware awareness. The position will lead a multidisciplinary effort, establish technical direction for LLM efficiency, and help shape how future models and computing platforms are designed together.
Responsibilities
- Lead cross-layer efforts to improve the efficiency of large language models across model architecture, training, and inference systems.
- Analyze how LLM workloads map to GPUs, memory systems, interconnects, and distributed infrastructure, identifying opportunities for model-system-hardware co-design.
- Establish a measurement-driven efficiency roadmap and lead projects from early investigation through production deployment.
- Partner with model researchers, systems engineers, compiler and kernel developers, and hardware architects to influence future model, software, and hardware roadmaps.
Requirements
- Master's or PhD degree, or equivalent experience, in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
- 5+ years of relevant experience in AI systems, model architecture, computer architecture, high-performance computing, or performance optimization.
- Strong understanding of LLM architectures, training and inference workloads, and the tradeoffs between model quality, computational cost, memory footprint, latency, throughput, and power.
- Strong background in performance analysis, roofline modeling, workload characterization, benchmarking, and hardware-aware optimization.
- Proven ability to provide technical leadership and drive complex optimization projects from concept to production.
Preferred Qualifications
- Track record of delivering measurable improvements in throughput, cost per token, energy per token, memory efficiency, or time to train.
- First-principles approach to improving LLM efficiency through measurement, modeling, optimization, and delivery.
- Familiarity with low-precision computation, quantization, sparsity, Mixture-of-Experts, long-context inference, and speculative decoding.
- Experience co-designing model architectures with training, inference, compiler, or hardware constraints.
- Experience influencing accelerator, system, or datacenter architecture based on future AI workload requirements.
Benefits
- Equity and benefits are provided.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Deep Learning Algorithm Engineering Intern - 2026
Nvidia · Zurich, Switzerland
PLN 117,800-204,100 per year
Senior Project Delivery Manager - NVIS
Nvidia · United States
USD 168,000-322,000 per year
Principal Research Scientist, Synthetic Data Generation
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Staff Software Engineer - Agentic Automation
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Firmware Engineer, NIC Firmware
Nvidia · Seattle, United States
USD 152,000-287,500 per year
Similar jobs
Principal High-Performance LLM Training Engineer
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Engineering Manager, Agentic GenAI Platform
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Principal Software Engineer, E2E Performance and Goodput — CSP Engagements
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, CUTLASS Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal System Software Engineer - AV Platform
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year