Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 3
Algorithms @ 4
Azure @ 3
Bash @ 6
CUDA @ 4
Cloud Computing @ 3
Communication @ 6
Debugging @ 4
Deep Learning @ 4
GCP @ 3
GPU
Go @ 6
HPC @ 4
InfiniBand @ 3
LLM
Machine Learning @ 7
NCCL @ 4
PyTorch @ 3
Python @ 6
Robotics
TensorFlow @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior AI/ML Performance and Efficiency Engineer focused on GPU clusters to support AI efficiency initiatives. The role involves improving efficiency across the technology stack, collaborating with researchers and engineering teams, and developing scalable solutions for AI and machine learning workloads.
Responsibilities
- Collaborate with AI/ML researchers to improve model efficiency, productivity, and cost savings.
- Build tools and frameworks and apply machine learning techniques to detect and analyze efficiency bottlenecks.
- Support innovative machine learning workloads across robotics, autonomous vehicles, large language models, video, and other areas.
- Collaborate across engineering organizations to improve the efficiency of hardware, software, and infrastructure usage.
- Monitor fleet-wide utilization patterns, analyze inefficiencies, identify new patterns, and deliver scalable solutions.
- Stay current with developments in AI/ML technologies, frameworks, and efficiency strategies, and advocate for their adoption.
Requirements
- Bachelor's degree or equivalent background in Computer Science or a related area, or equivalent experience.
- At least 5 years of experience designing and operating large-scale compute infrastructure.
- Strong understanding of modern machine learning techniques and tools.
- Experience investigating and resolving end-to-end training and inference performance issues.
- Debugging and optimization experience with Nsight Systems and Nsight Compute.
- Experience debugging large-scale distributed training using NCCL.
- Proficiency in Python, Go, and Bash.
- Familiarity with cloud computing platforms such as AWS, GCP, and Azure.
- Experience with parallel computing frameworks and paradigms.
- Commitment to ongoing learning about technologies and methods in AI/ML infrastructure.
- Excellent communication and collaboration skills.
Preferred Qualifications
- Experience with NVIDIA GPUs, CUDA programming, NCCL, and MLPerf benchmarking.
- Knowledge of machine learning and deep learning concepts, algorithms, and models.
- Familiarity with InfiniBand, IBOP, and RDMA.
- Understanding of distributed storage systems such as Lustre and GPFS for AI/HPC workloads.
- Familiarity with PyTorch and TensorFlow.
Compensation and Benefits
- Base salary range: $152,000–$241,500 USD for Level 3.
- Base salary range: $184,000–$287,500 USD for Level 4.
- Compensation depends on location, experience, and pay for similar positions.
- Eligible employees also receive equity and benefits.
- Applications will be accepted at least until March 23, 2026.
- NVIDIA is an equal opportunity employer and uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Senior AI Infrastructure Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Solutions Architect
Nebius · United States, Canada
USD 250,000-320,000 per year
Distinguished Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Senior Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year