Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 6
Deep Learning @ 7
GPU @ 3
Performance Analysis
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a senior engineer who is obsessed with performance analysis and optimization to help squeeze every last clock cycle out of AI training, the workload driving the design and construction of the largest and most powerful compute systems. The role involves working across all layers of the hardware/software stack—from GPU architecture to the application code—to achieve peak performance. The opportunity is to directly impact the hardware and software roadmap in a fast-growing technology company that leads the AI revolution.
Responsibilities
- Understand, analyze, profile, and optimize AI training workloads on state-of-the-art hardware and software platforms.
- Identify performance bottlenecks of AI training on GPUs, prioritize, and solve problems across key AI training workloads.
- Implement production-quality software across multiple layers of NVIDIA's deep learning platform stack, from drivers to DL frameworks.
- Build and support NVIDIA submissions for MLPerf Training benchmarks.
- Implement key DL training workloads in NVIDIA's proprietary processor and system simulators to enable future architecture studies.
- Develop tools to automate workload analysis, optimization, and other critical workflows.
Requirements
- PhD in CS, EE or CSEE (or equivalent experience) with 5+ years of relevant experience; or MS with 8+ years of experience.
- Strong background in deep learning and neural networks, particularly in training.
- Solid understanding of computer architecture and familiarity with GPU architecture fundamentals.
- Proven background in analyzing and tuning application performance.
- Proven experience with processor and system-level performance modeling.
- Proficiency in programming with C++, Python, and CUDA.
Benefits
- Eligible for equity and benefits.
More jobs at Nvidia
Senior Software Engineer, Linux Platform
Nvidia · United States
USD 168,000-270,200 per year
Technical Marketing Engineer
Nvidia · Santa Clara, United States
USD 136,000-253,000 per year
GPU Verification Engineer - New College Grad 2026
Nvidia · Westford, United States
USD 136,000-264,500 per year
Senior Developer Technology Engineer - Agentic SoC Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Engineer - Enterprise Content and AI Data Platform
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Similar jobs
Senior Deep Learning Systems Architect
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior AI Compiler Engineer, Algorithms and Code-Generation
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Software Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Senior Systems Software Engineer, AI Stack And Performance - DGX Station
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Deep Learning Frameworks CUDA Software Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer - AI Performance And Efficiency Tools
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Deep Learning Systems Engineer, Datacenters
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year