Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API
CUDA @ 3
Communication @ 3
Data Analysis @ 3
Deep Learning @ 3
Experimentation @ 3
GPU @ 5
HPC @ 3
Python @ 3
Robotics @ 3
Software Development @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The AI revolution advances when computation becomes fast, efficient, and economical enough to turn new ideas into products at global scale. CUDA is a critical layer beneath the frameworks, libraries, and applications used across AI, deep learning, high-performance computing, graphics, automotive, robotics, and other CUDA-powered products.
This role focuses on designing and shipping production C/C++ features and optimizations in the CUDA driver and runtime. The work includes tracing workloads across application, operating-system, CPU, interconnect, and GPU boundaries; bringing up new platforms; and using performance evidence to influence future software and hardware direction.
Responsibilities
- Design, implement, validate, and ship performance-centric features and programming-model capabilities in the CUDA driver and runtime using maintainable, well-tested production C/C++.
- Optimize critical execution paths, including kernel launch, synchronization, memory management and movement, CPU–GPU coordination, and system interconnect use, for latency, throughput, bandwidth, efficiency, and scalability.
- Own complex performance problems end to end by understanding workloads, forming hypotheses, creating focused measurements and models, isolating root causes across software and hardware boundaries, implementing production solutions, and validating application-level impact.
- Establish performance expectations for current and future platforms, characterize new silicon, close software and hardware gaps, and drive performance readiness through product release.
- Translate workload and platform evidence into CUDA API and programming-model improvements, systems-software direction, and measurement-backed recommendations for future hardware architecture and implementation.
- Partner with application, library, framework, operating-system, driver, runtime, firmware, GPU architecture, silicon, product, and customer-facing teams.
- Communicate findings clearly and raise engineering quality through design and code reviews.
- Lead complex feature development and cross-layer investigations, define performance requirements and technical direction for major subsystems, mentor engineers, and shape hardware/software decisions for future product generations.
Requirements
- A BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.
- At least 2 years of relevant systems-software development experience.
- Strong production C/C++ systems-programming experience, including delivery of substantial features, optimizations, or production fixes in a complex codebase.
- Strong operating systems and concurrency foundations, including threads, synchronization, processes, virtual memory, and user/kernel interactions.
- Strong computer-architecture foundations, including processors, memory hierarchy, caching and coherence, data movement, and system interconnects.
- Demonstrated success improving real software performance through measurement, identification of limiting mechanisms, implementation of effective solutions, and quantitative validation.
- Sound technical judgment, ownership of ambiguous problems, and clear communication across organizational and disciplinary boundaries.
- Direct CUDA or GPU experience is valuable but not required when accompanied by deep systems-software, operating-systems, computer-architecture, and performance-engineering foundations.
Preferred Qualifications
- Experience developing GPU or accelerator drivers, runtimes, kernel software, firmware, compilers, or other performance-critical low-level systems.
- Experience with pre-silicon analysis, platform bring-up, performance modeling, or hardware/software co-design.
- Systems-level performance experience with AI/deep learning, HPC, graphics, automotive, robotics, or similarly demanding workloads.
- Evidence of technical inventions, such as software-performance patents, novel production designs, or measurement-backed recommendations that influenced a hardware revision or future architecture.
- Python or another scripting language used for focused experimentation, data analysis, or visualization.
Compensation and Benefits
- Base salary range: USD 124,000–195,500 per year.
- Eligible for equity and benefits.
- Applications will be accepted at least until September 5, 2026.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
More jobs at Nvidia
Senior NPI Program Manager
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
GPU PCIe and Boot Architect - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior AI Engineer, High Performance AI
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Salesforce CPQ Developer
Nvidia · Santa Clara, United States
USD 176,000-276,000 per year
Senior Technical Program Manager - LLM Safety
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Compiler Verification Engineer, Compute Performance – GPU
Nvidia · Austin, United States
USD 140,000-224,200 per year
Senior Applied Research Scientist, Multimodal Foundation Models – Healthcare
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, CUDA Rust Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Radar Perception Engineer, Obstacle Foundation Models – Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, Metropolis Vision AI
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal Perception Engineer, Obstacle Foundation Models - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Linux Kernel Systems Software Engineer – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Distinguished Resiliency and Safety Architect, GPU Diagnostics
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year