Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Cloud Computing
Communication @ 7
GPU
HPC
InfiniBand @ 3
Linux @ 7
Mentoring @ 6
NVLink @ 3
Python @ 6
Security @ 7
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is developing data center platforms such as GB200 NVL72 for AI, HPC, and cloud computing. This role will lead innovation in diagnostics for NVIDIA's partner ecosystem, shaping how complex server platforms are validated, debugged, and optimized across ODM factories, Cloud Service Provider deployments, and field operations.
Responsibilities
- Develop diagnostic systems for NVIDIA data center platforms, including hardware and software tools for worst-case stress workloads involving CPUs, GPUs, memory, storage, and interconnects.
- Lead platform bring-up and integration, ensuring diagnostics are embedded effectively across the server lifecycle.
- Drive hardware validation strategy with architecture and hardware teams, creating validation plans for new server generations.
- Analyze the root causes of complex failures, act as a Level 2 engineering contact for critical issues, and provide scalable solutions across the stack.
- Develop diagnostic software to ensure quality and performance at scale across ODM and partner production lines.
- Mentor and grow engineering teams while providing technical leadership and fostering innovation and excellence.
- Influence long-term strategy by developing diagnostic architectures and roadmaps for upcoming NVIDIA and partner products.
Requirements
- Proven experience architecting diagnostics for complex server systems, particularly at the software and hardware interface.
- Deep systems knowledge covering x86 and ARM architectures, Linux and Windows operating system internals, UEFI/BIOS firmware, BMC, and platform security.
- Ability to evaluate system development tradeoffs and drive optimal solutions with customers and multidisciplinary teams.
- Expertise in C, C++, and Python for tool development and automation.
- Familiarity with high-speed interconnects such as PCIe, InfiniBand, NVLink, and Ethernet.
- Strong communication skills for engaging with technical and executive teams.
- Bachelor's or master's degree, or equivalent experience, in Computer Science, Electrical Engineering, or a related field.
- At least 8 years of engineering experience in diagnostics, embedded systems, or cloud platforms.
Preferred Qualifications
- Experience driving diagnostics across rack-level or cluster-level deployments.
- Background in cloud-scale infrastructure and partner engagement.
- Demonstrated success influencing product direction and vendor roadmaps.
- Passion for mentoring and building high-performing teams.
Compensation and Benefits
- Base salary range: USD 184,000–287,500 per year.
- Eligible for equity and benefits.
- NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.
- Applications will be accepted at least until February 15, 2026.
More jobs at Nvidia
Engineering Manager, Data Labeling Platform
Nvidia · Santa Clara, United States
USD 200,000-391,000 per year
Engineering Manager, Local AI Agents
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Deep Learning Software Engineer, DLSim
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Fleet Intelligence Backend
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Business Systems Analyst
Nvidia · Santa Clara, United States
USD 144,000-270,200 per year
Similar jobs
Senior System Software Engineer – Data Center Compute Diagnostics
Nvidia · Durham, United States
USD 224,000-356,500 per year
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior HPC AI Cluster Engineer
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Distinguished Engineer, Storage – AI Cloud
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior HPC Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior MLOps Engineer - DSX Enablement
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer – Platform Engineering
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year