Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Cloud Computing
Communication @ 7
GPU
HPC
InfiniBand @ 3
Linux @ 7
Mentoring @ 6
NVLink @ 3
Python @ 6
Security @ 7
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is developing data center platforms such as GB200 NVL72 for AI, HPC, and cloud computing. This role will lead innovation in diagnostics for NVIDIA's partner ecosystem, shaping how complex server platforms are validated, debugged, and optimized across ODM factories, Cloud Service Provider deployments, and field operations.
Responsibilities
- Develop diagnostic systems for NVIDIA data center platforms, including hardware and software tools for worst-case stress workloads involving CPUs, GPUs, memory, storage, and interconnects.
- Lead platform bring-up and integration, ensuring diagnostics are embedded effectively across the server lifecycle.
- Drive hardware validation strategy with architecture and hardware teams, creating validation plans for new server generations.
- Analyze the root causes of complex failures, act as a Level 2 engineering contact for critical issues, and provide scalable solutions across the stack.
- Develop diagnostic software to ensure quality and performance at scale across ODM and partner production lines.
- Mentor and grow engineering teams while providing technical leadership and fostering innovation and excellence.
- Influence long-term strategy by developing diagnostic architectures and roadmaps for upcoming NVIDIA and partner products.
Requirements
- Proven experience architecting diagnostics for complex server systems, particularly at the software and hardware interface.
- Deep systems knowledge covering x86 and ARM architectures, Linux and Windows operating system internals, UEFI/BIOS firmware, BMC, and platform security.
- Ability to evaluate system development tradeoffs and drive optimal solutions with customers and multidisciplinary teams.
- Expertise in C, C++, and Python for tool development and automation.
- Familiarity with high-speed interconnects such as PCIe, InfiniBand, NVLink, and Ethernet.
- Strong communication skills for engaging with technical and executive teams.
- Bachelor's or master's degree, or equivalent experience, in Computer Science, Electrical Engineering, or a related field.
- At least 8 years of engineering experience in diagnostics, embedded systems, or cloud platforms.
Preferred Qualifications
- Experience driving diagnostics across rack-level or cluster-level deployments.
- Background in cloud-scale infrastructure and partner engagement.
- Demonstrated success influencing product direction and vendor roadmaps.
- Passion for mentoring and building high-performing teams.
Compensation and Benefits
- Base salary range: USD 184,000–287,500 per year.
- Eligible for equity and benefits.
- NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.
- Applications will be accepted at least until February 15, 2026.
More jobs at Nvidia
Senior Math Libraries Engineer - LLM Integration and Developer Experience
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Architect, Networking AI
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Manager, CUDA Driver
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Custom SoC IP Verification Engineer
Nvidia · Santa Clara, United States
USD 168,000-310,500 per year
Research Intern, Fundamental Generative AI - 2027
Nvidia · Santa Clara, United States
USD 38-94 per hour
Similar jobs
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Senior Technical Marketing Engineer - DSX AI Infrastructure Software
Nvidia · Santa Clara, United States
USD 160,000-322,000 per year
Senior HPC AI Cluster Engineer
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Distinguished Engineer, Storage – AI Cloud
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior Software Engineer - Cluster Networking
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Site Reliability Engineer, DGX Cloud
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year