Distinguished Engineer – Data Center System Software Architect
at Nvidia
USD 320,000-488,800 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 3
GPU
HPC @ 3
InfiniBand @ 4
Linux @ 6
NVLink
Networking @ 4
Security @ 4
System Architecture @ 8
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA data center systems, including DGX and HGX, combine NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and an optimized NVIDIA AI and HPC software stack. The role owns the end-to-end architecture of these products at the system software level, including firmware, kernel drivers, operating systems, and user-mode drivers. You will work with internal component leads and engage with leading cloud service providers to bring these products to market.
Responsibilities
- Serve as the primary technical point of contact for major customers, leading technology discussions, defining KPIs, gathering requirements, and addressing complex technical queries.
- Lead technical innovation and strategic collaborations with major hyperscalers to architect next-generation data center products.
- Align NVIDIA's roadmap with major customers' requirements through direct engagement.
- Develop and drive adoption of new technologies and protocols.
- Make critical technical decisions in ambiguous situations and mitigate risks through left-shift strategies.
Requirements
- Deep expertise in scalable and performant server system architecture, with a focus on software/hardware interfaces.
- Extensive experience with complex system software for accelerators, including GPUs, DPUs, and FPGAs.
- Mastery of system firmware, including SBIOS and OpenBMC, embedded systems, and Linux kernel internals.
- Proficiency in out-of-band and in-band management architectures.
- Experience with device management protocols such as MCTP, PLDM, SPDM, and RDE, as well as system management protocols including Redfish and IPMI.
- Extensive knowledge of networking technologies and protocols, including TCP/IP, Ethernet, and InfiniBand, along with advanced switching and routing concepts.
- Experience collaborating with platform security experts to define tradeoffs between security and ease of use.
- Demonstrated success leading complex, cross-functional projects to completion and influencing outcomes without direct authority in large-scale, collaborative environments.
- Demonstrable experience implementing left-shift strategies to de-risk program execution.
- Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 20+ years of experience in system architecture and design.
Preferred Qualifications
- Knowledge of cloud- and cluster-level deployment and management systems.
- Participation in and contributions to standards bodies such as OCP and DMTF.
- Familiarity with NVIDIA HPC programming models and libraries, including CUDA, cuDNN, and DOCA.
- Knowledge of enterprise storage architectures and distributed parallel processing paradigms.
Benefits
- Equity and benefits are provided.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
- Applications will be accepted at least until August 3, 2026.
More jobs at Nvidia
Senior System Software Engineer - CPU SoC Boot Firmware
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
System Software Engineer – Data Center Compute Diagnostics
Nvidia · Durham, United States
USD 152,000-241,500 per year
Senior Research Scientist - Generative World Models for Autonomous Driving and Physical AI
Nvidia · Santa Clara, United States
USD 192,000-356,500 per year
Senior Technical Marketing Engineer - CAE Performance
Nvidia · Santa Clara, United States
USD 136,000-253,000 per year
Senior Software Technical Program Driver - OEM and NCP Escalations
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
Similar jobs
Principal Firmware Engineer – Server Manageability and Observability
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer - Manufacturing and Factory
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer - NVLink Rack Scale Stability and Reliability
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Software Engineer - Rack-Scale Systems Infrastructure
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Distinguished Engineer - Rack Scale Architecture
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year