Distinguished Engineer – Data Center System Software Architect

at Nvidia
USD 320,000-488,800 per year
SENIOR
✅ On-site

Tech Stack

AI CUDA @ 3 GPU HPC @ 3 InfiniBand @ 4 Linux @ 6 NVLink Networking @ 4 Security @ 4 System Architecture @ 8

Details

NVIDIA data center systems, including DGX and HGX, combine NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and an optimized NVIDIA AI and HPC software stack. The role owns the end-to-end architecture of these products at the system software level, including firmware, kernel drivers, operating systems, and user-mode drivers. You will work with internal component leads and engage with leading cloud service providers to bring these products to market.

Responsibilities

  • Serve as the primary technical point of contact for major customers, leading technology discussions, defining KPIs, gathering requirements, and addressing complex technical queries.
  • Lead technical innovation and strategic collaborations with major hyperscalers to architect next-generation data center products.
  • Align NVIDIA's roadmap with major customers' requirements through direct engagement.
  • Develop and drive adoption of new technologies and protocols.
  • Make critical technical decisions in ambiguous situations and mitigate risks through left-shift strategies.

Requirements

  • Deep expertise in scalable and performant server system architecture, with a focus on software/hardware interfaces.
  • Extensive experience with complex system software for accelerators, including GPUs, DPUs, and FPGAs.
  • Mastery of system firmware, including SBIOS and OpenBMC, embedded systems, and Linux kernel internals.
  • Proficiency in out-of-band and in-band management architectures.
  • Experience with device management protocols such as MCTP, PLDM, SPDM, and RDE, as well as system management protocols including Redfish and IPMI.
  • Extensive knowledge of networking technologies and protocols, including TCP/IP, Ethernet, and InfiniBand, along with advanced switching and routing concepts.
  • Experience collaborating with platform security experts to define tradeoffs between security and ease of use.
  • Demonstrated success leading complex, cross-functional projects to completion and influencing outcomes without direct authority in large-scale, collaborative environments.
  • Demonstrable experience implementing left-shift strategies to de-risk program execution.
  • Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 20+ years of experience in system architecture and design.

Preferred Qualifications

  • Knowledge of cloud- and cluster-level deployment and management systems.
  • Participation in and contributions to standards bodies such as OCP and DMTF.
  • Familiarity with NVIDIA HPC programming models and libraries, including CUDA, cuDNN, and DOCA.
  • Knowledge of enterprise storage architectures and distributed parallel processing paradigms.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.
  • Applications will be accepted at least until August 3, 2026.

More jobs at Nvidia

Similar jobs