Principal Firmware Engineer – Server Manageability and Observability

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI CUDA @ 3 GPU HPC @ 3 InfiniBand @ 4 Linux @ 6 NVLink Networking @ 4 Security @ 4 System Architecture @ 4

Details

NVIDIA data center systems, including DGX and HGX, combine NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and an optimized NVIDIA AI and HPC software stack. This role owns the end-to-end system software architecture of these products, including firmware, kernel drivers, operating systems, and user-mode drivers. The position involves collaboration with internal component leads and industry-leading cloud service providers to bring products to market.

Responsibilities

  • Serve as the primary technical point of contact for major customers by leading technology discussions, defining KPIs, gathering requirements, and addressing complex technical queries.
  • Lead technical innovation and strategic collaborations with major hyperscalers as a system software architect for next-generation data center products.
  • Align NVIDIA's roadmap with major customer requirements through direct engagement.
  • Develop and drive adoption of new technologies and protocols.
  • Make critical technical decisions in ambiguous situations and mitigate risks through left-shift strategies.
  • Lead complex, cross-functional projects to completion and influence outcomes without direct authority in large-scale, collaborative environments.
  • Implement left-shift strategies to de-risk program execution.

Requirements

  • Deep expertise in scalable and performant server system architecture, with a focus on software/hardware interfaces.
  • Extensive experience with complex system software for accelerators, including GPUs, DPUs, and FPGAs.
  • Mastery of system firmware, including SBIOS and OpenBMC, embedded systems, and Linux kernel internals.
  • Proficiency in out-of-band and in-band management architectures.
  • Experience with device management protocols such as MCTP, PLDM, SPDM, and RDE, as well as system management protocols including Redfish and IPMI.
  • Extensive knowledge of networking technologies and protocols, including TCP/IP, Ethernet, and InfiniBand, along with advanced switching and routing concepts.
  • Experience collaborating with platform security experts to define tradeoffs between security and ease of use.
  • Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 15 or more years of experience in system architecture and design.

Preferred Qualifications

  • Knowledge of cloud- and cluster-level deployment and management systems.
  • Participation in or contributions to standards bodies such as OCP and DMTF.
  • Familiarity with NVIDIA HPC programming models and libraries, including CUDA, cuDNN, and DOCA.
  • Knowledge of enterprise storage architectures and distributed parallel processing paradigms.

Benefits

The role includes eligibility for equity and benefits. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs