Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 6
GPU
HPC
InfiniBand @ 7
NVLink
Networking @ 7
Security
System Architecture @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a highly motivated technical leader to design, drive, and operationalize rack-scale factory and deployment flows for next-generation data center products. The role focuses on reliable, debuggable, and scalable manufacturing and deployment solutions for HGX, DGX, and large multi-node NVLink domain rack architectures. These systems combine NVIDIA GPUs, NVLink, InfiniBand networking, Grace CPUs, and an optimized AI and HPC software stack.
Responsibilities
- Lead and drive rack-scale/L11 flows for factory and initial data center deployment.
- Design and implement end-to-end factory workflows, including firmware flashing sequences, security provisioning, and deployment of software mitigations.
- Collaborate with data center architects, ODMs, and OEMs to define factory and data center requirements that support efficient and reliable production ramp-up.
- Champion reliability, debuggability, and optimization in firmware, diagnostic, and deployment tool design.
- Drive pre-silicon readiness for factory and manufacturing workflows for rack-scale products using NVIDIA simulation and emulation technology.
- Mentor architects and engineering teams to help develop future leaders.
- Make key technical decisions in ambiguous situations.
Requirements
- Bachelor’s or master’s degree in Computer Engineering, Computer Science, or a related field, or equivalent experience.
- At least 8 years of experience in system architecture and design.
- Deep experience designing architectures for scalable and high-performance server systems, particularly at the software/hardware interface.
- Strong understanding of networking technologies and protocols, including Ethernet and InfiniBand.
- Experience working with complex system software for accelerators such as GPUs, DPUs, or FPGAs.
- Expertise in out-of-band and in-band management architectures.
- Knowledge of system management protocols such as Redfish and IPMI.
- Demonstrable experience implementing left-shift strategies to de-risk program execution.
- Excellent written and verbal communication skills.
Preferred Qualifications
- Knowledge of large-scale cloud and cluster-level deployment and management systems.
- A demonstrated track record of leading data center products across their entire lifecycle, including inception, pre-silicon development, post-silicon bring-up, manufacturing, and deployment.
Compensation and Benefits
- Base salary range: $184,000–$287,500 for Level 4.
- Base salary range: $224,000–$356,500 for Level 5.
- Eligibility for equity and benefits.
- Applications will be accepted at least until July 22, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Distinguished Engineer - Rack Scale Architecture
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Director, Rack-Scale Software Architecture
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior Software Engineer - NVLink Rack Scale Stability and Reliability
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Director, Software TPM - Server Firmware and System Software
Nvidia · Santa Clara, United States
USD 272,000-425,500 per year
Principal Firmware Engineer – Server Manageability and Observability
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Distinguished Software Engineer - NVLink Fusion Software
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year