Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA
GPU @ 4
Leadership @ 4
Microservices
NVLink
Networking
Performance Optimization @ 4
Security
System Architecture @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA DGX systems are the foundation of advanced AI infrastructure, combining GPUs, CPUs, NVLink, NVIDIA Networking, and a fully optimized AI software stack. This role is responsible for the end-to-end delivery of DGX compute systems, from firmware through the AI stack to customer deployment.
Responsibilities
- Ensure every DGX platform is ready for the full NVIDIA software stack, including firmware, DGX OS, GPU drivers, CUDA Toolkit, DCGM, DOCA/OFED, and management tools.
- Own the general availability software and firmware release process, delivering firmware bundles, BaseOS ISOs, and release notes to OEM and OSV partners.
- Ensure platforms support AI agents such as NemoClaw and Hermes agents, NIM microservices, and expected customer workloads.
- Lead development of the manageability firmware stack, including BMC and BIOS, across all DGX platforms.
- Integrate firmware from GPU, CPU, and networking partner teams at the system level.
- Manage third-party vendors and drive platform requirements across firmware areas.
- Define validation strategies covering firmware regression, NVQual certification, deep-learning workload performance, OS/CUDA stack testing, multi-user scenarios, power and thermal validation, and field upgrade reliability.
- Establish quality gates and maintain a zero ship-stopper discipline.
- Drive platform bring-up for each new DGX system, coordinating first boot across new CPU and GPU silicon, board design, and firmware teams.
- Own architectural strategy for next-generation platforms, including firmware update mechanisms, system security posture, and AI application readiness.
- Ensure firmware release processes meet cloud service provider and enterprise deployment requirements.
- Represent DGX platform readiness in executive reviews and strategic planning with VP and SVP leadership.
- Engage with industry standards bodies, including DMTF Redfish and OCP.
- Own the complete DGX delivery lifecycle, including system architecture, firmware development, integration, full-stack validation, general availability release, and customer deployment.
- Align GPU, CPU, networking, security, operating system, and AI software teams across NVIDIA.
- Own root cause and corrective action processes for field issues.
- Manage external vendor partnerships, including AMI for SBIOS and BMC contributors.
- Build and lead an engineering organization, mentor leaders, and foster technical excellence, intellectual honesty, and customer focus.
Requirements
- BS or MS in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 12 or more years of experience in systems firmware or software engineering.
- 5 or more years of engineering leadership experience.
- Deep expertise in the server system stack, including SBIOS, BMC, operating systems, applications, and system-level integration of complex multi-component products.
- Proven experience delivering multi-generation server or data center platforms from architecture through customer deployment.
- Experience managing engineering organizations across multiple geographies in a matrix environment.
- Strong understanding of server hardware, including CPUs, GPUs, interconnects, memory, PCIe, and power delivery.
- Experience owning end-to-end product quality, from firmware validation through full-stack system testing and field deployment.
Preferred Qualifications
- Experience with NVIDIA DGX or GPU-accelerated server platforms.
- Experience driving server bring-up for new silicon and system architecture redesigns.
- Familiarity with DMTF Redfish, OCP standards, and server manageability ecosystems.
- Experience with AI or deep-learning workload validation and platform-level performance optimization.
- Ability to operate at the VP/SVP level and influence cross-business-unit strategic decisions.
Benefits
The base salary range is USD 320,000 to USD 488,750. The role also includes eligibility for equity and benefits. Applications will be accepted at least until August 1, 2026. NVIDIA is an equal opportunity employer.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Principal Architect, System Software - Orbital Data Center
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Firmware Engineer – Server Manageability and Observability
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Distinguished Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Distinguished Engineer – Data Center System Software Architect
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Senior MLOps Engineer - DSX Enablement
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer – Data Center Compute Diagnostics
Nvidia · Durham, United States
USD 224,000-356,500 per year