Senior Software Engineer, SRE and Production Engineering - DGX Cloud
at Nvidia
USD 152,000-287,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 6
Debugging @ 4
GPU @ 4
Go @ 7
IaaS
InfiniBand @ 6
Kubernetes @ 4
Linux @ 4
NVLink @ 6
Networking @ 6
Observability @ 4
Python @ 7
SRE @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. The team is seeking software engineers with SRE or production engineering experience and hands-on experience with bare-metal NVIDIA systems. The team builds software and operational tooling that moves GPU capacity from installed hardware to production service in an IaaS production environment supporting BMaaS and VMaaS.
Responsibilities
- Build automation for bare-metal provisioning, hardware validation, firmware and software upgrades, repair, and cluster lifecycle management.
- Develop tools that interact with BMC and Redfish interfaces to monitor hardware health, manage server state, and support recovery workflows.
- Operate and advance NVIDIA NVL72 systems and BlueField-3 or later DPUs across cloud partner and on-premises environments.
- Diagnose failures across servers, DPUs, GPU systems, CPU systems, networking, Linux, and Kubernetes, and turn recurring issues into automated detection and repair.
- Define validation and handoff criteria so new capacity enters production safely and consistently.
- Participate in on-call duties, incident response, root-cause analysis, and follow-up work to implement permanent solutions.
- Collaborate with hardware, networking, platform, data center operations, and partner teams to resolve issues across ownership boundaries.
Requirements
- 5+ years of experience building software for or operating production infrastructure, including substantial hands-on bare-metal experience.
- Strong Go or Python skills, with a record of delivering production automation and services.
- Direct experience with BMC and Redfish for server provisioning, health inspection, power management, or fault diagnosis.
- Practical experience working directly with NVIDIA GPU hardware, such as NVL72 systems, and BlueField-3 or newer DPUs.
- Experience with Linux, firmware and driver management, network boot, and the server lifecycle from initial provisioning through repair.
- Experience managing production reliability through on-call duties, incident response, observability, and durable solutions.
- Ability to debug failures across hardware, host operating systems, networking, and distributed services.
- Clear communication skills and demonstrated ownership of problems spanning multiple teams.
- BS/MS in Computer Science or equivalent practical experience.
Preferred Qualifications
- Experience operating BlueField DPUs in DPU mode, including host-to-DPU connectivity and lifecycle debugging. Equivalent DPU experience is also considered.
- Background with NVLink, InfiniBand, Spectrum-X, or GPU cluster performance validation.
- Experience building safe, repeatable workflows for rack-scale bring-up, firmware upgrades, hardware replacement, and customer handoff.
- Experience with Kubernetes, GitOps, Argo CD, SLOs, and fleet-wide automation.
Compensation and Benefits
The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. The role is also eligible for equity and benefits.
Applications will be accepted at least until October 3, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Software Engineer, DGX Cloud AI Infrastructure - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Nvidia · Durham, United States
USD 272,000-431,200 per year
Tegra System Software Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Developer Tools for Cloud
Nvidia · United States
USD 152,000-287,500 per year
Data and Platform Engineer
Nvidia · United States
USD 200,000-322,000 per year
Similar jobs
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Engineer, NCX
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Senior Technical Marketing Engineer - DSX AI Infrastructure Software
Nvidia · Santa Clara, United States
USD 160,000-322,000 per year
Senior MLOps Engineer - DSX Enablement
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year