Senior Software Engineer, SRE and Production Engineering - DGX Cloud

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI Communication @ 6 Debugging @ 4 GPU @ 4 Go @ 7 IaaS InfiniBand @ 6 Kubernetes @ 4 Linux @ 4 NVLink @ 6 Networking @ 6 Observability @ 4 Python @ 7 SRE @ 4

Details

NVIDIA DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. The team is seeking software engineers with SRE or production engineering experience and hands-on experience with bare-metal NVIDIA systems. The team builds software and operational tooling that moves GPU capacity from installed hardware to production service in an IaaS production environment supporting BMaaS and VMaaS.

Responsibilities

  • Build automation for bare-metal provisioning, hardware validation, firmware and software upgrades, repair, and cluster lifecycle management.
  • Develop tools that interact with BMC and Redfish interfaces to monitor hardware health, manage server state, and support recovery workflows.
  • Operate and advance NVIDIA NVL72 systems and BlueField-3 or later DPUs across cloud partner and on-premises environments.
  • Diagnose failures across servers, DPUs, GPU systems, CPU systems, networking, Linux, and Kubernetes, and turn recurring issues into automated detection and repair.
  • Define validation and handoff criteria so new capacity enters production safely and consistently.
  • Participate in on-call duties, incident response, root-cause analysis, and follow-up work to implement permanent solutions.
  • Collaborate with hardware, networking, platform, data center operations, and partner teams to resolve issues across ownership boundaries.

Requirements

  • 5+ years of experience building software for or operating production infrastructure, including substantial hands-on bare-metal experience.
  • Strong Go or Python skills, with a record of delivering production automation and services.
  • Direct experience with BMC and Redfish for server provisioning, health inspection, power management, or fault diagnosis.
  • Practical experience working directly with NVIDIA GPU hardware, such as NVL72 systems, and BlueField-3 or newer DPUs.
  • Experience with Linux, firmware and driver management, network boot, and the server lifecycle from initial provisioning through repair.
  • Experience managing production reliability through on-call duties, incident response, observability, and durable solutions.
  • Ability to debug failures across hardware, host operating systems, networking, and distributed services.
  • Clear communication skills and demonstrated ownership of problems spanning multiple teams.
  • BS/MS in Computer Science or equivalent practical experience.

Preferred Qualifications

  • Experience operating BlueField DPUs in DPU mode, including host-to-DPU connectivity and lifecycle debugging. Equivalent DPU experience is also considered.
  • Background with NVLink, InfiniBand, Spectrum-X, or GPU cluster performance validation.
  • Experience building safe, repeatable workflows for rack-scale bring-up, firmware upgrades, hardware replacement, and customer handoff.
  • Experience with Kubernetes, GitOps, Argo CD, SLOs, and fleet-wide automation.

Compensation and Benefits

The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. The role is also eligible for equity and benefits.

Applications will be accepted at least until October 3, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs