Senior Production Engineer - DGX Cloud

at Nvidia
USD 168,000-333,500 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 Algorithms @ 4 Communication @ 7 Data Structures @ 4 DevOps @ 4 Distributed Systems @ 6 GPU Go @ 4 Hiring @ 4 Kubernetes @ 7 Mathematics @ 4 Observability Python @ 4 SRE @ 4 Slurm @ 7

Details

NVIDIA is hiring experienced Senior Production Engineers to help scale its AI infrastructure. Production Engineering is treated as a software engineering discipline, with significant contributions expected to the codebase. The role focuses on reliability assessments, incident management, production system observability, monitoring and alerting, automated deployments, and toil elimination.

NVIDIA's DGX Cloud team develops production systems that enable large-scale GPU clusters for a variety of AI workloads, including custom software for GPU asset provisioning, configuration, and lifecycle management across cloud providers.

Responsibilities

  • Develop and operate production systems supporting scalable GPU clusters for AI workloads.
  • Build custom software for GPU asset provisioning, configuration, and lifecycle management across cloud providers.
  • Implement monitoring and health management capabilities to improve the reliability, availability, and scalability of GPU assets.
  • Analyze data from GPU hardware diagnostics, cluster telemetry, and network telemetry.
  • Work with teams across NVIDIA to ensure production AI clusters operate reliably, consistently, and at maximum performance.
  • Evaluate system failures and improve services through a well-defined incident management process.
  • Contribute substantially to NVIDIA's production engineering codebase.

Requirements

  • Direct experience in a Production Engineering, DevOps, or Site Reliability Engineering role within a highly technical organization, with demonstrable impact.
  • Strong communication skills and the ability to work with multifunctional teams, principals, and architects across organizational boundaries and geographies.
  • 8 or more years of experience in a similar role and experience operating large-scale production systems.
  • Experience with Production Engineering, DevOps, or SRE principles, tools, and techniques.
  • A bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a comparable discipline, or equivalent experience.
  • Technical knowledge of a systems programming language such as Go or Python.
  • Solid understanding of data structures and algorithms.

Preferred Qualifications

  • Technical competency managing and automating large-scale distributed systems independently of cloud providers.
  • Advanced hands-on experience with cluster management systems, including Kubernetes, Slurm, and Bright Cluster Manager.
  • Proven operational excellence maintaining reliable and performant AI infrastructure.

Compensation and Benefits

  • Level 4 base salary: USD 168,000–270,250 per year.
  • Level 5 base salary: USD 208,000–333,500 per year.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until July 10, 2026.
  • This posting is for an existing vacancy.
  • NVIDIA uses AI tools in its recruiting processes.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs