Principal Software Engineer, DGX Cloud Production Engineering

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API @ 6 Distributed Systems @ 8 GPU @ 4 Go @ 7 Kubernetes @ 7 Linux @ 7 Machine Learning Networking @ 6 Observability @ 4 Python @ 7 Security @ 6

Details

NVIDIA DGX Cloud is scaling GPU infrastructure across internal, partner, and cloud environments. This role will help shape the technical direction for production engineering, Kubernetes-based operations, automation, and reliability across large-scale GPU clusters.

The position is for a senior technical leader who can define architecture, lead through influence, build critical systems, and turn ambiguous infrastructure problems into durable software and operating models.

Responsibilities

  • Define and execute the technical strategy for DGX Cloud cluster operations, including automation, GitOps, and Day 2 reliability for large-scale GPU clusters across NVIDIA Cloud Partners (NCPs) and on-premises environments.
  • Lead the design and implementation of systems for cluster lifecycle management, validation, repair, upgrades, observability, and readiness.
  • Establish patterns for Kubernetes-based GPU cluster operations across partner and on-premises environments.
  • Identify and eliminate operational toil through software, APIs, automation, and agent-assisted workflows.
  • Set technical standards for production readiness, SLOs, incident response, handoff gates, and operational acceptance.
  • Mentor engineers and influence platform, infrastructure, storage, networking, security, and workload teams.

Requirements

  • 15+ years of experience building and operating large-scale distributed systems or cloud infrastructure.
  • Deep experience with Kubernetes, Linux, infrastructure automation, and production operations.
  • Strong programming experience in Go, Python, or similar languages.
  • Proven ability to lead complex cross-organizational technical initiatives.
  • Experience designing reliable systems with clear SLOs, observability, incident response, and automation.
  • Bachelor’s or master’s degree in Computer Science, or equivalent experience.

Preferred Qualifications

  • Experience with GPU clusters, AI/ML infrastructure, Kubernetes operators, GitOps, BMaaS/VMaaS, managed Kubernetes, or multi-cloud fleet operations.
  • Experience building internal platforms, control planes, lifecycle automation, or production readiness frameworks.
  • Track record of turning operational pain into reusable software, APIs, and engineering standards.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.

More jobs at Nvidia

Similar jobs