Distinguished Engineer, Production Engineering, Cluster Management

at Nvidia
USD 320,000-488,800 per year
SENIOR
✅ Hybrid

Tech Stack

API @ 4 Distributed Systems @ 4 GPU Go @ 7 Kubernetes @ 7 Leadership @ 6 Linux @ 7 Networking @ 7 Python @ 7 SRE @ 6 Technical Leadership @ 6

Details

NVIDIA is seeking a Distinguished Engineer to serve as a senior technical leader in the Production Engineering organization. The role focuses on leading cluster activities within DGX Cloud GPU capacity. Production Engineering ensures that large-scale production systems remain reliable, manageable, and progressively automated across DGX Cloud resources by combining software engineering, systems engineering, and production expertise.

The position focuses on the operational structure for DGX Cloud clusters across on-premises environments, hyperscalers, and NVIDIA Cloud Partner environments. The scope includes Kubernetes service management, provider and hardware readiness, on-premises infrastructure handling, deployment and operational preparedness, service reliability collaboration, and the processes that integrate these areas into a unified production system.

This is a hands-on Distinguished Engineer role for a deeply technical leader who will define architectural direction for cluster operations throughout DGX Cloud, set technical strategy and operating standards, guide the evolution of production operations, and drive cross-organizational capabilities that keep DGX Cloud capacity usable, supportable, and improving at scale.

Responsibilities

  • Define the long-range technical strategy for operating DGX Cloud clusters consistently across local data centers, hyperscalers, and NeoCloud environments.
  • Define architectural direction and operating standards for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability.
  • Guide the roadmap and execution of cross-organizational investments that improve production readiness, operational safety, performance, and coordination.
  • Make and influence high-impact technical decisions across platform, hardware, provider, and service teams operating DGX Cloud resources in production.
  • Build durable workflows, interfaces, and engineering handshakes across Kubernetes production services, provider and hardware preparation, on-premises infrastructure, and bare-metal environments.
  • Restructure operations and service-layer reliability domains.
  • Build and evolve automation, APIs, operating workflows, and readiness gates required to move new capacity into stable production and keep existing capacity balanced.
  • Implement production operating approaches that reduce manual effort, establish clear ownership, and increase consistency, traceability, and release safety.
  • Partner with platform teams, hardware and provider engineering, service owners, and Production Engineering leaders to identify recurring friction and convert it into durable improvements in software, processes, and operational interfaces.
  • Raise engineering standards for operability, resilience, scalability, and performance through build leadership, architecture reviews, and technical standards.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience.
  • 18 or more years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments.
  • Company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms.
  • A consistent track record of defining operating models, architectural direction, and engineering standards across multiple technical domains and organizations.
  • Experience leading large, cross-team technical efforts from concept through production, including aligning collaborators, creating clarity, and delivering measurable outcomes.
  • Deep experience in one or more of the following areas: Kubernetes-based production systems, infrastructure automation, or distributed systems operations.
  • Strong software engineering skills in languages such as Python, Go, or similar low-level programming languages.
  • Deep understanding of distributed systems, Linux, networking, containers, and production reliability concerns.
  • Experience creating operational workflows, APIs, service interfaces, or automation frameworks that become standard ways for teams to run production systems.
  • Strong architectural judgment and a demonstrated history of simplifying complex operational problems through reusable software, clear technical strategy, and durable engineering direction.

Preferred Experience

  • Defining the structural foundation for a large, heterogeneous infrastructure environment spanning multiple platforms or providers.
  • Establishing widely used operating standards, architectures, APIs, or workflows that improve reliability, operability, or performance at company scale.
  • Building automation and engineering interfaces that connect platform teams, infrastructure teams, and service owners into a consistent production system.
  • Improving production readiness, restoration, runtime safety, or release quality for large-scale infrastructure.
  • Combining technical strategy, large-scale architecture review, coding, and adoption leadership across organizational boundaries.

Benefits

The base salary range is USD 320,000 to USD 488,750, determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits. Applications will be accepted at least until September 7, 2026.

NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs