Senior Software Engineer, Kubernetes Runtime and Release

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ Remote

Tech Stack

AI API Distributed Systems @ 7 GPU @ 6 Go @ 7 Helm Kubernetes @ 4 Networking @ 6

Details

NVIDIA researchers depend on GPU clusters for large-scale AI workloads. The DGX Cloud Kubernetes Runtime & Release team builds and maintains the supported Kubernetes runtime, automates its delivery, and brings new providers and GPU platforms into production across major public clouds and specialized GPU providers.

The team is expanding its ownership of NVIDIA’s cluster software delivery across Runtime, Release Engineering, and Provider Integration. Each role focuses on the candidate’s strengths, and experience across every area is not required.

Responsibilities

The primary focus will be one of the following areas, with collaboration across the team:

  • Runtime: Build Go controllers and APIs to install, upgrade, and validate GPU cluster software. Integrate components, define API contracts, and evolve Helm and Argo CD delivery toward controller-driven automation.
  • Release Engineering: Build validation pipelines that inform release decisions across providers and GPU platforms. Develop systems to allocate GPU capacity across validation runs and account for cloud reservations and quotas. Improve qualification efficiency through reusable tests and clear failure reports.
  • Provider Integration: Bring new providers and GPU hardware into production, potentially among the first engineers working with new silicon. Resolve integration failures with partner teams and turn initial provisioning, upgrade, and operational checks into repeatable automation.

Requirements

  • 6+ years of experience building production infrastructure software or distributed systems.
  • Strong programming skills in Go or another language for building production systems, with willingness to work primarily in Go.
  • Kubernetes experience and depth in at least one area: controllers and operators, release automation, test and validation systems, or cloud integration.
  • Experience delivering engineering projects, diagnosing complex failures, and collaborating across teams.
  • BS or MS in Computer Science, Engineering, or equivalent experience.

Preferred Qualifications

  • Go development with controller-runtime, CRDs, and reconcilers.
  • Release qualification across multiple environments or platforms.
  • GPU infrastructure, accelerated networking, or GPU scheduling.
  • Bringing new hardware, regions, or cloud providers into production.
  • Resource allocation, leasing, or fair-share scheduling.
  • Upstream integration, compatibility, or software supply chain integrity.

Benefits

The role includes eligibility for equity and benefits. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. Applications will be accepted at least until October 3, 2026.

More jobs at Nvidia

Similar jobs