Senior Software Engineer, DGX Cloud Production Engineering

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API @ 4 Communication @ 6 Distributed Systems @ 7 GPU @ 4 Go @ 7 HPC @ 4 Kubernetes @ 4 Leadership @ 6 Networking Python @ 7 Rust @ 7 Technical Leadership @ 6

Details

NVIDIA is seeking a Senior Software Engineer to build the next generation of its Kubernetes platform. The team develops foundational capabilities for self-service GPU infrastructure, managed Kubernetes control planes, cluster operations, and automation for large-scale AI and improved computational environments.

The role operates at the intersection of platform architecture and production-grade Kubernetes lifecycle systems, contributing to the technical strategy and execution of systems that provision, upgrade, operate, and scale clusters across cloud and on-premises environments.

Responsibilities

  • Collaborate on the architecture and development of core Kubernetes platform capabilities, including cluster management, control plane services, fleet lifecycle, and day-2 operations.
  • Design and build highly reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale.
  • Define technical requirements, validation criteria, production-readiness practices, and the direction for declarative workflows and automation across the Kubernetes stack.
  • Collaborate across engineering teams to create cohesive platform experiences spanning management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
  • Diagnose and resolve complex platform issues spanning infrastructure, runtime, networking, hardware, and operations.
  • Improve the scalability, resilience, and operability of systems supporting large-scale AI deployments.

Requirements

  • BS or MS degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 5+ years of relevant software engineering experience, including experience building and operating large-scale production systems.
  • Deep expertise in Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management.
  • Strong background in distributed systems design, reliability, scalability, and failure recovery.
  • Proven experience building platform software, infrastructure control planes, or foundations for managed services.
  • Strong programming skills in one or more systems or cloud-native languages, such as Go, Python, Rust, or C++.
  • Experience designing clear APIs and abstractions for platform consumers and engineering teams.
  • Demonstrated ability to provide technical leadership across team boundaries and drive ambiguous, cross-functional initiatives to completion.
  • Excellent communication and collaboration skills, supported by significant technical contributions and recognized expertise influencing department-level architecture and high-priority company initiatives.

Preferred Qualifications

  • Experience building Kubernetes platforms or managed Kubernetes services.
  • Experience with fleet management, cluster upgrades, node lifecycle, remediation, or day-2 operations.
  • Experience with declarative infrastructure, Kubernetes controllers, GitOps, or policy-driven platform automation.
  • Familiarity with public-cloud and bare-metal infrastructure environments.
  • Experience supporting AI, GPU, HPC, or other large-scale accelerated computing platforms.

Compensation and Benefits

  • Base salary range: USD 152,000–241,500 for Level 3.
  • Base salary range: USD 184,000–287,500 for Level 4.
  • Compensation is determined based on location, experience, and the pay of employees in similar positions.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until August 22, 2026.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs