Senior Manager, Kubernetes Runtime Engineering

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI API @ 7 Compliance @ 4 GPU @ 4 Kubernetes @ 7 NVLink @ 4 Networking @ 4 Security @ 4

Details

The NVIDIA Kubernetes Engine (NKE) team is looking for a technical leader to lead the Runtime Engineering team responsible for the full configuration lifecycle of NKE tenant workload clusters. The team develops software components that keep GPU workloads reliable and secure at scale, including cluster bootstrapping, node configuration, and the container execution environment, including NVIDIA's AI Container Runtime (AICR). The role covers networking, storage, GPU resource management, and cluster security for a production-grade, multi-tenant Kubernetes platform.

Responsibilities

  • Own the build, implementation, and operational reliability of cluster configurations for NKE tenant workloads across all supported topologies.
  • Manage a team of engineers coordinating the container runtime stack, including AICR, GPU Management Operator, DCGM, and related node-level components.
  • Drive architecture decisions for cluster networking (CNI), storage (CSI), cluster high availability, and GPU resource partitioning, including MIG, MPS, and time-slicing.
  • Define and implement cluster hardening standards, RBAC models, pod security policies, and multi-tenancy isolation boundaries.
  • Partner with NKE platform, infrastructure, and cybersecurity teams to integrate new capabilities and resolve cross-cutting runtime concerns.
  • Build and maintain tooling for AICR lifecycle management, including provisioning, upgrades, configuration drift detection, and remediation.
  • Represent the runtime team in architecture reviews, roadmap planning, and customer communications with NVIDIA leadership.
  • Contribute to open-source communities wherever NKE has upstream dependencies or influence.

Requirements

  • Bachelor's or master's degree in Computer Science or a related field, or equivalent experience.
  • 12 or more years of relevant experience designing and delivering large-scale distributed software systems.
  • 5 or more years of people-management experience leading, developing, and scaling high-performing software engineering teams responsible for complex, production-critical software.
  • Experience leading engineers with varying specializations and seniority levels, including runtime, networking, and security disciplines.
  • Deep knowledge of Kubernetes internals, including interactions between the scheduler, kubelet, API server, and admission controllers.
  • Experience with cluster lifecycle management using Cluster API, kubeadm, or equivalent tools, including fleet-scale cluster provisioning and upgrades.
  • Experience with security and compliance practices, including the CIS Kubernetes Benchmark, Pod Security Admission, image signing, and software supply chain integrity.
  • Proven ability to design and implement maintainable APIs for consumers.
  • Familiarity with Identity and Access Management approaches.
  • Ability to manage up, down, and across organizations and reach cross-organizational consensus in ambiguous situations.

Preferred Qualifications

  • Experience with NVIDIA GPU Operator, DCGM Exporter, or NVLink-aware scheduling.
  • Experience running Kubernetes at hyperscale with GPU node pools.
  • Track record of upstream open-source contributions in Kubernetes or another open-source runtime ecosystem.
  • Strong written, verbal, interpersonal, coaching, analytical, problem-solving, and technical-planning skills.

Benefits

NVIDIA offers competitive salaries, a comprehensive benefits package, equity, and benefits for employees and their families. NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs