Engineering Manager, Kubernetes Customer Delivery and Self-Service
at Nvidia
USD 224,000-356,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Communication @ 7
Distributed Systems @ 7
GPU @ 4
Kubernetes @ 7
Leadership @ 7
Mentoring
Networking @ 4
Prioritization @ 7
Security
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA’s DGX Cloud Kubernetes Platform & Production Engineering team is seeking an Engineering Manager to develop and guide the Customer Delivery and Self-Service function. This group manages the engineering delivery process from an approved customer request through platform enablement, cluster build, validation, and handoff. The leader will transform the current cross-team process into a scalable, automated, self-service solution.
Responsibilities
- Build and lead a team of software and production engineers focused on Kubernetes customer delivery, onboarding, and self-service.
- Own end-to-end delivery of production Kubernetes clusters for AI workloads, from accepted request through enablement, qualification, validation, and customer handoff.
- Drive coordinated delivery plans with clear owners, dependencies, readiness gates, timelines, risks, status, and blocking issues.
- Partner across platform, runtime, release, fleet operations, CSE, product, TPM, security, and infrastructure teams.
- Build integrations connecting customer intake and status systems with Kubernetes provisioning, access, validation, and production acceptance.
- Turn recurring delivery tasks into detailed, automated self-service workflows using APIs, AI tools, and agents.
- Define service interfaces and measure and improve delivery speed, readiness, automation, recovery, and customer visibility.
- Set the team’s roadmap, staffing, and operational ownership while hiring, mentoring, and developing technical leaders.
Requirements
- 8+ years of overall industry experience, including 2+ years leading or managing engineers.
- Experience building platform APIs, self-service infrastructure, workflow automation, developer platforms, or customer onboarding systems.
- Strong understanding of Kubernetes, cloud infrastructure, distributed systems, or production engineering.
- Hands-on experience using AI coding tools and AI-enabled engineering workflows.
- Experience integrating multiple systems and teams into a reliable end-to-end workflow.
- Ability to translate customer and operational requirements into clear technical interfaces and automated solutions.
- Strong cross-functional leadership, communication, customer empathy, prioritization, and judgment.
- Bachelor’s or master’s degree in Computer Science, Engineering, or equivalent experience.
Preferred Qualifications
- Experience building Kubernetes provisioning, infrastructure-as-code, service catalog, or internal developer platform capabilities.
- Familiarity with Terraform, GitOps, identity and access management, RBAC, APIs, workflow engines, and production-readiness automation.
- Experience with GPU infrastructure and AI-optimized Kubernetes clusters, including accelerated networking, high-performance storage, GPU scheduling, workload qualification, or large-scale fleet operations.
- Track record of reducing onboarding time and operational toil through automation and self-service.
- Experience combining strong platform engineering with a product development and customer-focused approach.
Compensation and Benefits
- Base salary range: USD 224,000–356,500 per year, determined by location, experience, and the pay of employees in similar positions.
- Eligible for equity and benefits.
- Full-time position.
- Applications will be accepted at least until August 15, 2026.
NVIDIA uses AI tools in its recruiting processes and is committed to fostering an inclusive work environment and providing equal employment opportunities.
More jobs at Nvidia
Senior System Software Engineer, Platform - OpenBMC
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Security Engineer, RTOS and Virtualization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer - NVIDIA Warp
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software And System Architect
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Similar jobs
Senior Security Engineer, AI Security
Reddit · United States
USD 190,800-267,100 per year
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Product Manager, Cloud Networking
Confluent · United States
USD 231,500-272,000 per year
Senior Software Engineer, Attestation Services – DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Principal Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Staff AI Platform Engineer, Infrastructure Services
SentinelOne · United States
USD 156,000-215,000 per year
Senior AI Platform Engineer, Infrastructure Services
SentinelOne · United States
USD 132,000-182,000 per year