Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Communication @ 6
Distributed Systems @ 7
GPU @ 4
Go @ 7
HPC @ 4
Kubernetes @ 4
Leadership @ 6
Networking
Python @ 7
Rust @ 7
Technical Leadership @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Software Engineer to build the next generation of its Kubernetes platform. The team develops foundational capabilities for self-service GPU infrastructure, managed Kubernetes control planes, cluster operations, and automation for large-scale AI and improved computational environments.
The role operates at the intersection of platform architecture and production-grade Kubernetes lifecycle systems, contributing to the technical strategy and execution of systems that provision, upgrade, operate, and scale clusters across cloud and on-premises environments.
Responsibilities
- Collaborate on the architecture and development of core Kubernetes platform capabilities, including cluster management, control plane services, fleet lifecycle, and day-2 operations.
- Design and build highly reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale.
- Define technical requirements, validation criteria, production-readiness practices, and the direction for declarative workflows and automation across the Kubernetes stack.
- Collaborate across engineering teams to create cohesive platform experiences spanning management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
- Diagnose and resolve complex platform issues spanning infrastructure, runtime, networking, hardware, and operations.
- Improve the scalability, resilience, and operability of systems supporting large-scale AI deployments.
Requirements
- BS or MS degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 5+ years of relevant software engineering experience, including experience building and operating large-scale production systems.
- Deep expertise in Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management.
- Strong background in distributed systems design, reliability, scalability, and failure recovery.
- Proven experience building platform software, infrastructure control planes, or foundations for managed services.
- Strong programming skills in one or more systems or cloud-native languages, such as Go, Python, Rust, or C++.
- Experience designing clear APIs and abstractions for platform consumers and engineering teams.
- Demonstrated ability to provide technical leadership across team boundaries and drive ambiguous, cross-functional initiatives to completion.
- Excellent communication and collaboration skills, supported by significant technical contributions and recognized expertise influencing department-level architecture and high-priority company initiatives.
Preferred Qualifications
- Experience building Kubernetes platforms or managed Kubernetes services.
- Experience with fleet management, cluster upgrades, node lifecycle, remediation, or day-2 operations.
- Experience with declarative infrastructure, Kubernetes controllers, GitOps, or policy-driven platform automation.
- Familiarity with public-cloud and bare-metal infrastructure environments.
- Experience supporting AI, GPU, HPC, or other large-scale accelerated computing platforms.
Compensation and Benefits
- Base salary range: USD 152,000–241,500 for Level 3.
- Base salary range: USD 184,000–287,500 for Level 4.
- Compensation is determined based on location, experience, and the pay of employees in similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until August 22, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior DevTech Compute Engineer, Compression and Data Processing
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Senior DFX Software Engineer - Machine Learning
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Technical Program Manager, AI Infrastructure and Capacity Operations
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Program Manager – Chip System Software
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Similar jobs
Principal Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Principal Engineer, AI Tooling and Workflows
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, Attestation Services – DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · London, United Kingdom
GBP 325,000-485,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year