Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 4
CI/CD @ 4
Distributed Systems @ 7
Go @ 6
IaC
Kubernetes @ 7
Networking @ 7
Observability
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. The team deploys, integrates, and operates the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA’s global network across data centers, colocation facilities, and cloud environments. CFR owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. The team builds software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments.
This is a hands-on senior individual contributor role responsible for the lifecycle and automation of the Kubernetes platform supporting GNI network systems. The role also provides production support for network services running on the platform and involves partnering with engineering owners when issues or changes cross the platform boundary. The engineer will take complex problems from design through production and remain accountable for the outcome, while helping establish consistent engineering practices across the US and Bangalore teams.
Responsibilities
- Design, build, and operate the Kubernetes platform powering GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.
- Own the lifecycle management of GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.
- Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.
- Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.
- Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.
- Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.
- Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
- 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
- Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.
- Proficiency in at least one general-purpose programming language, such as Go or Python.
- Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.
- Experience deploying and supporting network automation or telemetry services on Kubernetes.
- Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.
Preferred Qualifications
- Strong knowledge of IP routing, data center fabrics, and cloud networking.
- Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery.
- Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades.
- Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns.
- Experience designing or operating network automation and telemetry services on Kubernetes at global scale.
- Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects.
Compensation and Additional Information
The base salary range is USD 176,000–276,000 for Level 4 and USD 208,000–333,500 for Level 5. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until August 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.