Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 7
AWS @ 4
Azure @ 4
CI/CD
Communication @ 6
Compliance @ 4
Debugging @ 4
GCP @ 4
GPU @ 4
Go @ 6
Kubernetes @ 7
Linux @ 4
Python @ 6
Security @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's DGX Cloud division combines hardware and software innovations to provide accelerated computing solutions for challenging AI workloads worldwide.
The team is seeking a Senior Systems Software Engineer with strong experience in Kubernetes node engineering, OS image packaging, and cloud infrastructure. The role requires deep hyperscaler-level knowledge across the complete node lifecycle, including Cluster API (CAPI) providers, bring-your-own-node onboarding, OS image build pipelines, packaging, and nodepool management. The engineer will manage the node layer within NVIDIA Kubernetes Engine (NKE), supporting internal researchers and NVIDIA Cloud Partners (NCPs) at frontier AI scale.
Responsibilities
- Build and refine CAPI providers for NVIDIA Kubernetes Engine to maintain consistent and scalable node provisioning across DGX Cloud and NCP environments.
- Develop and maintain bring-your-own-node workflows for integrating different NVIDIA hardware into NKE clusters.
- Coordinate OS image generation, packaging, deployment, and update processes for NKE nodes, ensuring images are optimized for NVIDIA GPU workloads and meet enterprise- and cloud-grade security and compliance requirements.
- Develop and maintain node image hardening pipelines incorporating CIS benchmarks, automated CVE remediation, and security-posture promotion gates.
- Develop and maintain automated test suites for node images across Kubernetes versions and NVIDIA hardware configurations, with continuous validation through CI/CD pipelines.
- Manage nodepool lifecycles at scale, including provisioning, upgrades, drain and cordon workflows, and seamless node replacement across large clusters with diverse NVIDIA hardware.
- Investigate and resolve production node-layer issues in NKE clusters, including image configuration, driver packaging, kubelet operation, and hardware activation problems.
- Review and optimize the node layer in high-scale production environments.
- Partner with upstream communities, including Cluster API, Kubernetes, and CNCF projects, to establish node provisioning and lifecycle standards.
- Communicate progress and findings at internal and external events such as KubeCon and GTC.
Requirements
- 8 years of experience in systems software, cloud infrastructure, or Kubernetes node engineering.
- Bachelor's or Master's degree in Engineering, Electrical Engineering, Computer Engineering, Computer Science, or equivalent experience.
- Deep expertise in Cluster API (CAPI), including provider development and the full machine lifecycle from provisioning through deletion.
- Extensive experience with OS image build pipelines, node image packaging, and delivery systems for Kubernetes nodes, such as image-builder, containerd, cloud-init, and Packer.
- Practical experience with bring-your-own-node models and integrating diverse hardware into live Kubernetes environments.
- Experience with large-scale nodepool lifecycle management and upgrades.
- Strong understanding of kubelet configuration, node bootstrap, and the Kubernetes node registration lifecycle.
- Experience with node image security, including vulnerability scanning, patch automation, and compliance gating in image build pipelines.
- Proficiency in Go and/or Python.
- Hands-on experience with at least one major public cloud provider, such as GCP, AWS, Azure, OCI, or equivalent.
Preferred Qualifications
- Experience building or maintaining node image pipelines for a hyperscaler Kubernetes distribution, such as GKE, EKS, AKS, OKE, or equivalent.
- Experience with supply chain security and node image hardening, including image signing, provenance attestation, SBOM generation, CIS benchmark compliance, and automated CVE remediation.
- Experience with automated node provisioning and optimal sizing at scale, such as Karpenter or GKE NAP, and understanding of how these systems interact with GPU workload scheduling.
- Operational experience with immutable OS image distributions, such as Flatcar, Bottlerocket, or Azure Linux.
- Experience debugging node-layer failures in large Kubernetes clusters.
- Upstream contributions to Cluster API, Kubernetes, or related CNCF projects.
- Excellent communication and interpersonal skills.
Compensation And Benefits
The base salary depends on location, experience, and the pay of employees in similar positions. The base salary ranges are:
- Level 4: USD 184,000–287,500 per year
- Level 5: USD 224,000–356,500 per year
The role is also eligible for equity and benefits. Applications will be accepted at least until June 14, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.