Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
Algorithms @ 4
CI/CD @ 4
Cloud Computing @ 7
Data Structures @ 4
Distributed Systems @ 4
GPU @ 7
Go @ 4
Kubernetes @ 7
Linux @ 4
Security @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Join NVIDIA's DGX Cloud organization as a Principal Software Engineer on the DSX Kubernetes Fleet team. The team builds automation, lifecycle management, and deployment safety for a large-scale, GPU-accelerated container platform that supports AI teams training and deploying at AI Factory scale, with resiliency and security as top priorities.
The role focuses on designing and building APIs and workflows that convert high-level deployment intent into production-ready AI infrastructure management. You will own software across the full lifecycle of AI Factory Kubernetes clusters, from provisioning and upgrades through decommissioning.
Responsibilities
- Simplify the development, deployment, and monitoring of GPU- and DPU-accelerated applications.
- Design and develop software for managing fleets of Kubernetes clusters for GPUs and DPUs.
- Work with cloud-native technologies to improve NVIDIA accelerators in Kubernetes environments.
- Collaborate with engineering teams across NVIDIA to ensure seamless software integration.
- Automate and optimize build, testing, integration, and release processes for cloud-native applications.
- Multitask across projects and address evolving priorities effectively.
Requirements
- BS or MS in Computer Science, a related field, or equivalent experience.
- 15 or more years of professional experience in large-scale environments.
- Expert-level knowledge of systems programming languages, including Go and C.
- Solid understanding of data structures and algorithms.
- Strong knowledge of Kubernetes and container technologies.
- In-depth experience with Unix or Unix-like kernel internals, particularly Linux.
- Hands-on experience with modern infrastructure automation tools and technologies.
- Proven experience setting up, maintaining, and automating continuous deployment systems.
- Strong background in cloud computing and distributed software design and development.
- Understanding of performance, security, and reliability in complex distributed systems.
Preferred Qualifications
- Extensive experience with Go.
- Deep understanding of rack-scale GPU systems.
- Experience with GitLab, Argo, Flux, and other CI/CD systems.
- Significant hands-on experience with containers and Kubernetes.
- Experience with container workload isolation and confidential computing.
Compensation And Benefits
The base salary range is USD 272,000 to USD 431,250 per year, based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until August 11, 2026. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.