Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
Algorithms @ 4
CI/CD @ 7
Cloud Computing @ 7
Data Structures @ 4
Distributed Systems @ 4
GPU @ 7
Go @ 4
Kubernetes @ 7
Linux @ 4
Security @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's DSX Kubernetes Fleet team within the DGX Cloud organization builds automation, lifecycle management, and deployment safety for a large-scale, GPU-accelerated container platform. The team develops APIs and workflows that turn high-level deployment intent into production-ready AI infrastructure management and owns the full lifecycle of AI Factory Kubernetes clusters, from provisioning through upgrades and decommissioning.
Responsibilities
- Explore innovative ways to simplify the development, deployment, and monitoring of GPU- and DPU-accelerated applications.
- Design and develop software for managing fleets of Kubernetes clusters for GPUs and DPUs.
- Work with Cloud Native technologies to improve NVIDIA accelerators in Kubernetes environments.
- Collaborate with engineering teams across NVIDIA to ensure seamless software integration.
- Automate and optimize build, test, integration, and release processes for cloud-native applications.
- Multitask across different projects and address evolving priorities effectively.
Requirements
- Bachelor's or master's degree in Computer Science or a related field, or equivalent experience.
- 10+ years of proven work experience in large-scale environments.
- Expert-level knowledge of systems programming languages, including Go and C, with a solid understanding of data structures and algorithms.
- Strong understanding of container orchestration systems, particularly Kubernetes, and container technology.
- In-depth knowledge of Unix/Unix-like kernel internals, particularly Linux.
- Hands-on automation experience with modern infrastructure tools and technologies.
- Proven experience setting up, maintaining, and automating continuous deployment systems.
- Strong background in cloud computing and distributed software design and development.
- Understanding of performance, security, and reliability in complex distributed systems.
Preferred Qualifications
- Extensive experience with Go.
- Deep understanding of rack-scale GPU systems.
- Strong background with GitLab, Argo, Flux, and other CI/CD systems.
- Significant hands-on experience with containers and Kubernetes.
- Hands-on experience with container workload isolation and confidential computing.
Benefits
NVIDIA offers a comprehensive benefits package, equity, and competitive compensation. The base salary depends on location, experience, and compensation for employees in similar positions. Applications will be accepted at least until August 11, 2026. This posting is for an existing vacancy.
More jobs at Nvidia
Senior Software Engineer, Unified Access Management Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Technical Program Manager, AV System Integration
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
Senior System Software Engineer, Automotive Performance
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Embedded System Software Engineer – Platform Execution Lead
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Principal Software Engineer - DGX Cloud
Nvidia · Seattle, United States
USD 272,000-431,200 per year
Senior Storage Software Engineer, DGXC Data Services
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Cloud Software Engineer, DGXC Data Services
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year