Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
at Nvidia
USD 272,000-431,200 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Algorithms @ 4
Communication @ 7
Data Structures @ 4
Deep Learning
Distributed Systems @ 6
GPU @ 4
Go @ 4
Hiring @ 4
Kubernetes @ 4
Mathematics @ 4
Python @ 4
Slurm @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is hiring experienced software engineers with Kubernetes experience to help scale its AI infrastructure. The role focuses on building and operating production systems that enable large-scale GPU clusters to support a variety of AI workloads. You will work on custom Kubernetes software for GPU resource scheduling, monitoring, health management, and distributed systems reliability.
NVIDIA develops visual computing and GPU deep learning technologies for applications including video games, movie production, product design, medical diagnosis, scientific research, and AI computing.
Responsibilities
- Work as part of the DGX Cloud team responsible for production systems supporting large-scale GPU clusters and AI workloads.
- Develop custom software related to scheduling GPU resources on Kubernetes.
- Implement monitoring and health management capabilities to improve the reliability, availability, and scalability of GPU assets.
- Work with data streams from GPU hardware diagnostics, cluster telemetry, and network telemetry.
- Collaborate with teams across NVIDIA to ensure production AI clusters operate reliably, consistently, and with maximum performance.
- Evaluate system failures and improve services through a well-defined incident management process.
Requirements
- Significant software engineering experience with Kubernetes, including cluster operations, operator development, node health monitoring, and GPU resource scheduling.
- Direct experience in a software engineering role within a highly technical organization, with demonstrable impact from your work.
- Experience developing software with Kubernetes APIs and frameworks, beyond merely operating a cluster.
- Strong communication skills and the ability to work with multifunctional teams, principals, and architects across organizational boundaries and geographies.
- 15 or more years of experience in a similar role and experience with large-scale production systems.
- Knowledge of common software engineering principles, tools, and techniques.
- A bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a comparable discipline, or equivalent experience.
- Technical knowledge of a systems programming language such as Go or Python.
- Solid understanding of data structures and algorithms.
Preferred Qualifications
- Technical competency in managing and automating large-scale distributed systems independently of cloud providers.
- Advanced hands-on experience and deep understanding of cluster management systems, including Kubernetes, Slurm, and Bright Cluster Manager.
- Proven operational excellence in maintaining reliable and performant AI infrastructure.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until October 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Software Engineer, DGX Cloud AI Infrastructure - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Tegra System Software Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Developer Tools for Cloud
Nvidia · United States
USD 152,000-287,500 per year
Data and Platform Engineer
Nvidia · United States
USD 200,000-322,000 per year
Principal Engineer, Compilers and Formal Methods
Nvidia · Seattle, United States
USD 248,000-391,000 per year
Similar jobs
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure
Nvidia · United States
USD 152,000-287,500 per year
NVIDIA 2027 Internships: Software Engineering
Nvidia · Santa Clara, United States
USD 20-71 per hour
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Systems Software Engineer – EDA Infrastructure
Nvidia · United States
USD 184,000-356,500 per year
Senior Systems Software Engineer – GeForce NOW Cloud
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year