Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Algorithms @ 4
Communication @ 7
Data Structures @ 4
Deep Learning
Distributed Systems @ 6
GPU
Go @ 4
Hiring @ 4
Kubernetes @ 4
Mathematics @ 4
Python @ 4
Slurm @ 7
Software Development @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is hiring experienced software engineers with Kubernetes experience to help scale its AI infrastructure. The role focuses on building and operating production systems for large-scale GPU clusters and AI workloads. You will work on custom software for Kubernetes GPU resource scheduling, cluster health management, monitoring, reliability, availability, and scalability.
NVIDIA is advancing infrastructure solutions for AI-based applications using GPU computing and deep learning.
Responsibilities
- Work as part of the DGX Cloud team responsible for production systems that enable large-scale GPU clusters to support a variety of AI workloads.
- Develop custom software related to scheduling GPU resources on Kubernetes.
- Implement monitoring and health management capabilities for reliable, available, and scalable GPU assets.
- Analyze data streams from GPU hardware diagnostics, cluster telemetry, and network telemetry.
- Collaborate with teams across NVIDIA to ensure production AI clusters operate reliably, consistently, and with maximum performance.
- Evaluate system failures and improve services through a well-defined incident management process.
Requirements
- Direct software engineering experience in a highly technical organization, with demonstrable impact from previous work.
- Software development experience with Kubernetes APIs and frameworks, beyond operating a cluster.
- Strong communication skills and the ability to work with multifunctional teams, principals, and architects across organizational boundaries and geographies.
- 15 or more years of experience in similar roles and experience with large-scale production systems.
- Knowledge of common software engineering principles, tools, and techniques.
- A bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a comparable discipline, or equivalent experience.
- Technical knowledge of a systems programming language such as Go or Python.
- Solid understanding of data structures and algorithms.
Preferred Qualifications
- Technical competency in managing and automating large-scale distributed systems independently of cloud providers.
- Advanced hands-on experience and deep understanding of cluster management systems, including Kubernetes, Slurm, and Bright Cluster Manager.
- Proven operational excellence maintaining reliable and performant AI infrastructure.
Benefits
The role includes eligibility for equity and benefits. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until June 29, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.