Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
at Nvidia
USD 272,000-431,200 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Algorithms @ 4
Communication @ 7
Data Structures @ 4
Distributed Systems @ 7
GPU
Go @ 4
Hiring @ 4
Kubernetes @ 4
Mathematics @ 4
Python @ 4
Slurm @ 7
Software Development @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is hiring experienced software engineers with Kubernetes experience to help scale up its AI Infrastructure. You will help advance NVIDIA's capacity to build and deploy leading infrastructure solutions for a broad range of AI-based applications.
Responsibilities
- Be part of an DGX Cloud team responsible for production systems that enable large scalable GPU clusters to be used for a variety of AI workloads, including working on custom software related to scheduling GPU resources on Kubernetes.
- Implement monitoring and health management capabilities that enable reliability, availability, and scalability of GPU assets by harnessing multiple data streams (GPU hardware diagnostics, cluster and network telemetry).
- Work with teams across NVIDIA to ensure production AI clusters run reliably and consistently with maximum performance; evaluate system failures and improve services via a well-defined incident management process.
Requirements
- Direct experience in a software engineering role within a highly technical organization with demonstrable impact; software development experience with Kubernetes APIs and frameworks (not just operating a cluster).
- Strong communication skills; ability to work successfully with multi-functional teams, principles and architects, and coordinate across organizational boundaries and geographies.
- 15+ years in similar role and experience on large-scale production systems; experience with common software engineering principles, tools, and techniques.
- BS in Computer Science, Engineering, Physics, Mathematics, or comparable degree/equivalent experience.
- Technical knowledge including a systems programming language (Go, Python) and a solid understanding of data structures and algorithms.
Ways to Stand Out
- Technical competency in managing and automating large-scale distributed systems independent of cloud providers; advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, Slurm, Bright Cluster Manager).
- Proven operational excellence in maintaining reliable and performant AI infrastructure.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.
Applications for this job will be accepted at least until June 29, 2026.
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior GPU and HPC Infrastructure Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Software Engineer - Analytics Platform
Bloomberg · New York City, United States
USD 160,000-240,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Senior Storage Software Engineer, DGXC Data Services
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Ai Infrastructure Engineer - Dgx Cloud
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, GoLang - DSX MaxQ
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer - HPC
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year