Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
at Nvidia
USD 152,000-287,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Algorithms @ 4
Communication @ 7
Data Structures @ 4
Distributed Systems @ 6
GPU
Go @ 4
Hiring @ 4
Kubernetes @ 7
Mathematics @ 4
Python @ 4
Slurm @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is hiring experienced software engineers to help scale its AI infrastructure. The role focuses on cluster operations, operator development, node health monitoring, GPU resource scheduling, and building reliable infrastructure solutions for AI-based applications.
Responsibilities
- Join the DGX Cloud team responsible for production systems that enable large-scale GPU clusters to support a variety of AI workloads.
- Design and develop a massively distributed, scalable platform to identify, diagnose, and remediate non-performant GPU assets.
- Work with teams across NVIDIA to ensure production AI clusters run reliably and consistently with maximum performance.
- Evaluate system failures and improve services through a well-defined incident management process.
Requirements
- Direct experience in a software engineering role within a highly technical organization, with demonstrable impact from your work.
- Strong communication skills and the ability to collaborate with multifunctional teams, principals, and architects across organizational boundaries and geographies.
- 5+ years of experience in a similar role and experience with large-scale production systems.
- Experience with common software engineering principles, tools, and techniques.
- A bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a comparable discipline, or equivalent experience.
- Technical knowledge of a systems programming language such as Go or Python.
- Solid understanding of data structures and algorithms.
Preferred Qualifications
- Technical competency in managing and automating large-scale distributed systems independently of cloud providers.
- Advanced hands-on experience and deep understanding of cluster management systems, including Kubernetes, Slurm, and Base Command Manager.
- Experience with asynchronous workflows and/or event-driven architecture.
- Proven operational excellence in maintaining reliable and performant infrastructure.
Benefits
The position includes eligibility for equity and NVIDIA benefits. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
More jobs at Nvidia
Senior Localization and Planning Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Customer Success Insights Engineer
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Relational Foundation Model Engineer, Modern Data Stack
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Developer Experience
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Insider Threat Detection Engineer
Nvidia · United States
USD 168,000-310,500 per year
Similar jobs
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Nvidia · Durham, United States
USD 272,000-431,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Software Engineering Intern, Dynamo – Fall 2026
Nvidia · Santa Clara, United States
USD 20-71 per hour
Senior Full Stack Software Engineer - DGX Cloud
Nvidia · United States
USD 184,000-356,500 per year