Principal Software Engineer, DGX Cloud Production Engineering
at Nvidia
USD 272,000-431,200 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 6
Distributed Systems @ 8
GPU @ 4
Go @ 7
Kubernetes @ 7
Linux @ 7
Machine Learning
Networking @ 6
Observability @ 4
Python @ 7
Security @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA DGX Cloud is scaling GPU infrastructure across internal, partner, and cloud environments. This role will help shape the technical direction for production engineering, Kubernetes-based operations, automation, and reliability across large-scale GPU clusters.
The position is for a senior technical leader who can define architecture, lead through influence, build critical systems, and turn ambiguous infrastructure problems into durable software and operating models.
Responsibilities
- Define and execute the technical strategy for DGX Cloud cluster operations, including automation, GitOps, and Day 2 reliability for large-scale GPU clusters across NVIDIA Cloud Partners (NCPs) and on-premises environments.
- Lead the design and implementation of systems for cluster lifecycle management, validation, repair, upgrades, observability, and readiness.
- Establish patterns for Kubernetes-based GPU cluster operations across partner and on-premises environments.
- Identify and eliminate operational toil through software, APIs, automation, and agent-assisted workflows.
- Set technical standards for production readiness, SLOs, incident response, handoff gates, and operational acceptance.
- Mentor engineers and influence platform, infrastructure, storage, networking, security, and workload teams.
Requirements
- 15+ years of experience building and operating large-scale distributed systems or cloud infrastructure.
- Deep experience with Kubernetes, Linux, infrastructure automation, and production operations.
- Strong programming experience in Go, Python, or similar languages.
- Proven ability to lead complex cross-organizational technical initiatives.
- Experience designing reliable systems with clear SLOs, observability, incident response, and automation.
- Bachelor’s or master’s degree in Computer Science, or equivalent experience.
Preferred Qualifications
- Experience with GPU clusters, AI/ML infrastructure, Kubernetes operators, GitOps, BMaaS/VMaaS, managed Kubernetes, or multi-cloud fleet operations.
- Experience building internal platforms, control planes, lifecycle automation, or production readiness frameworks.
- Track record of turning operational pain into reusable software, APIs, and engineering standards.
Benefits
- Equity and benefits are provided.
- NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.
More jobs at Nvidia
Senior Software Engineer, Unified Access Management Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Technical Program Manager, AV System Integration
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
Senior System Software Engineer, Automotive Performance
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Embedded System Software Engineer – Platform Execution Lead
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year