Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
ArgoCD @ 4
Communication @ 6
Debugging
Distributed Systems @ 6
GPU @ 4
Go @ 7
Kubernetes @ 4
Linux @ 4
Networking
Observability @ 4
Python @ 7
Security
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI research and production workloads. The production engineering team focuses on Kubernetes-based infrastructure, GPU cluster operations, reliability, automation, GitOps, and Day 2 operability across DGX Cloud environments.
Responsibilities
- Build and operate automation for large-scale GPU clusters across NVIDIA Cloud Partners and on-premises environments.
- Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations.
- Improve Day 0, Day 1, and Day 2 workflows for cluster bring-up, handoff, and production operations.
- Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows.
- Participate in on-call rotations, incident response, debugging, and durable follow-up work.
- Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready.
Requirements
- 8+ years of experience building or operating production infrastructure.
- Strong programming skills in Python, Go, or similar languages.
- Experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
- Ability to troubleshoot distributed systems in production.
- Clear communication skills and the ability to work across teams.
- Bachelor's or master's degree in Computer Science, or equivalent experience.
Preferred Qualifications
- Experience with GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation.
- Experience with SLOs, on-call operations, incident response, observability, and reliability practices.
- Exposure to bare-metal-as-a-service, virtual-machine-as-a-service, managed Kubernetes, or multi-cloud infrastructure.
Compensation and Benefits
The base salary range is $184,000–$287,500 for Level 4 and $224,000–$356,500 for Level 5. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until September 27, 2026. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
More jobs at Nvidia
Senior Data Backend Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Engineer, Performance
Nvidia · Santa Clara, United States
USD 136,000-212,800 per year
Senior System Software Engineer, Robotics
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - GPU Power and Performance Management
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
PhD Research Intern, Electronic Design Automation - 2027
Nvidia · Santa Clara, United States
USD 38-94 per hour
Similar jobs
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Principal Software Engineer – Infrastructure
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year