Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Algorithms @ 4
Communication @ 7
Data Structures @ 4
DevOps @ 4
Distributed Systems @ 6
GPU
Go @ 4
Hiring @ 4
Kubernetes @ 7
Mathematics @ 4
Observability
Python @ 4
SRE @ 4
Slurm @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is hiring experienced Senior Production Engineers to help scale its AI infrastructure. Production Engineering is treated as a software engineering discipline, with significant contributions expected to the codebase. The role focuses on reliability assessments, incident management, production system observability, monitoring and alerting, automated deployments, and toil elimination.
NVIDIA's DGX Cloud team develops production systems that enable large-scale GPU clusters for a variety of AI workloads, including custom software for GPU asset provisioning, configuration, and lifecycle management across cloud providers.
Responsibilities
- Develop and operate production systems supporting scalable GPU clusters for AI workloads.
- Build custom software for GPU asset provisioning, configuration, and lifecycle management across cloud providers.
- Implement monitoring and health management capabilities to improve the reliability, availability, and scalability of GPU assets.
- Analyze data from GPU hardware diagnostics, cluster telemetry, and network telemetry.
- Work with teams across NVIDIA to ensure production AI clusters operate reliably, consistently, and at maximum performance.
- Evaluate system failures and improve services through a well-defined incident management process.
- Contribute substantially to NVIDIA's production engineering codebase.
Requirements
- Direct experience in a Production Engineering, DevOps, or Site Reliability Engineering role within a highly technical organization, with demonstrable impact.
- Strong communication skills and the ability to work with multifunctional teams, principals, and architects across organizational boundaries and geographies.
- 8 or more years of experience in a similar role and experience operating large-scale production systems.
- Experience with Production Engineering, DevOps, or SRE principles, tools, and techniques.
- A bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a comparable discipline, or equivalent experience.
- Technical knowledge of a systems programming language such as Go or Python.
- Solid understanding of data structures and algorithms.
Preferred Qualifications
- Technical competency managing and automating large-scale distributed systems independently of cloud providers.
- Advanced hands-on experience with cluster management systems, including Kubernetes, Slurm, and Bright Cluster Manager.
- Proven operational excellence maintaining reliable and performant AI infrastructure.
Compensation and Benefits
- Level 4 base salary: USD 168,000–270,250 per year.
- Level 5 base salary: USD 208,000–333,500 per year.
- Eligible for equity and benefits.
- Applications will be accepted at least until July 10, 2026.
- This posting is for an existing vacancy.
- NVIDIA uses AI tools in its recruiting processes.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Director, Global Risk and Compliance
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior System Software Engineer - CPU SoC Boot Firmware
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Forward-Deployed Engineer, Enterprise AI and Automation
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Software Engineer, Distributed Systems Engineer - DGX Cloud
Nvidia · Durham, United States
USD 272,000-431,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Software Engineer, DGXC Data Services
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Cloud Software Engineer, DGXC Data Services
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year