Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
GPU @ 4
Go @ 7
Grafana @ 6
HPC @ 4
Leadership @ 4
Observability @ 6
OpenTelemetry @ 6
Prometheus @ 6
Python @ 7
SRE @ 4
Technical Leadership @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a seasoned engineer to join the DGX Cloud team as a Senior Software Engineer specializing in resilience engineering. The role focuses on redefining reliability practices, building large-scale reliability systems, and driving operational excellence in a 24/7 environment.
Responsibilities
- Build an organization-wide reliability strategy and guide the maturation of NVIDIA's operational practices.
- Establish and maintain a rigorous service-level objective (SLO) program across teams.
- Lead incident response for high-severity incidents, ensuring low-drama and high-signal resolution.
- Build and improve production code daily, enhancing the data platform and related tooling.
- Implement chaos engineering, failure injection, and resilience testing.
- Improve engineering standards through hands-on technical leadership and practical experience.
Requirements
- Deep, hands-on experience operating large-scale production systems with a proven track record.
- Detailed understanding of failure modes in large systems, including cascading dependencies and retry storms.
- Strong software engineering skills with current, hands-on experience in Go, Python, or similar programming languages.
- Proven experience establishing and maintaining an SLO program with operational rigor.
- Practical experience in reliability engineering, including chaos engineering and failure injection.
- Ability to influence across team boundaries through credibility and technical expertise.
- 8 or more years of industry experience.
- Bachelor's or master's degree, or equivalent experience operating systems at scale.
Preferred Qualifications
- Experience in a world-class reliability function, such as Google SRE or Meta production engineering.
- Expertise operating GPU, HPC, or AI training infrastructure and understanding its failure modes.
- Track record of measurable reliability improvements within an organization.
- Proficiency with modern observability and operational tools such as Prometheus, OpenTelemetry, Grafana, PagerDuty, and Rootly.
Benefits
The role offers competitive salaries, equity, and benefits. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Applications will be accepted at least until August 1, 2026. NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Software Engineer, Resilience Engineering - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Engineer System Software, SDN Operations
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
Senior Manager, Storage Production Engineering
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Engineering Manager, Agentic GenAI Platform
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Staff Backend Engineer - Adaptive Telemetry, Databases
Grafana Labs · United States
USD 175,000-210,000 per year