Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
GPU @ 4
Go @ 7
Grafana @ 6
HPC @ 4
Observability @ 6
OpenTelemetry @ 6
Prometheus @ 6
Python @ 7
SRE @ 4
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a seasoned engineer to join the DGX Cloud Resilience Engineering team and help redefine reliability and operational excellence for large-scale systems in a 24/7 environment.
Responsibilities
- Build an organization-wide reliability strategy and guide the maturation of operational practices.
- Establish and maintain a rigorous SLO program across teams.
- Lead incident response for high-severity incidents, ensuring low-drama and high-signal resolution.
- Build and improve production code daily, enhancing the data platform and related tooling.
- Implement chaos engineering, failure injection, and resilience testing.
- Improve engineering standards through hands-on technical leadership and example.
Requirements
- Deep, hands-on experience running large-scale production systems with a proven track record.
- Detailed understanding of failure modes in large systems, including cascading dependencies and retry storms.
- Strong software engineering skills with current, hands-on experience in Go, Python, or similar languages.
- Proven experience establishing and maintaining an SLO program with operational rigor.
- Practical experience with reliability practices such as chaos engineering and failure injection.
- Ability to influence across team boundaries through credibility and expertise.
- 10 or more years of industry experience with a bachelor's or master's degree, or equivalent experience operating systems at scale.
Preferred Qualifications
- Experience in a world-class reliability function, such as Google SRE or Meta production engineering.
- Expertise operating GPU, HPC, or AI training infrastructure and understanding its failure modes.
- Track record of measurable reliability improvements within an organization.
- Proficiency with modern observability and operational tools such as Prometheus, OpenTelemetry, Grafana, PagerDuty, and Rootly.
Compensation and Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Compensation is determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until June 27, 2026. This posting is for an existing vacancy.
More jobs at Nvidia
Engineering Manager, Data Labeling Platform
Nvidia · Santa Clara, United States
USD 200,000-391,000 per year
Engineering Manager, Local AI Agents
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Deep Learning Software Engineer, DLSim
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Fleet Intelligence Backend
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Business Systems Analyst
Nvidia · Santa Clara, United States
USD 144,000-270,200 per year
Similar jobs
Senior Software Engineer, Resilience Engineering - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Fleet Intelligence Agent Systems
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Engineer System Software, SDN Operations
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year