Senior Software Engineer, Resilience Engineering - DGX Cloud

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 GPU @ 4 Go @ 7 Grafana @ 6 HPC @ 4 Leadership @ 4 Observability @ 6 OpenTelemetry @ 6 Prometheus @ 6 Python @ 7 SRE @ 4 Technical Leadership @ 4

Details

NVIDIA is seeking a seasoned engineer to join the DGX Cloud team as a Senior Software Engineer specializing in resilience engineering. The role focuses on redefining reliability practices, building large-scale reliability systems, and driving operational excellence in a 24/7 environment.

Responsibilities

  • Build an organization-wide reliability strategy and guide the maturation of NVIDIA's operational practices.
  • Establish and maintain a rigorous service-level objective (SLO) program across teams.
  • Lead incident response for high-severity incidents, ensuring low-drama and high-signal resolution.
  • Build and improve production code daily, enhancing the data platform and related tooling.
  • Implement chaos engineering, failure injection, and resilience testing.
  • Improve engineering standards through hands-on technical leadership and practical experience.

Requirements

  • Deep, hands-on experience operating large-scale production systems with a proven track record.
  • Detailed understanding of failure modes in large systems, including cascading dependencies and retry storms.
  • Strong software engineering skills with current, hands-on experience in Go, Python, or similar programming languages.
  • Proven experience establishing and maintaining an SLO program with operational rigor.
  • Practical experience in reliability engineering, including chaos engineering and failure injection.
  • Ability to influence across team boundaries through credibility and technical expertise.
  • 8 or more years of industry experience.
  • Bachelor's or master's degree, or equivalent experience operating systems at scale.

Preferred Qualifications

  • Experience in a world-class reliability function, such as Google SRE or Meta production engineering.
  • Expertise operating GPU, HPC, or AI training infrastructure and understanding its failure modes.
  • Track record of measurable reliability improvements within an organization.
  • Proficiency with modern observability and operational tools such as Prometheus, OpenTelemetry, Grafana, PagerDuty, and Rootly.

Benefits

The role offers competitive salaries, equity, and benefits. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Applications will be accepted at least until August 1, 2026. NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs