Senior Software Engineer, Resilience Engineering - DGX Cloud

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 GPU @ 4 Go @ 7 Grafana @ 6 HPC @ 4 Observability @ 6 OpenTelemetry @ 6 Prometheus @ 6 Python @ 7 SRE @ 4 Technical Leadership

Details

NVIDIA is seeking a seasoned engineer to join the DGX Cloud Resilience Engineering team and help redefine reliability and operational excellence for large-scale systems in a 24/7 environment.

Responsibilities

  • Build an organization-wide reliability strategy and guide the maturation of operational practices.
  • Establish and maintain a rigorous SLO program across teams.
  • Lead incident response for high-severity incidents, ensuring low-drama and high-signal resolution.
  • Build and improve production code daily, enhancing the data platform and related tooling.
  • Implement chaos engineering, failure injection, and resilience testing.
  • Improve engineering standards through hands-on technical leadership and example.

Requirements

  • Deep, hands-on experience running large-scale production systems with a proven track record.
  • Detailed understanding of failure modes in large systems, including cascading dependencies and retry storms.
  • Strong software engineering skills with current, hands-on experience in Go, Python, or similar languages.
  • Proven experience establishing and maintaining an SLO program with operational rigor.
  • Practical experience with reliability practices such as chaos engineering and failure injection.
  • Ability to influence across team boundaries through credibility and expertise.
  • 10 or more years of industry experience with a bachelor's or master's degree, or equivalent experience operating systems at scale.

Preferred Qualifications

  • Experience in a world-class reliability function, such as Google SRE or Meta production engineering.
  • Expertise operating GPU, HPC, or AI training infrastructure and understanding its failure modes.
  • Track record of measurable reliability improvements within an organization.
  • Proficiency with modern observability and operational tools such as Prometheus, OpenTelemetry, Grafana, PagerDuty, and Rootly.

Compensation and Benefits

The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Compensation is determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.

Applications will be accepted at least until June 27, 2026. This posting is for an existing vacancy.

More jobs at Nvidia

Similar jobs