Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI Communication @ 7 GPU @ 4 HPC @ 4 Leadership @ 7 NVLink @ 3 Observability @ 4 Statistics @ 4

Details

NVIDIA is seeking a Principal Software Engineer to join the CSP Engagements team as the technical focal point for fleet-scale reliability. The role works directly with engineering teams at key CSP and hyperscale customers to ensure NVIDIA platforms achieve target MTBI (Mean Time Between Interruptions) in production. You will augment NVIDIA’s internal software, firmware, and quality teams with a dedicated CSP-facing focus, using customer fleet telemetry and failure data to improve reliability and validate that laboratory improvements translate to real customer environments.

Responsibilities

  • Drive reliability work streams with CSP engineering teams, including shared understanding of MTBI measurement methodology, failure classification, and health-monitoring architecture.
  • Gather and synthesize CSP fleet reliability data, identify cross-customer failure patterns, and champion improvements with NVIDIA firmware, driver, and hardware teams.
  • Define consistent MTBI measurement methodology across different CSP monitoring environments and operational practices.
  • Conduct fleet-scale failure-pattern analysis using Pareto analysis, survival analysis, and Weibull analysis to classify failures as systemic, environmental, or configuration-specific.
  • Drive fleet health-monitoring integration architecture, ensuring NVIDIA health agents, telemetry, and reporting align with CSP operational workflows and automation.
  • Define burn-in reliability test environments and cluster certification criteria with quality teams, and validate that the criteria are meaningful to customers.
  • Collaborate with CSPs to complete reliability-related integration work, including health-monitoring deployment, telemetry pipelines, and alerting configuration, before at-scale launch.
  • Develop predictive failure models using fleet telemetry and validate their effectiveness in customer environments.

Requirements

  • 15+ years of experience in systems software at datacenter scale or reliability engineering focused on at-scale challenges.
  • Bachelor’s or master’s degree in Computer Science, Electrical Engineering, Statistics, or a related field, or equivalent experience.
  • Deep expertise in multi-NUMA and rack-scale system software and firmware.
  • Expertise in statistical failure-analysis methods, including MTBF/MTBI calculation, Pareto analysis, and root-cause classification.
  • Experience with fleet-level telemetry and observability systems, including time-series databases, anomaly detection, health scoring, and event correlation.
  • Understanding of hardware failure modes in large-scale GPU and accelerator deployments, including compute, interconnect, memory, power, and thermal domains.
  • Experience defining or operating burn-in, stress-testing, or certification frameworks for complex hardware systems.
  • Familiarity with predictive maintenance or anomaly-detection approaches applied to fleet health data.
  • Strong customer focus and the ability to translate fleet reliability challenges into actionable engineering priorities.
  • Strong communication skills, including presenting statistical reliability findings to technical audiences and executive leadership.
  • Demonstrated success driving cross-functional improvements across hardware, firmware, and software teams without direct authority.

Preferred Qualifications

  • Experience with fleet reliability at a hyperscaler or leading CSP, including hardware health and fleet reliability.
  • Familiarity with NVIDIA GPU error taxonomy, including Xid errors, NVLink error counters, thermal events, and CPER records.
  • Experience building health-scoring or predictive-failure models for accelerator or HPC infrastructure.
  • Background defining MTBI/MTBF measurement standards or certification programs for complex multi-component systems.
  • Understanding of how reliability data flows from device firmware through telemetry pipelines to fleet-level dashboards and automated remediation.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

Applications will be accepted at least until July 24, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs