Principal Software Engineer, At-Scale Reliability And Fleet Intelligence — CSP Engagements

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

Communication @ 7 GPU @ 4 Leadership @ 7 Observability @ 4 Statistics @ 4 Stress Testing @ 3

Details

Principal Software Engineer to join the CSP Engagements team as the technical focal point for fleet-scale reliability, working directly with engineering teams of key CSP / hyperscale customers to ensure NVIDIA platforms achieve target MTBI (Mean Time Between Interruptions) in production.

Responsibilities

  • Drive reliability work streams with CSP engineering teams — ensuring shared understanding of MTBI measurement methodology, failure classification, and health monitoring architecture
  • Gather and synthesize CSP fleet reliability data — identify failure patterns that appear across multiple customers and champion improvements back into NVIDIA's firmware, driver, and hardware teams
  • Define consistent MTBI measurement methodology that works across different CSP monitoring environments and operational practices
  • Conduct fleet-scale failure pattern analysis using statistical methods (Pareto, survival analysis, Weibull) to classify failures as systemic, environmental, or configuration-specific
  • Drive fleet health monitoring integration architecture — ensure NVIDIA's health agents, telemetry, and reporting align with CSP operational workflows and automation
  • Define burn-in reliability test environment and cluster certification criteria in collaboration with quality teams, validating with customers that criteria are meaningful
  • Collaborate with CSPs to ensure reliability-related integration work (health monitoring deployment, telemetry pipeline, alerting configuration) is complete ahead of at-scale launch
  • Develop predictive failure models using fleet telemetry and validate their effectiveness in customer environments

Requirements

  • 15+ years of experience in systems software at datacenter scale, or reliability engineering with focus on at-scale challenges.
  • BS or MS in Computer Science, Electrical Engineering, Statistics, or related field (or equivalent experience)
  • Deep expertise in multi-NUMA, rack-scale system software and firmware; statistical failure analysis methods: MTBF/MTBI calculation, Pareto analysis, root cause classification
  • Experience with fleet-level telemetry and observability systems: time-series databases, anomaly detection, health scoring, event correlation
  • Understanding of hardware failure modes in large-scale GPU/accelerator deployments — ability to classify and prioritize across compute, interconnect, memory, power, and thermal domains
  • Experience defining or operating burn-in, stress testing, or certification frameworks for complex hardware systems; familiarity with predictive maintenance or anomaly detection approaches applied to fleet health data
  • Customer obsession — passion for understanding fleet reliability challenges at scale and translating them into actionable engineering priorities
  • Strong communication — ability to present statistical reliability findings to both deep technical audiences and executive leadership; demonstrated success driving cross-functional improvements across hardware, firmware, and software teams without direct authority

Benefits

  • Eligible for equity and benefits (NVIDIA benefits)

More jobs at Nvidia

Similar jobs