Principal Software Engineer, At-Scale Reliability And Fleet Intelligence — CSP Engagements
at Nvidia
USD 272,000-431,200 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 7
GPU @ 4
Leadership @ 7
Observability @ 4
Statistics @ 4
Stress Testing @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Principal Software Engineer to join the CSP Engagements team as the technical focal point for fleet-scale reliability, working directly with engineering teams of key CSP / hyperscale customers to ensure NVIDIA platforms achieve target MTBI (Mean Time Between Interruptions) in production.
Responsibilities
- Drive reliability work streams with CSP engineering teams — ensuring shared understanding of MTBI measurement methodology, failure classification, and health monitoring architecture
- Gather and synthesize CSP fleet reliability data — identify failure patterns that appear across multiple customers and champion improvements back into NVIDIA's firmware, driver, and hardware teams
- Define consistent MTBI measurement methodology that works across different CSP monitoring environments and operational practices
- Conduct fleet-scale failure pattern analysis using statistical methods (Pareto, survival analysis, Weibull) to classify failures as systemic, environmental, or configuration-specific
- Drive fleet health monitoring integration architecture — ensure NVIDIA's health agents, telemetry, and reporting align with CSP operational workflows and automation
- Define burn-in reliability test environment and cluster certification criteria in collaboration with quality teams, validating with customers that criteria are meaningful
- Collaborate with CSPs to ensure reliability-related integration work (health monitoring deployment, telemetry pipeline, alerting configuration) is complete ahead of at-scale launch
- Develop predictive failure models using fleet telemetry and validate their effectiveness in customer environments
Requirements
- 15+ years of experience in systems software at datacenter scale, or reliability engineering with focus on at-scale challenges.
- BS or MS in Computer Science, Electrical Engineering, Statistics, or related field (or equivalent experience)
- Deep expertise in multi-NUMA, rack-scale system software and firmware; statistical failure analysis methods: MTBF/MTBI calculation, Pareto analysis, root cause classification
- Experience with fleet-level telemetry and observability systems: time-series databases, anomaly detection, health scoring, event correlation
- Understanding of hardware failure modes in large-scale GPU/accelerator deployments — ability to classify and prioritize across compute, interconnect, memory, power, and thermal domains
- Experience defining or operating burn-in, stress testing, or certification frameworks for complex hardware systems; familiarity with predictive maintenance or anomaly detection approaches applied to fleet health data
- Customer obsession — passion for understanding fleet reliability challenges at scale and translating them into actionable engineering priorities
- Strong communication — ability to present statistical reliability findings to both deep technical audiences and executive leadership; demonstrated success driving cross-functional improvements across hardware, firmware, and software teams without direct authority
Benefits
- Eligible for equity and benefits (NVIDIA benefits)
More jobs at Nvidia
Senior Software Technical Program Manager
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Product Architect
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Manager, Storage Production Engineering
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Intellectual Property Security Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Technical Program Manager, Compute
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 290,000-365,000 per year
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Agentic GenAI Platform
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Principal Software Engineer, Rack-Scale System Software — CSP Engagements
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Staff Software Engineer — AI Applications And Platform Foundations
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Data Scientist - Cloud Gaming And Ai
Nvidia · Santa Clara, United States
USD 248,000-379,500 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior Technical Program Manager, Cloud Infrastructure Npi
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year