Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements
at Nvidia
USD 272,000-431,200 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 7
GPU @ 4
HPC @ 4
Leadership @ 7
NVLink @ 3
Observability @ 4
Statistics @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Principal Software Engineer to join the CSP Engagements team as the technical focal point for fleet-scale reliability. The role works directly with engineering teams at key CSP and hyperscale customers to ensure NVIDIA platforms achieve target MTBI (Mean Time Between Interruptions) in production. You will augment NVIDIA’s internal software, firmware, and quality teams with a dedicated CSP-facing focus, using customer fleet telemetry and failure data to improve reliability and validate that laboratory improvements translate to real customer environments.
Responsibilities
- Drive reliability work streams with CSP engineering teams, including shared understanding of MTBI measurement methodology, failure classification, and health-monitoring architecture.
- Gather and synthesize CSP fleet reliability data, identify cross-customer failure patterns, and champion improvements with NVIDIA firmware, driver, and hardware teams.
- Define consistent MTBI measurement methodology across different CSP monitoring environments and operational practices.
- Conduct fleet-scale failure-pattern analysis using Pareto analysis, survival analysis, and Weibull analysis to classify failures as systemic, environmental, or configuration-specific.
- Drive fleet health-monitoring integration architecture, ensuring NVIDIA health agents, telemetry, and reporting align with CSP operational workflows and automation.
- Define burn-in reliability test environments and cluster certification criteria with quality teams, and validate that the criteria are meaningful to customers.
- Collaborate with CSPs to complete reliability-related integration work, including health-monitoring deployment, telemetry pipelines, and alerting configuration, before at-scale launch.
- Develop predictive failure models using fleet telemetry and validate their effectiveness in customer environments.
Requirements
- 15+ years of experience in systems software at datacenter scale or reliability engineering focused on at-scale challenges.
- Bachelor’s or master’s degree in Computer Science, Electrical Engineering, Statistics, or a related field, or equivalent experience.
- Deep expertise in multi-NUMA and rack-scale system software and firmware.
- Expertise in statistical failure-analysis methods, including MTBF/MTBI calculation, Pareto analysis, and root-cause classification.
- Experience with fleet-level telemetry and observability systems, including time-series databases, anomaly detection, health scoring, and event correlation.
- Understanding of hardware failure modes in large-scale GPU and accelerator deployments, including compute, interconnect, memory, power, and thermal domains.
- Experience defining or operating burn-in, stress-testing, or certification frameworks for complex hardware systems.
- Familiarity with predictive maintenance or anomaly-detection approaches applied to fleet health data.
- Strong customer focus and the ability to translate fleet reliability challenges into actionable engineering priorities.
- Strong communication skills, including presenting statistical reliability findings to technical audiences and executive leadership.
- Demonstrated success driving cross-functional improvements across hardware, firmware, and software teams without direct authority.
Preferred Qualifications
- Experience with fleet reliability at a hyperscaler or leading CSP, including hardware health and fleet reliability.
- Familiarity with NVIDIA GPU error taxonomy, including Xid errors, NVLink error counters, thermal events, and CPER records.
- Experience building health-scoring or predictive-failure models for accelerator or HPC infrastructure.
- Background defining MTBI/MTBF measurement standards or certification programs for complex multi-component systems.
- Understanding of how reliability data flows from device firmware through telemetry pipelines to fleet-level dashboards and automated remediation.
Benefits
- Equity and benefits are provided.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
Applications will be accepted at least until July 24, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Systems Software Engineer, Low Latency Streaming Technology - Automotive
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Reinforcement Learning Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Manager, Storage Engineering
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Senior Software and System Architect
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Customer Technical Program Manager
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Similar jobs
Principal Software Engineer, Rack-Scale System Software — CSP Engagements
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Technical Program Manager, Compute
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 290,000-365,000 per year
Director, Software Engineering
Nvidia · Santa Clara, United States
USD 292,000-442,800 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Distinguished Engineer, End-to-End Scaling Performance Architecture
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Principal Software Engineer, E2E Performance and Goodput — CSP Engagements
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Systems Software Engineer, Data Center Platform Enablement
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year