Principal Software Engineer, Rack-Scale System Software — CSP Engagements

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI API @ 4 Communication @ 7 Design Patterns @ 4 Distributed Systems @ 8 GPU @ 7 HPC InfiniBand Leadership @ 6 NVLink Networking Observability @ 4 Technical Leadership @ 6

Details

We're looking for a Principal Software Engineer to join NVIDIA's CSP Engagements team as the technical focal point for rack-scale system software and firmware. You will work with cloud service provider engineering teams to ensure they can deploy, monitor, and operate rack-scale systems reliably at fleet scale.

The role focuses on system-level software that manages, monitors, and recovers the rack as a whole, including fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and software-driven serviceability. You will drive work streams with CSP engineering teams, build shared understanding of the architecture, incorporate operational feedback, and ensure integration readiness.

Responsibilities

  • Drive rack-scale software and firmware architecture alignment across CSP engagements, including fabric management software, link health monitoring, GPU/NVSwitch error handling, software- and firmware-based serviceability features, hot-plug support, component isolation, firmware-driven recovery, and multi-component firmware orchestration.
  • Drive technical work streams with CSP engineering teams on rack-scale system software, including fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and software- and firmware-controlled recovery operations.
  • Capture and synthesize CSP engineering feedback on health monitoring APIs, software-driven serviceability workflows, firmware update orchestration, and error recovery behavior, and champion that feedback in NVIDIA's architecture decisions.
  • Collaborate with cross-functional teams to ensure customer operational requirements are reflected in system software and firmware development.
  • Identify cross-CSP patterns in rack-scale software and firmware issues, error handling behavior, and system configuration practices, and drive documentation, tooling, and test strategy improvements.
  • Collaborate with execution teams on a left-shift strategy, ensuring customer-side software and firmware integration work is identified early and completed ahead of hardware availability.
  • Make critical technical decisions on rack-scale system software and firmware trade-offs and mitigate execution risks through early engagement with CSP engineering teams.

Requirements

  • 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering.
  • Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • Deep understanding of rack-scale system software challenges, including multi-component coordination, error propagation, health monitoring, serviceability, and reliability.
  • Experience with fabric management software, cluster management, or system-level orchestration frameworks.
  • Familiarity with firmware architectures and update lifecycle management, including multi-component update sequencing, rollback, and recovery.
  • Understanding of distributed-systems error handling and recovery design patterns, including fault isolation, retry policies, and graceful degradation.
  • Experience with health monitoring and telemetry systems, including health scoring, event correlation, and API design for fleet-level observability.
  • Understanding of GPU or accelerator system software, including drivers, device management, and power management, is a strong plus.
  • Customer-focused approach and genuine interest in understanding how CSPs operate sophisticated systems at fleet scale.
  • Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority.
  • Strong communication skills, including the ability to translate complex system software architecture into actionable mentorship for customer engineering teams.

Preferred Qualifications

  • Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software.
  • Background in system software for large-scale clusters at a hyperscaler, including cluster management, fleet orchestration, or health platforms.
  • Experience creating error handling and recovery frameworks for multi-component systems with hundreds or thousands of coordinating devices.
  • Familiarity with GPU or accelerator fleet operations, including driver lifecycle management, firmware rollout strategies, and health-based scheduling.
  • Understanding of how system software decisions affect serviceability, availability, and operational cost at fleet scale.

Company Information

NVIDIA data center systems, including DGX and HGX, combine NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and an optimized NVIDIA AI and HPC software stack.

Compensation And Benefits

The base salary range is USD 272,000–431,250 per year, determined by location, experience, and pay for employees in similar positions. The role is also eligible for equity and benefits.

NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. Applications will be accepted at least until August 1, 2026. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs