Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI GPU @ 4 HPC NVLink @ 4 Observability Security @ 4

Details

We're looking for a Principal Software Engineer to join the CSP Engagements team as the technical focal point for GPU firmware and GPU system software. The role works directly with engineering teams at key CSP and hyperscale customers to ensure they can reliably manage, update, and operate NVIDIA GPU firmware at fleet scale.

You will drive work streams with CSP and hyperscale customer engineering teams to build shared understanding of GPU firmware and system software integration, incorporate customer feedback into NVIDIA's feature roadmap and delivery plan, and ensure customer-side automation and recovery procedures are ready before each firmware release. Cross-CSP visibility enables you to identify patterns in GPU firmware operational challenges and drive systemic improvements.

Responsibilities

  • Drive GPU firmware and software work streams with CSP engineering teams, ensuring they understand GPU firmware architecture, including VBIOS, InfoROM, microcontroller firmware, update sequencing, recovery procedures, and GPU power management.
  • Gather and synthesize CSP feedback on GPU firmware and software, including manageability, observability, security requirements such as multi-tenancy isolation, secure boot, and attestation, and performance.
  • Champion customer priorities in NVIDIA's GPU firmware and software feature roadmap and delivery plan.
  • Drive GPU firmware update orchestration for large-scale deployments, including multi-GPU update sequencing, rollback strategy, failure handling, and validation across hundreds of GPUs per rack.
  • Serve as the technical focal point between NVIDIA and CSP firmware and software engineering teams.
  • Ensure GPU behaviors, including error recovery flows, thermal protection, and power state transitions, are well documented and accessible for customer integration.
  • Identify cross-CSP GPU software and firmware issue patterns, including common update failures, recovery gaps, and configuration problems.
  • Drive improvements to documentation, tooling, and test strategies.

Requirements

  • 15 or more years of experience in GPU system software, GPU firmware, or accelerator platform engineering.
  • Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • Deep understanding of GPU architecture internals, including streaming multiprocessors, GEMM execution, compute kernels, memory hierarchy, and the impact of firmware and driver decisions on GPU compute performance.
  • Understanding of multi-GPU fabric architectures such as NVLink and how firmware coordinates across multiple GPUs in rack-scale systems.
  • Understanding of GPU firmware architecture, including VBIOS, GPU microcontroller firmware, InfoROM, and interaction with the GPU driver stack.
  • Experience with firmware update lifecycle management at scale, including multi-device update sequencing, A/B updates, rollback, staged rollout, and emergency recovery.
  • Understanding of GPU error handling and recovery flows, including how firmware-level errors propagate through the driver stack to application-visible failures.
  • Experience with GPU health monitoring and telemetry, including Xid errors, thermal events, power events, and ECC counters.
  • Customer obsession and a proven ability to influence engineering teams to improve quality and fleet manageability.

Preferred Qualifications

  • Direct experience with NVIDIA GPU VBIOS, GPU microcontroller firmware, or GPU driver internals.
  • Background in GPU fleet management at a scale of more than 10,000 GPUs, including firmware rollout, health-based remediation, and fleet-wide configuration management.
  • Experience with GPU error taxonomy, including Xid classification, NVLink error counters, and ECC events, as well as building runbooks around GPU firmware behavior.
  • Understanding of GPU security, including secure boot chains, code signing, attestation, debug authentication, and firmware-level multi-tenancy isolation.
  • Familiarity with GPU power management architecture and its impact on workload performance at fleet scale.

Company Information

NVIDIA develops technologies in artificial intelligence, high-performance computing, and visualization. The company is committed to fostering an inclusive work environment and is an equal opportunity employer. NVIDIA uses AI tools in its recruiting processes.

Compensation and Benefits

The base salary range is USD 272,000–431,250 per year. Salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until August 1, 2026.

More jobs at Nvidia

Similar jobs