Principal Software Engineer, GPU Firmware And GPU System Software — CSP Engagements

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

GPU @ 4 NVLink @ 4 Observability Security

Details

Responsibilities

  • Drive GPU firmware & siftware work streams with CSP engineering teams — ensuring they understand GPU firmware architecture (VBIOS, InfoROM, microcontroller firmware), update sequencing, recovery procedures, and GPU power management
  • Gather and synthesize CSP feedback on GPU firmware/software — covering manageability, observability, security requirements (e.g., multi-tenancy isolation, secure boot, attestation), and performance — and champion those priorities into NVIDIA's GPU firmware/software feature roadmap and delivery plan
  • Drive GPU firmware update orchestration for large-scale deployments — multi-GPU update sequencing, rollback strategy, failure handling, and validation across hundreds of GPUs per rack
  • Serve as the technical focal point between NVIDIA and CSP firmware/software engineering — ensuring GPU behaviors (error recovery flows, thermal protection, power state transitions) are well-documented and accessible for customer integration
  • Identify cross-CSP GPU SW/FW issue patterns — common update failures, recovery gaps, and configuration problems — and drive documentation, tooling, and test strategy improvements

Requirements

  • 15+ years of experience in GPU system software, GPU firmware, or accelerator platform engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience)
  • Deep understanding of GPU architecture internals: streaming multiprocessors, GEMM execution, compute kernels, memory hierarchy, and how firmware/driver decisions impact GPU compute performance
  • Understanding of multi-GPU fabric architectures (NVLink, or similar) and how firmware coordinates across multiple GPUs in a rack-scale system
  • Understanding of GPU firmware architecture: VBIOS, GPU microcontroller firmware, InfoROM, and their interaction with the GPU driver stack
  • Experience with firmware update lifecycle management at scale: multi-device update sequencing, A/B updates, rollback, staged rollout, emergency recovery
  • Understanding of GPU error handling and recovery flows — how firmware-level errors propagate through the driver stack to application-visible failures
  • Experience with GPU health monitoring and telemetry: Xid errors, thermal events, power events, ECC counters, and their significance for firmware/software teams
  • Customer obsession — genuine passion for simplifying GPU firmware integration for fleet-scale customers. Proven success influencing engineering teams to improve quality and fleet manageability

Benefits

  • Eligible for equity and benefits

More jobs at Nvidia

Similar jobs