Principal Software Engineer - Rack-Scale Systems Infrastructure

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API @ 4 Debugging @ 4 Distributed Systems @ 8 GPU Go @ 4 HPC InfiniBand @ 7 Kubernetes @ 4 Linux @ 4 NVLink Networking @ 7 Observability @ 4 Rust @ 4 Security @ 6

Details

NVIDIA is developing AI and accelerated computing technologies for the next era of computing. As a Principal Rack-Scale Systems Infrastructure Engineer, you will build and guide software systems supporting rack-scale infrastructure products and services. The role sits at the intersection of software and hardware, covering control planes, state machines, orchestration systems, firmware, operating system lifecycle, and networking fabrics. You will compose infrastructure-as-a-service control plane software that makes complex rack-scale hardware dependable, manageable, and programmable for NVIDIA, partners, cloud service providers, and enterprise customers.

Responsibilities

  • Define the complete software architecture for rack-scale infrastructure products and services, including control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.
  • Use Kubernetes and cloud-native primitives, including controllers, operators, reconciliation loops, and open-source components, to manage infrastructure safely at rack and fleet scale.
  • Build open-source infrastructure software in forms such as libraries, services, controllers, operators, and integration APIs for internal deployments and cloud service provider environments.
  • Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces.
  • Translate infrastructure roadmaps into software requirements, architecture specifications, and execution plans.
  • Partner with hyperscalers, cloud service providers, enterprise customers, internal component leads, vendors, and business partners on deployment and integration requirements.
  • Establish reliability, security, validation, and left-shift strategies to reduce risk before hardware reaches production.
  • Mentor senior engineers and technical leads in large-scale networked systems, foundational software, and rack-scale control plane development.
  • Make technical decisions in ambiguous environments while balancing customer needs, schedules, hardware constraints, maintainability, open-source adoption, and long-term infrastructure evolution.

Requirements

  • BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 15+ years of experience in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.
  • Architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed-systems tradeoffs.
  • Practical production-quality coding experience with Go, C++, or Rust. Rust experience is highly valued.
  • Experience with Kubernetes or similar orchestration systems for managing infrastructure, hardware resources, or large-scale infrastructure services.
  • Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.
  • Strong understanding of data center networking technologies and protocols, including Ethernet, InfiniBand, RDMA, and fabric-level manageability.
  • Experience with complex accelerator-based systems such as GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.
  • Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols.
  • Ability to work with security experts on secure boot, attestation, access control, update safety, serviceability, and operational usability.
  • Experience crafting software for open-source release, including API stability, modularity, documentation, community usability, and separation of shared software from deployment-specific integrations.
  • Experience using AI-assisted development tools for coding, test generation, debugging, build iteration, and documentation.
  • Strong ability to specify requirements, guide architecture, manage delivery across engineering teams, and communicate complex hardware and software tradeoffs to leaders, customers, partners, and executives.

Preferred Qualifications

  • Strong Rust skills in systems, infrastructure, or hardware-adjacent software.
  • Experience building software for internal services, cloud service provider-integrated offerings, reusable libraries, and customer-extensible APIs.
  • Experience with fleet-scale provisioning, updates, rollback, observability, health monitoring, and remediation.
  • Experience leading across the data center product lifecycle, including inception, pre- and post-silicon, manufacturing, deployment, and operations.
  • Familiarity with open-source ecosystems and contribution models.
  • Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infrastructure management.

Compensation and Benefits

  • Base salary: USD 272,000–431,250 per year, determined by location, experience, and comparable employee compensation.
  • Eligible for equity and benefits.
  • Applications accepted at least until August 1, 2026.
  • NVIDIA is an equal opportunity employer.

More jobs at Nvidia

Similar jobs