Principal Software Engineer - Rack-Scale Systems Infrastructure
at Nvidia
USD 272,000-431,200 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Debugging @ 4
Distributed Systems @ 8
GPU
Go @ 4
HPC
InfiniBand @ 7
Kubernetes @ 4
Linux @ 4
NVLink
Networking @ 7
Observability @ 4
Rust @ 4
Security @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is developing AI and accelerated computing technologies for the next era of computing. As a Principal Rack-Scale Systems Infrastructure Engineer, you will build and guide software systems supporting rack-scale infrastructure products and services. The role sits at the intersection of software and hardware, covering control planes, state machines, orchestration systems, firmware, operating system lifecycle, and networking fabrics. You will compose infrastructure-as-a-service control plane software that makes complex rack-scale hardware dependable, manageable, and programmable for NVIDIA, partners, cloud service providers, and enterprise customers.
Responsibilities
- Define the complete software architecture for rack-scale infrastructure products and services, including control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.
- Use Kubernetes and cloud-native primitives, including controllers, operators, reconciliation loops, and open-source components, to manage infrastructure safely at rack and fleet scale.
- Build open-source infrastructure software in forms such as libraries, services, controllers, operators, and integration APIs for internal deployments and cloud service provider environments.
- Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces.
- Translate infrastructure roadmaps into software requirements, architecture specifications, and execution plans.
- Partner with hyperscalers, cloud service providers, enterprise customers, internal component leads, vendors, and business partners on deployment and integration requirements.
- Establish reliability, security, validation, and left-shift strategies to reduce risk before hardware reaches production.
- Mentor senior engineers and technical leads in large-scale networked systems, foundational software, and rack-scale control plane development.
- Make technical decisions in ambiguous environments while balancing customer needs, schedules, hardware constraints, maintainability, open-source adoption, and long-term infrastructure evolution.
Requirements
- BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 15+ years of experience in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.
- Architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed-systems tradeoffs.
- Practical production-quality coding experience with Go, C++, or Rust. Rust experience is highly valued.
- Experience with Kubernetes or similar orchestration systems for managing infrastructure, hardware resources, or large-scale infrastructure services.
- Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.
- Strong understanding of data center networking technologies and protocols, including Ethernet, InfiniBand, RDMA, and fabric-level manageability.
- Experience with complex accelerator-based systems such as GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.
- Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols.
- Ability to work with security experts on secure boot, attestation, access control, update safety, serviceability, and operational usability.
- Experience crafting software for open-source release, including API stability, modularity, documentation, community usability, and separation of shared software from deployment-specific integrations.
- Experience using AI-assisted development tools for coding, test generation, debugging, build iteration, and documentation.
- Strong ability to specify requirements, guide architecture, manage delivery across engineering teams, and communicate complex hardware and software tradeoffs to leaders, customers, partners, and executives.
Preferred Qualifications
- Strong Rust skills in systems, infrastructure, or hardware-adjacent software.
- Experience building software for internal services, cloud service provider-integrated offerings, reusable libraries, and customer-extensible APIs.
- Experience with fleet-scale provisioning, updates, rollback, observability, health monitoring, and remediation.
- Experience leading across the data center product lifecycle, including inception, pre- and post-silicon, manufacturing, deployment, and operations.
- Familiarity with open-source ecosystems and contribution models.
- Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infrastructure management.
Compensation and Benefits
- Base salary: USD 272,000–431,250 per year, determined by location, experience, and comparable employee compensation.
- Eligible for equity and benefits.
- Applications accepted at least until August 1, 2026.
- NVIDIA is an equal opportunity employer.
More jobs at Nvidia
Senior Software Engineer, Unified Access Management Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Technical Program Manager, AV System Integration
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
Senior System Software Engineer, Automotive Performance
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Embedded System Software Engineer – Platform Execution Lead
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Distinguished Engineer, Storage – AI Cloud
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Attestation Services – DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Perplexity AI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 250,000-485,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year