Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
AWS @ 4
ArgoCD @ 4
GCP @ 4
GPU
Go @ 6
Kubernetes @ 6
Leadership @ 7
Mathematics @ 4
Microservices @ 4
Networking @ 4
OpenShift @ 6
Python @ 6
Terraform @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Principal Software Engineer to lead the architectural vision for a massive global compute platform and operationalize an internal frontier-class AI inference system. The role focuses on driving efficiency, defining platform architecture, and optimizing infrastructure performance across on-premises and cloud environments.
Responsibilities
- Architect and transform a global enterprise compute platform running thousands of nodes and tens of thousands of virtual machines and containers through OpenShift and KubeVirt.
- Define service tiers, SLAs, and automated cluster lifecycles.
- Build the operational foundation for an internal AI inference platform supporting frontier-class models.
- Develop automated remediation pipelines, hardware watchdogs, and telemetry for pre-release, rack-scale GPU systems, including Blackwell and upcoming architectures.
- Collect and review system data for capacity planning amid hardware supply constraints.
- Develop capacity strategies involving public cloud bursting, hardware dogfooding, and alternative compute architectures such as ARM.
- Design self-service architectures, APIs, and Terraform/OpenTofu providers to drive adoption of standard platforms across autonomous engineering teams.
- Evaluate application architectures and lead migrations of large legacy workloads, including long-running VDI environments, to modern Kubernetes orchestration.
Requirements
- Bachelor's degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
- 15+ years of experience in compute platform engineering, site reliability, or systems architecture, with a strong focus on automation at massive scale.
- Deep expertise in Kubernetes architecture and virtualization architectures, specifically operating virtual machines inside Kubernetes with KubeVirt and OpenShift.
- In-depth knowledge of hardware technologies, including GPUs and high-speed backplane networking, with experience mitigating hardware-level failures, silent data corruption, and anomalies in large-scale environments.
- Experience operating large global environments spanning bare metal, virtualized infrastructure, and cloud with a unified GitOps approach using ArgoCD or similar tools.
- Proficiency in Go and/or Python, along with expert-level infrastructure-as-code development using Terraform and configuration management tools.
- Strong leadership skills and the ability to influence technical direction across highly autonomous teams without relying on top-down mandates.
Preferred Qualifications
- Hands-on experience managing bleeding-edge, pre-release hardware in production environments.
- Deep understanding of advanced storage migrations and protocols, including NFSv4, NVMe/TCP, and hyperconverged storage.
- Understanding of microservices architecture and multi-cloud deployment strategies involving AWS and GCP.
- Proven experience building Day 2 operational maturity, including self-service, advanced auto-remediation, and strict SLAs, on existing foundations.
Benefits
- Base salary range of USD 248,000 to USD 391,000, determined by location, experience, and comparable employee compensation.
- Eligibility for equity and benefits.
- NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until September 3, 2026.
More jobs at Nvidia
Deep Learning Algorithm Engineering Intern - 2026
Nvidia · Zurich, Switzerland
PLN 117,800-204,100 per year
Senior Project Delivery Manager - NVIS
Nvidia · United States
USD 168,000-322,000 per year
Principal Research Scientist, Synthetic Data Generation
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Staff Software Engineer - Agentic Automation
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Firmware Engineer, NIC Firmware
Nvidia · Seattle, United States
USD 152,000-287,500 per year
Similar jobs
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Staff+ Software Engineer, Node Infra
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Linux Systems Engineer - EDA Infrastructure
Nvidia · United States
USD 184,000-356,500 per year
Senior Engineer System Software, SDN Operations
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Staff Backend Software Engineer, Agent Platform
SentinelOne · United States
USD 156,000-215,000 per year
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale – DGX Cloud
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year