Principal Software Engineer - Compute Infrastructure

at Nvidia
USD 248,000-391,000 per year
SENIOR
✅ Hybrid

Tech Stack

AI API AWS @ 4 ArgoCD @ 4 GCP @ 4 GPU Go @ 6 Kubernetes @ 6 Leadership @ 7 Mathematics @ 4 Microservices @ 4 Networking @ 4 OpenShift @ 6 Python @ 6 Terraform @ 6

Details

NVIDIA is seeking a Principal Software Engineer to lead the architectural vision for a massive global compute platform and operationalize an internal frontier-class AI inference system. The role focuses on driving efficiency, defining platform architecture, and optimizing infrastructure performance across on-premises and cloud environments.

Responsibilities

  • Architect and transform a global enterprise compute platform running thousands of nodes and tens of thousands of virtual machines and containers through OpenShift and KubeVirt.
  • Define service tiers, SLAs, and automated cluster lifecycles.
  • Build the operational foundation for an internal AI inference platform supporting frontier-class models.
  • Develop automated remediation pipelines, hardware watchdogs, and telemetry for pre-release, rack-scale GPU systems, including Blackwell and upcoming architectures.
  • Collect and review system data for capacity planning amid hardware supply constraints.
  • Develop capacity strategies involving public cloud bursting, hardware dogfooding, and alternative compute architectures such as ARM.
  • Design self-service architectures, APIs, and Terraform/OpenTofu providers to drive adoption of standard platforms across autonomous engineering teams.
  • Evaluate application architectures and lead migrations of large legacy workloads, including long-running VDI environments, to modern Kubernetes orchestration.

Requirements

  • Bachelor's degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
  • 15+ years of experience in compute platform engineering, site reliability, or systems architecture, with a strong focus on automation at massive scale.
  • Deep expertise in Kubernetes architecture and virtualization architectures, specifically operating virtual machines inside Kubernetes with KubeVirt and OpenShift.
  • In-depth knowledge of hardware technologies, including GPUs and high-speed backplane networking, with experience mitigating hardware-level failures, silent data corruption, and anomalies in large-scale environments.
  • Experience operating large global environments spanning bare metal, virtualized infrastructure, and cloud with a unified GitOps approach using ArgoCD or similar tools.
  • Proficiency in Go and/or Python, along with expert-level infrastructure-as-code development using Terraform and configuration management tools.
  • Strong leadership skills and the ability to influence technical direction across highly autonomous teams without relying on top-down mandates.

Preferred Qualifications

  • Hands-on experience managing bleeding-edge, pre-release hardware in production environments.
  • Deep understanding of advanced storage migrations and protocols, including NFSv4, NVMe/TCP, and hyperconverged storage.
  • Understanding of microservices architecture and multi-cloud deployment strategies involving AWS and GCP.
  • Proven experience building Day 2 operational maturity, including self-service, advanced auto-remediation, and strict SLAs, on existing foundations.

Benefits

  • Base salary range of USD 248,000 to USD 391,000, determined by location, experience, and comparable employee compensation.
  • Eligibility for equity and benefits.
  • NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.

Applications will be accepted at least until September 3, 2026.

More jobs at Nvidia

Similar jobs