Principal Software Engineer - DGX Cloud

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 3 API @ 4 AWS @ 7 Azure @ 7 CUDA @ 3 Cloud Computing @ 7 Communication @ 6 Data Pipelines @ 4 Distributed Systems @ 4 Docker @ 7 GCP @ 7 GPU Go @ 6 Grafana Java @ 6 Kubernetes @ 7 Machine Learning OpenTelemetry Prometheus Python @ 6 Security @ 4 Slurm Technical Leadership

Details

NVIDIA is looking for a Principal Software Engineer to join the DGX Cloud team and build foundational systems for high-performance GPU infrastructure. The role focuses on scalable automation, system integration, and seamless workflows across global cloud operations. As a Principal Engineer, you will provide technical leadership and help shape the platform supporting AI and cloud computing.

Responsibilities

  • Lead the development of next-generation APIs, state management, and workflow orchestration systems that automate fleet lifecycle operations at massive scale.
  • Drive technical alignment across dependent systems and partner teams to ensure cohesive integration, clear interfaces, and reliable end-to-end workflows.
  • Coach and mentor senior engineers while elevating technical standards and guidelines across the organization.
  • Maintain a strong focus on customer experience and product requirements, translating technical insight into high-impact business solutions.
  • Partner with executive and engineering leadership to codify critical business processes into self-measuring, scalable, and operationally consistent platforms that reduce manual effort.
  • Direct integration strategies for technologies including Kubernetes, Slurm, Prometheus, OpenTelemetry, and Grafana.

Requirements

  • 16 or more years of progressive industry experience.
  • Master's or bachelor's degree, or equivalent experience defining and shipping complex distributed systems.
  • Deep hands-on expertise establishing, operating, and scaling services in fast-paced, high-reliability environments.
  • Ability to work effectively in ambiguous, fast-paced environments by testing ideas, iterating toward working solutions, and hardening successful approaches into reliable, scalable systems.
  • Outstanding proficiency in modern systems programming languages such as Go, Java, or Python.
  • Proven experience defining, owning, and evolving the architecture of high-scale distributed systems, including advanced patterns for APIs, control planes, and data pipelines.
  • Deep understanding of global cloud infrastructure, including AWS, GCP, and Azure, as well as container ecosystems such as Docker and Kubernetes.
  • Ability to drive technical strategy and influence outcomes across organizational boundaries.
  • Excellent communication skills, including the ability to explain complex technical concepts, build organizational consensus, and mentor high-performing engineers.

Preferred Qualifications

  • Experience leading the development and adoption of organization-wide workflow orchestration systems for petabyte-scale infrastructure.
  • Experience working in a Principal or Staff+ capacity and delivering measurable improvements in operational efficiency, reliability, and security across a large engineering organization.
  • Familiarity with the operational and deployment aspects of the NVIDIA AI/ML software stack, including CUDA, cuDNN, and containerization.
  • Patent contributions or a strong publication record in distributed systems, cloud computing, or infrastructure automation.

Benefits

NVIDIA offers competitive salaries, equity, and a comprehensive benefits package.

The base salary range is USD 272,000–431,250. Applications will be accepted at least until May 3, 2026. NVIDIA is an equal opportunity employer.

More jobs at Nvidia

Similar jobs