Senior Software Engineer - DGX Cloud

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI API Algorithms @ 4 CI/CD @ 7 Cloud Computing @ 7 Data Structures @ 4 Distributed Systems @ 4 GPU @ 7 Go @ 4 Kubernetes @ 7 Linux @ 4 Security @ 4

Details

NVIDIA's DSX Kubernetes Fleet team within the DGX Cloud organization builds automation, lifecycle management, and deployment safety for a large-scale, GPU-accelerated container platform. The team develops APIs and workflows that turn high-level deployment intent into production-ready AI infrastructure management and owns the full lifecycle of AI Factory Kubernetes clusters, from provisioning through upgrades and decommissioning.

Responsibilities

  • Explore innovative ways to simplify the development, deployment, and monitoring of GPU- and DPU-accelerated applications.
  • Design and develop software for managing fleets of Kubernetes clusters for GPUs and DPUs.
  • Work with Cloud Native technologies to improve NVIDIA accelerators in Kubernetes environments.
  • Collaborate with engineering teams across NVIDIA to ensure seamless software integration.
  • Automate and optimize build, test, integration, and release processes for cloud-native applications.
  • Multitask across different projects and address evolving priorities effectively.

Requirements

  • Bachelor's or master's degree in Computer Science or a related field, or equivalent experience.
  • 10+ years of proven work experience in large-scale environments.
  • Expert-level knowledge of systems programming languages, including Go and C, with a solid understanding of data structures and algorithms.
  • Strong understanding of container orchestration systems, particularly Kubernetes, and container technology.
  • In-depth knowledge of Unix/Unix-like kernel internals, particularly Linux.
  • Hands-on automation experience with modern infrastructure tools and technologies.
  • Proven experience setting up, maintaining, and automating continuous deployment systems.
  • Strong background in cloud computing and distributed software design and development.
  • Understanding of performance, security, and reliability in complex distributed systems.

Preferred Qualifications

  • Extensive experience with Go.
  • Deep understanding of rack-scale GPU systems.
  • Strong background with GitLab, Argo, Flux, and other CI/CD systems.
  • Significant hands-on experience with containers and Kubernetes.
  • Hands-on experience with container workload isolation and confidential computing.

Benefits

NVIDIA offers a comprehensive benefits package, equity, and competitive compensation. The base salary depends on location, experience, and compensation for employees in similar positions. Applications will be accepted at least until August 11, 2026. This posting is for an existing vacancy.

More jobs at Nvidia

Similar jobs