Senior Software Engineer, Distributed Systems Engineering - DGX Cloud

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API @ 4 Algorithms @ 4 Communication @ 7 Data Structures @ 4 Distributed Systems @ 6 GPU Go @ 4 Hiring @ 4 Kubernetes @ 4 Mathematics @ 4 Python @ 4 Slurm @ 7 Software Development @ 4

Details

NVIDIA is hiring experienced software engineers with Kubernetes experience to help scale its AI infrastructure. The role focuses on developing and operating production systems that enable large-scale GPU clusters for a variety of AI workloads. Candidates should be creative, passionate about Kubernetes and GPUs, and able to execute effectively in a highly technical environment.

Responsibilities

  • Contribute to the DGX Cloud team responsible for production systems supporting large-scale GPU clusters and AI workloads.
  • Develop custom software for scheduling GPU resources on Kubernetes.
  • Implement monitoring and health management capabilities for the reliability, availability, and scalability of GPU assets.
  • Work with data streams including GPU hardware diagnostics, cluster telemetry, and network telemetry.
  • Collaborate with teams across NVIDIA to ensure production AI clusters operate reliably, consistently, and at maximum performance.
  • Evaluate system failures and improve services through a defined incident management process.

Requirements

  • Direct software engineering experience in a highly technical organization, with demonstrable impact from previous work.
  • Software development experience with Kubernetes APIs and frameworks, beyond simply operating a cluster.
  • Strong communication skills and the ability to work with multifunctional teams, principals, and architects across organizational boundaries and geographies.
  • At least 5 years of experience in a similar role and experience working with large-scale production systems.
  • Knowledge of common software engineering principles, tools, and techniques.
  • A bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or a comparable field, or equivalent experience.
  • Technical knowledge of a systems programming language such as Go or Python.
  • Solid understanding of data structures and algorithms.

Preferred Qualifications

  • Technical competency in managing and automating large-scale distributed systems independently of cloud providers.
  • Advanced hands-on experience and deep understanding of cluster management systems such as Kubernetes, Slurm, and Bright Cluster Manager.
  • Proven operational excellence in maintaining reliable and performant AI infrastructure.

Benefits

  • Base salary is determined based on location, experience, and the pay of employees in similar positions.
  • Base salary range: $152,000–$241,500 for Level 3, or $184,000–$287,500 for Level 4.
  • Eligible for equity and benefits.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

Applications will be accepted at least until October 12, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs