Senior Site Reliability Engineer, BCM - DGX Cloud

at Nvidia
USD 168,000-333,500 per year
SENIOR
✅ On-site

Tech Stack

AI GPU HPC InfiniBand @ 6 Kubernetes @ 4 Linux @ 4 Networking @ 4 Python @ 6 SRE Slurm Software Development @ 7 System Administration @ 4

Details

NVIDIA Base Command Manager powers thousands of clusters worldwide, ranging from a few to several thousand nodes, and streamlines cluster provisioning, workload management, and infrastructure monitoring. The role will contribute to the success of external customers running NVIDIA solutions and internal clusters used for research, operations, and next-generation projects.

Responsibilities

  • Contribute to deployments and daily operations of large-scale, next-generation GPU platforms.
  • Handle incidents in GPU clusters, bridging the gap between cluster operations and development.
  • Design and implement small features in the Base Command Manager product to develop an in-depth understanding of the product.
  • Validate complex cluster configurations, including Slurm and Kubernetes orchestrators, for performance, scalability, and resilience against real-world customer scenarios.

Requirements

  • Bachelor's degree or equivalent experience in Computer Science or a related field.
  • 8+ years of experience in site reliability engineering and/or software development roles.
  • Fluency in Python.
  • In-depth knowledge of Linux and networking.

Preferred Qualifications

  • Experience with C++, high-performance computing, Kubernetes, and/or system administration.
  • Previous experience as a system administrator running BCM, Bright Cluster Manager, or Base Command Manager clusters.
  • Proficiency with cluster networking, including InfiniBand and Spectrum-X.

Benefits

  • Equity.
  • NVIDIA benefits.
  • Inclusive and equal-opportunity work environment.

Applications will be accepted at least until August 21, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs