Senior Software Engineer, Cloud-Native Stack – CSP Engagements

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI Ansible CI/CD @ 3 CUDA @ 4 Communication @ 6 Debugging @ 4 Deep Learning @ 4 Distributed Systems @ 8 GPU @ 4 GitHub @ 3 GitHub Actions @ 3 Go @ 8 Helm IaC InfiniBand Kubernetes @ 7 Machine Learning Microservices Networking @ 4 Observability @ 3 OpenTelemetry @ 3 Prometheus @ 3 Python @ 8 Rust @ 8 Slurm @ 7 Software Development @ 8 Terraform

Details

NVIDIA is developing advanced multi-rack, multi-tenant AI/ML datacenters using NVIDIA GB200 and upcoming GB300 GPUs. The CSP Engagements team is seeking a Senior Software Engineer to focus on the cloud-native stack for datacenter products such as GB200. The role involves defining customer workflows, prototyping stack enhancements, and debugging complex Kubernetes and Slurm issues across multi-rack, multi-tenant AI datacenters.

Responsibilities

  • Perform deep-dive debugging of multi-rack, multi-tenant clusters, including scheduler behavior, container runtime issues, device-plugin crashes, and RDMA/InfiniBand fabric anomalies.
  • Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom microservices that expose new GPU capabilities.
  • Drive joint architecture reviews and whiteboard sessions with cloud service provider and internal platform teams; convert findings into RFCs and upstream pull requests.
  • Create reproducible testbeds using Helm, Ansible, and Terraform to mirror customer environments; automate validation and benchmark suites.
  • Deliver technical collateral, including design documents, how-to guides, and demo scripts, and present at customer on-sites, KubeCon, and SlurmUG.
  • Collaborate with account executive, field application engineer, and solution architect teams to deliver integrated customer solutions and technical documentation.

Requirements

  • Strong source-level expertise in Kubernetes internals, including the scheduler, CRI, CNI, CSI, and operators.
  • Strong expertise in Slurm, including federation, power-save, and plugins.
  • Hands-on experience integrating next-generation GPUs, including Blackwell, GB200, or GB300, or comparable accelerators into containerized clusters.
  • Proven experience debugging large-scale, cloud-native stacks across networking, including RDMA/RoCE, storage, and control planes.
  • Customer-facing engineering or solutions architecture experience, including requirements gathering, proof-of-concept ownership, and roadmap influence.
  • Familiarity with CI/CD technologies such as GitHub Actions and Tekton, observability tools such as Prometheus and OpenTelemetry, and infrastructure as code.
  • Excellent communication skills, with the ability to switch between deep technical detail and high-level business impact.
  • At least 10 years of professional software development experience in distributed systems using Go, Rust, C, C++, or Python for tooling.
  • Bachelor’s or master’s degree, or equivalent experience, in Computer Engineering, Computer Science, or a related field.

Preferred Qualifications

  • Upstream contributions to Kubernetes, Slurm, Volcano, or similar projects.
  • Experience with GPU computing, including CUDA, and deep learning workloads.

Benefits

  • Base salary range of $184,000–$287,500 for Level 4 or $224,000–$356,500 for Level 5, determined by location, experience, and compensation for comparable positions.
  • Eligibility for equity and benefits.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.
  • Applications will be accepted at least until August 3, 2026.

More jobs at Nvidia

Similar jobs