Senior Software Engineer, Cloud-Native Stack – CSP Engagements
at Nvidia
USD 184,000-356,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Ansible
CI/CD @ 3
CUDA @ 4
Communication @ 6
Debugging @ 4
Deep Learning @ 4
Distributed Systems @ 8
GPU @ 4
GitHub @ 3
GitHub Actions @ 3
Go @ 8
Helm
IaC
InfiniBand
Kubernetes @ 7
Machine Learning
Microservices
Networking @ 4
Observability @ 3
OpenTelemetry @ 3
Prometheus @ 3
Python @ 8
Rust @ 8
Slurm @ 7
Software Development @ 8
Terraform
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is developing advanced multi-rack, multi-tenant AI/ML datacenters using NVIDIA GB200 and upcoming GB300 GPUs. The CSP Engagements team is seeking a Senior Software Engineer to focus on the cloud-native stack for datacenter products such as GB200. The role involves defining customer workflows, prototyping stack enhancements, and debugging complex Kubernetes and Slurm issues across multi-rack, multi-tenant AI datacenters.
Responsibilities
- Perform deep-dive debugging of multi-rack, multi-tenant clusters, including scheduler behavior, container runtime issues, device-plugin crashes, and RDMA/InfiniBand fabric anomalies.
- Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom microservices that expose new GPU capabilities.
- Drive joint architecture reviews and whiteboard sessions with cloud service provider and internal platform teams; convert findings into RFCs and upstream pull requests.
- Create reproducible testbeds using Helm, Ansible, and Terraform to mirror customer environments; automate validation and benchmark suites.
- Deliver technical collateral, including design documents, how-to guides, and demo scripts, and present at customer on-sites, KubeCon, and SlurmUG.
- Collaborate with account executive, field application engineer, and solution architect teams to deliver integrated customer solutions and technical documentation.
Requirements
- Strong source-level expertise in Kubernetes internals, including the scheduler, CRI, CNI, CSI, and operators.
- Strong expertise in Slurm, including federation, power-save, and plugins.
- Hands-on experience integrating next-generation GPUs, including Blackwell, GB200, or GB300, or comparable accelerators into containerized clusters.
- Proven experience debugging large-scale, cloud-native stacks across networking, including RDMA/RoCE, storage, and control planes.
- Customer-facing engineering or solutions architecture experience, including requirements gathering, proof-of-concept ownership, and roadmap influence.
- Familiarity with CI/CD technologies such as GitHub Actions and Tekton, observability tools such as Prometheus and OpenTelemetry, and infrastructure as code.
- Excellent communication skills, with the ability to switch between deep technical detail and high-level business impact.
- At least 10 years of professional software development experience in distributed systems using Go, Rust, C, C++, or Python for tooling.
- Bachelor’s or master’s degree, or equivalent experience, in Computer Engineering, Computer Science, or a related field.
Preferred Qualifications
- Upstream contributions to Kubernetes, Slurm, Volcano, or similar projects.
- Experience with GPU computing, including CUDA, and deep learning workloads.
Benefits
- Base salary range of $184,000–$287,500 for Level 4 or $224,000–$356,500 for Level 5, determined by location, experience, and compensation for comparable positions.
- Eligibility for equity and benefits.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
- Applications will be accepted at least until August 3, 2026.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, Golang - DSX MaxQ
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year