Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 4
Debugging @ 6
Distributed Systems @ 4
GPU @ 4
Go @ 7
HPC @ 4
InfiniBand @ 4
Kubernetes @ 4
Linux @ 4
Machine Learning
NVLink @ 4
Networking @ 3
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Are you passionate about Kubernetes and AI and want to help build the best platform for ML/AI infrastructure? NVIDIA's collaborative group of engineers, architects, and SREs builds and nurtures a declarative, Kubernetes-native control plane that powers GPU-accelerated infrastructure across multiple cloud providers.
The team is building a platform that gathers topology-related information from multiple sources and systems, aggregates and normalizes that data, and makes it available to provisioning systems and workload schedulers. This role will help maintain the critical open-source Topograph project and interface with NVIDIA hardware to ensure GPU-to-GPU communication is optimized for large-scale workloads across multiple providers.
Responsibilities
- Build a system that gathers topology-related information from multiple sources.
- Aggregate and normalize collected data for provisioning systems and workload schedulers.
- Contribute directly to the critical open-source Topograph project.
- Interact with current hardware to ensure new product launches have efficient scheduling capabilities.
Requirements
- At least 8 years of relevant experience.
- Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or a related technical field, or equivalent experience.
- Strong production engineering experience in Go or another systems language.
- Experience with distributed systems, Kubernetes, Slurm/Slinky, Linux, containers, APIs, and CI.
- Ability to design clean interfaces between discovery logic, data models, and scheduler output.
- Familiarity with networking, cluster topology, cloud infrastructure, or large-scale compute systems.
- Excellent testing, debugging, documentation, and code review habits.
Preferred Qualifications
- Experience with GPU clusters, NVLink, InfiniBand, Ethernet fabrics, or HPC.
- Hands-on experience with Kubernetes scheduling, Slurm/Slinky topology, DRA, Kueue, Slinky, or device plugins.
- Experience integrating with cloud provider topology APIs or cluster metadata systems.
Benefits
- Base salary range of $184,000–$287,500 USD, determined by location, experience, and compensation for similar positions.
- Eligibility for equity and benefits.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
- This is an existing vacancy, and applications will be accepted at least until July 5, 2026.
- NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Principal Software Engineer - Rack-Scale Systems Infrastructure
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · Palo Alto, United States, San Francisco, United States
USD 220,000-405,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · London, United Kingdom
GBP 325,000-485,000 per year