Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 4
Debugging @ 6
Distributed Systems @ 4
GPU @ 4
Go @ 7
HPC @ 4
InfiniBand @ 4
Kubernetes @ 4
Linux @ 4
Machine Learning
NVLink @ 4
Networking @ 3
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Are you passionate about Kubernetes and AI and want to help build the best platform for ML/AI infrastructure? Do you thrive when your work directly empowers teams to push the boundaries of what's possible? We're a collaborative group of engineers, architects, and SREs who are passionate about building and nurturing the declarative, Kubernetes-native control plane that powers GPU-accelerated infrastructure across multiple cloud providers.
We are building a platform that gathers topology related information from multiple sources and systems, aggregates and normalizes that data, and makes it available to provisioning systems and workload schedulers. We are looking for Senior Software Engineer who will be directly involved in not only helping maintain this critical open-source project for the community, but interfacing with bleeding edge NVIDIA hardware to ensure GPU to GPU communication is optimized for large-scale workloads across multiple providers.
What you'll be doing:
Responsibilities
- Building a system that gathers topology related information from multiple sources
- Taking data collected to aggregate and normalize the data to make it available for provisioning systems and workload schedulers
- Direct contributor in a critical open-source project, Topograph
- Interacting with the latest and greatest hardware to ensure new product launches have the most efficient scheduling capabilities
What we need to see:
Requirements
- At least 8 years of relevant experience
- Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or a related technical field, or equivalent experience
- Strong production engineering experience in Go or another systems language
- Experience with distributed systems, Kubernetes, Slurm/Slinky, Linux, containers, APIs, and CI
- Ability to design clean interfaces between discovery logic, data models, and scheduler output
- Familiarity with networking, cluster topology, cloud infrastructure, or large-scale compute systems
- Excellent testing, debugging, documentation, and code review habits
Ways to stand out from the crowd
- Experience with GPU clusters, NVLink, InfiniBand, Ethernet fabrics, or HPC
- Hands-on work with Kubernetes scheduling, Slurm/Slinky topology, DRA, Kueue, Slinky, or device plugins
- Experience integrating with cloud provider topology APIs or cluster metadata systems
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services.
#LI-Remote
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until July 5, 2026.