Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 4
ArgoCD @ 7
Bash @ 6
Data Structures @ 4
Distributed Systems @ 4
GPU
Go @ 4
HTTP @ 4
Java
Kubernetes @ 7
Linux @ 4
Python @ 6
Rust @ 4
Security @ 4
gRPC @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA Cloud Functions (NVCF) is an open-source platform that links workloads to GPUs. It enables teams to deploy, manage, and serve GPU-accelerated, containerized applications across regions and clusters worldwide. The platform routes inference, streaming, and batch jobs across decentralized GPU clusters, allowing endpoints to scale consistently on-premises or in the cloud.
The team is seeking a Senior Systems Software Engineer to improve the performance, reliability, and scaling behavior of a system that routes AI workloads onto distributed GPU fleets. The role involves working on a polyglot platform with control-plane and edge deployments, using systems performance engineering, distributed systems, and Kubernetes-based runtimes.
Responsibilities
- Develop GPU- and DPU-accelerated applications and make them easier to develop, deploy, and monitor on NVIDIA hardware.
- Design and ship services in Java, Go, and Rust in a public open-source repository.
- Automate and optimize build, test, integration, and release processes for cloud-native systems.
- Collaborate with engineering teams across NVIDIA to integrate with technologies including KAI Scheduler, NVIDIA NIM, Grove, and Dynamo.
- Help steward the open-source project by triaging community issues and pull requests and writing contributor documentation.
Requirements
- Bachelor's or Master's degree in Computer Science or equivalent experience.
- At least 3 years of hands-on software engineering experience.
- Expert-level knowledge of a systems programming language such as Go, C, or Rust.
- Proven understanding of data structures, algorithms, and distributed software architecture.
- Strong understanding of Kubernetes and container technologies, with hands-on automation experience using continuous integration frameworks such as GitLab and ArgoCD.
- Expertise in a scripting language such as Bash or Python.
- Knowledge of and experience with Unix or Unix-like kernel internals, including Linux.
- Understanding of performance, security, and reliability in complex distributed systems.
Preferred Qualifications
- Background with publish-subscribe models and message queues.
- Experience optimizing high-throughput network paths.
- Working understanding of unary, streaming, and bidirectional protocols across HTTP/2 and gRPC.
- Experience developing Kubernetes Custom Resources and Operators deployed with cloud service providers.
Benefits
The role includes equity and benefits. NVIDIA offers a competitive salary and benefits package.
The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. Applications will be accepted at least until July 28, 2026.