Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
CI/CD @ 4
CUDA @ 6
Debugging @ 4
Distributed Systems @ 4
Docker @ 6
GPU @ 6
GitHub @ 4
Go @ 7
Grafana @ 4
HPC @ 4
Helm @ 4
Jenkins @ 4
Kubernetes @ 7
Linux @ 6
Node.js @ 4
Prometheus @ 4
Python @ 4
React @ 4
Rust @ 4
Slurm @ 4
Software Development @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for outstanding software engineers to help expand its enterprise GPU management and monitoring tools. The role involves designing and building cloud-native management agents, Kubernetes integrations, and end-to-end integration solutions that combine GPUs with datacenter software management ecosystems. The work supports NVIDIA products across HPC, cloud, and enterprise environments on bare-metal and virtualized platforms.
Contributions will span GPU system integration areas including telemetry and metrics, health checks, diagnostics, configuration, and system management. These tools support both passive background monitoring and active online management, with an emphasis on operational transparency and seamless integration in customer environments. The code will support single-node developer systems through large clusters with thousands of nodes.
Responsibilities
- Develop and maintain distributed, robust, and scalable Go programs deployed to Kubernetes environments that manage large datacenters.
- Develop and maintain user-space applications, containers, Go bindings, and CLI tools.
- Enable GPU management integration with open-source ecosystems including Kubernetes and Docker.
- Support internal and external users through bug fixes, documentation, and feature improvements.
- Maintain high-quality products through robust test coverage.
Requirements
- Bachelor's degree or higher in Computer Science or equivalent experience.
- Five or more years of meaningful industry experience with a strong Go and Kubernetes development background.
- User-space development and debugging expertise in Linux environments.
- Experience with APIs and interface design.
- Outstanding written and verbal interpersonal skills, including business-level English.
- Strong motivation and commitment to learning new skills.
- Ability to execute all aspects of the software development lifecycle and manage time in a fast-paced, heavily multitasked environment.
- Development experience with Rust, Python, and/or C/C++.
- Experience with distributed systems and concurrent applications, especially in Kubernetes environments.
- Experience developing and maintaining enterprise software.
- Experience deploying, managing, and debugging applications in Kubernetes environments.
Preferred Qualifications
- Background with containers such as Docker and OCI, orchestration frameworks, and logging or telemetry backends.
- Experience with Kubernetes monitoring stacks and tools such as Prometheus, Loki, and Grafana.
- Experience with modern UI development in React and Node.js or similar frameworks.
- Experience developing Kubernetes operators or Helm charts.
- Experience with HPC job schedulers such as Slurm or Run:ai.
- Familiarity with Kubernetes internals.
- Exposure to GPU programming with CUDA.
- Experience with Jenkins and GitHub or GitLab CI/CD pipelines.
Compensation and Benefits
The base salary is determined based on location, experience, and the pay of employees in similar positions. The base salary range is $152,000–$241,500 for Level 3 and $184,000–$287,500 for Level 4. Employees are also eligible for equity and benefits.
Applications will be accepted at least until July 30, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.