Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
GPU
HPC
InfiniBand @ 6
Kubernetes @ 4
Linux @ 4
Networking @ 4
Python @ 6
SRE
Slurm
Software Development @ 7
System Administration @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA Base Command Manager powers thousands of clusters worldwide, ranging from a few to several thousand nodes, and streamlines cluster provisioning, workload management, and infrastructure monitoring. The role will contribute to the success of external customers running NVIDIA solutions and internal clusters used for research, operations, and next-generation projects.
Responsibilities
- Contribute to deployments and daily operations of large-scale, next-generation GPU platforms.
- Handle incidents in GPU clusters, bridging the gap between cluster operations and development.
- Design and implement small features in the Base Command Manager product to develop an in-depth understanding of the product.
- Validate complex cluster configurations, including Slurm and Kubernetes orchestrators, for performance, scalability, and resilience against real-world customer scenarios.
Requirements
- Bachelor's degree or equivalent experience in Computer Science or a related field.
- 8+ years of experience in site reliability engineering and/or software development roles.
- Fluency in Python.
- In-depth knowledge of Linux and networking.
Preferred Qualifications
- Experience with C++, high-performance computing, Kubernetes, and/or system administration.
- Previous experience as a system administrator running BCM, Bright Cluster Manager, or Base Command Manager clusters.
- Proficiency with cluster networking, including InfiniBand and Spectrum-X.
Benefits
- Equity.
- NVIDIA benefits.
- Inclusive and equal-opportunity work environment.
Applications will be accepted at least until August 21, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Localization and Planning Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Customer Success Insights Engineer
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Relational Foundation Model Engineer, Modern Data Stack
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Developer Experience
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Insider Threat Detection Engineer
Nvidia · United States
USD 168,000-310,500 per year
Similar jobs
Service Reliability Engineer
Nvidia · United States
USD 168,000-333,500 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior HPC Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Engineer System Software, SDN Operations
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Deep Learning Software Infrastructure Engineer
Nvidia · United States
USD 224,000-431,200 per year