Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
at Nvidia
PLN 221,200-507,000 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Ansible @ 6
Bash @ 7
CI/CD
Deep Learning @ 4
Docker @ 4
GPU @ 4
Grafana
HPC @ 6
IaC @ 6
InfiniBand @ 4
JAX @ 4
Kubernetes @ 4
Linux @ 6
MLOps
Machine Learning
NVLink
Networking @ 4
Observability
Prometheus
PyTorch @ 4
Python @ 7
Slurm @ 4
Terraform @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Responsibilities
- Own the full lifecycle of GPU compute clusters — procurement, provisioning, configuration management, monitoring, and deprecation — across heterogeneous Linux environments (DGX, HGX, embedded systems)
- Design and scale storage solutions (NFS, Lustre, WekaFS, or equivalent) with a clear roadmap for capacity and performance growth
- Lead automation of infrastructure using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab)
- Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies
- Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents
- Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads
- Evaluate and introduce new technologies — networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, container runtimes — to improve performance and reliability
- Mentor junior engineers and contribute to team-wide engineering standards
Requirements
- BS/MS in CS, EE, CE, or equivalent hands-on experience
- 5+ years of experience deploying and administering large-scale HPC or ML training clusters
- Deep expertise in Linux systems administration at scale
- Strong scripting and automation skills in Python and/or bash
- Hands-on experience with Slurm (scheduling, accounting, cgroup configuration)
- Proficiency with configuration management and IaC (Ansible required; Terraform a plus)
- Experience with container technologies (Docker, Apptainer/Singularity, Kubernetes)
- Solid understanding of high-speed networking (InfiniBand, RoCE, RDMA, EFA)
- Experience with distributed/parallel filesystems and storage architecture
- Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders
Ways to stand out from the crowd
- Experience with NVIDIA GPU infrastructure tools (DCGM, nvidia-smi, MIG, NVSwitch diagnostics)
- Familiarity with cluster management platforms (Colossus, Bright Cluster Manager, xCAT, or similar)
- Experience supporting large-scale distributed deep learning workloads (PyTorch, JAX, Megatron)
- Knowledge of BMC/IPMI/Redfish for out-of-band management and hardware lifecycle
- Background in MLOps tooling or ML platform engineering
Join our team of world-class engineers and be part of the groundbreaking work we do at NVIDIA. We are committed to encouraging a collaborative and inclusive environment, where every team member has the opportunity to thrive and make a significant impact!
More jobs at Nvidia
Senior System Software Engineer - Halos Core And Robotics Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Director, Autonomous Vehicles Platform
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior Developer Technology Engineer - Edge Agentic Ai
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - Halos
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Drive Os Communication Infrastructure
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer - Datacenter Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
AI/Ml Specialist Solutions Architect
Nebius · United States, Canada
USD 250,000-320,000 per year
Solutions Architect
Nebius · United States, Canada
USD 250,000-320,000 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior SRE Engineer
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Senior HPC Performance Engineer
Nvidia · Germany
PLN 221,200-507,000 per year