Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Ansible @ 6
Bash @ 7
CI/CD
Deep Learning @ 4
Docker @ 4
GPU @ 4
Grafana
HPC @ 6
IaC
InfiniBand @ 4
JAX @ 4
Kubernetes @ 4
Linux @ 6
MLOps @ 6
Machine Learning
NVLink
Networking @ 4
Observability
Prometheus
PyTorch @ 4
Python @ 7
Slurm @ 4
Terraform @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Deep Learning Frameworks (DLFW) Infrastructure team is looking for a deeply technical Senior HPC Cluster Administrator to lead the design, deployment, and reliability of large-scale GPU compute clusters. These systems support demanding deep learning training, inference, and high-performance computing workloads, from DGX/HGX platforms to Grace Blackwell systems. You will drive architectural decisions across compute, networking, and storage, and partner closely with software, research, and product teams to keep the infrastructure ahead of the workloads it supports.
Responsibilities
- Own the full lifecycle of GPU compute clusters, including procurement, provisioning, configuration management, monitoring, and deprecation, across heterogeneous Linux environments such as DGX, HGX, and embedded systems.
- Design and scale storage solutions using NFS, Lustre, WekaFS, or equivalent technologies, with a clear roadmap for capacity and performance growth.
- Lead infrastructure automation using Ansible and Terraform, along with CI/CD pipelines using GitLab.
- Manage and optimize job scheduling with Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies.
- Maintain and improve observability stacks using Prometheus, Grafana, and DCGM, and drive proactive resolution of hardware and software incidents.
- Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads.
- Evaluate and introduce networking fabrics such as InfiniBand, NVLink, EFA, and RDMA, as well as storage tiers and container runtimes, to improve performance and reliability.
- Mentor junior engineers and contribute to team-wide engineering standards.
Requirements
- BS or MS in Computer Science, Electrical Engineering, Computer Engineering, or equivalent hands-on experience.
- 5+ years of experience deploying and administering large-scale HPC or ML training clusters.
- Deep expertise in Linux systems administration at scale.
- Strong scripting and automation skills in Python and/or Bash.
- Hands-on experience with Slurm, including scheduling, accounting, and cgroup configuration.
- Proficiency with configuration management and infrastructure as code; Ansible is required and Terraform is a plus.
- Experience with container technologies such as Docker, Apptainer/Singularity, and Kubernetes.
- Solid understanding of high-speed networking, including InfiniBand, RoCE, RDMA, and EFA.
- Experience with distributed or parallel filesystems and storage architecture.
- Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders.
Additional Qualifications
- Experience with NVIDIA GPU infrastructure tools such as DCGM, nvidia-smi, MIG, and NVSwitch diagnostics.
- Familiarity with cluster management platforms such as Colossus, Bright Cluster Manager, xCAT, or similar tools.
- Experience supporting large-scale distributed deep learning workloads using PyTorch, JAX, or Megatron.
- Knowledge of BMC, IPMI, and Redfish for out-of-band management and hardware lifecycle operations.
- Background in MLOps tooling or ML platform engineering.
Benefits
NVIDIA is committed to encouraging a collaborative and inclusive environment where every team member has the opportunity to thrive and make a significant impact.
Salary
For Poland, the base salary range is 221,250 PLN–383,500 PLN for Level 3, and 292,500 PLN–507,000 PLN for Level 4. The base salary is determined based on location, experience, and the pay of employees in similar positions.