Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

at Nvidia
📍 Warsaw, Poland
PLN 221,200-507,000 per year
SENIOR
✅ On-site

Tech Stack

Ansible @ 6 Bash @ 7 CI/CD Deep Learning @ 4 Docker @ 4 GPU @ 4 Grafana HPC @ 6 IaC InfiniBand @ 4 JAX @ 4 Kubernetes @ 4 Linux @ 6 MLOps @ 6 Machine Learning NVLink Networking @ 4 Observability Prometheus PyTorch @ 4 Python @ 7 Slurm @ 4 Terraform @ 6

Details

NVIDIA's Deep Learning Frameworks (DLFW) Infrastructure team is looking for a deeply technical Senior HPC Cluster Administrator to lead the design, deployment, and reliability of large-scale GPU compute clusters. These systems support demanding deep learning training, inference, and high-performance computing workloads, from DGX/HGX platforms to Grace Blackwell systems. You will drive architectural decisions across compute, networking, and storage, and partner closely with software, research, and product teams to keep the infrastructure ahead of the workloads it supports.

Responsibilities

  • Own the full lifecycle of GPU compute clusters, including procurement, provisioning, configuration management, monitoring, and deprecation, across heterogeneous Linux environments such as DGX, HGX, and embedded systems.
  • Design and scale storage solutions using NFS, Lustre, WekaFS, or equivalent technologies, with a clear roadmap for capacity and performance growth.
  • Lead infrastructure automation using Ansible and Terraform, along with CI/CD pipelines using GitLab.
  • Manage and optimize job scheduling with Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies.
  • Maintain and improve observability stacks using Prometheus, Grafana, and DCGM, and drive proactive resolution of hardware and software incidents.
  • Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads.
  • Evaluate and introduce networking fabrics such as InfiniBand, NVLink, EFA, and RDMA, as well as storage tiers and container runtimes, to improve performance and reliability.
  • Mentor junior engineers and contribute to team-wide engineering standards.

Requirements

  • BS or MS in Computer Science, Electrical Engineering, Computer Engineering, or equivalent hands-on experience.
  • 5+ years of experience deploying and administering large-scale HPC or ML training clusters.
  • Deep expertise in Linux systems administration at scale.
  • Strong scripting and automation skills in Python and/or Bash.
  • Hands-on experience with Slurm, including scheduling, accounting, and cgroup configuration.
  • Proficiency with configuration management and infrastructure as code; Ansible is required and Terraform is a plus.
  • Experience with container technologies such as Docker, Apptainer/Singularity, and Kubernetes.
  • Solid understanding of high-speed networking, including InfiniBand, RoCE, RDMA, and EFA.
  • Experience with distributed or parallel filesystems and storage architecture.
  • Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders.

Additional Qualifications

  • Experience with NVIDIA GPU infrastructure tools such as DCGM, nvidia-smi, MIG, and NVSwitch diagnostics.
  • Familiarity with cluster management platforms such as Colossus, Bright Cluster Manager, xCAT, or similar tools.
  • Experience supporting large-scale distributed deep learning workloads using PyTorch, JAX, or Megatron.
  • Knowledge of BMC, IPMI, and Redfish for out-of-band management and hardware lifecycle operations.
  • Background in MLOps tooling or ML platform engineering.

Benefits

NVIDIA is committed to encouraging a collaborative and inclusive environment where every team member has the opportunity to thrive and make a significant impact.

Salary

For Poland, the base salary range is 221,250 PLN–383,500 PLN for Level 3, and 292,500 PLN–507,000 PLN for Level 4. The base salary is determined based on location, experience, and the pay of employees in similar positions.

More jobs at Nvidia

Similar jobs