Senior HPC Cluster Engineer

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Ansible @ 6 Bash @ 6 CUDA @ 6 CentOS @ 6 Communication @ 6 Debugging Docker @ 4 GPU @ 4 Grafana @ 3 HPC @ 4 InfiniBand @ 3 Kubernetes @ 4 Linux @ 6 MPI @ 4 NCCL @ 4 Networking @ 3 Observability Performance Analysis Prometheus @ 3 Python @ 6 Slurm @ 4 Technical Leadership

Details

NVIDIA is seeking a highly skilled and experienced HPC Cluster Engineer to design, deploy, and operate GPU compute clusters for Electronic Design Automation (EDA) and high-performance computing workloads used across multiple teams and projects. The role involves collaborating with researchers and infrastructure teams to ensure GPU clusters are highly performant, scalable, and reliable.

Responsibilities

  • Develop and enhance the ecosystem around GPU-accelerated computing, including scalable automation solutions.
  • Continuously improve infrastructure provisioning, management, observability, and day-to-day operations through automation.
  • Provide technical leadership and strategic guidance for managing large-scale HPC systems, including the deployment of compute, networking, and storage.
  • Foster strong customer and cross-functional partnerships to ensure consistent cluster support and rapidly adapt to evolving user needs.
  • Support researchers running EDA workloads, including performance analysis and optimization.
  • Conduct root cause analysis, suggest corrective actions, and proactively identify and resolve issues.
  • Build innovative tooling to accelerate researchers' productivity, debugging, and software performance at scale.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • At least 5 years of experience building and operating large-scale compute infrastructure, including cluster configuration management tools such as BCM or Ansible.
  • Experience with AI/HPC job schedulers and orchestrators such as Slurm, LSF, PBS, or Kubernetes.
  • Applied experience with AI/HPC workflows using MPI and NCCL.
  • Proficiency with Linux, including Rocky Linux, CentOS, RHEL, and/or Ubuntu distributions.
  • Solid understanding of container technologies such as Enroot and Docker.
  • Proficiency in Python and Bash.
  • Experience analyzing and tuning performance for a variety of EDA workloads.
  • Excellent problem-solving skills, including the ability to analyze complex systems, identify bottlenecks, and implement scalable solutions.
  • Excellent communication and collaboration skills.
  • Passion for continual learning and staying current with new technologies and effective approaches in HPC infrastructure.

Preferred Qualifications

  • Background with NVIDIA GPUs, CUDA programming, NCCL, and MLPerf benchmarking.
  • Experience supporting EDA workloads and tools.
  • Familiarity with high-speed HPC networking, including InfiniBand, RDMA, and RoCE.
  • Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC workloads.
  • Familiarity with metrics collection and visualization at scale using Prometheus, OpenSearch, and Grafana.

Compensation and Benefits

The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. The role also includes eligibility for equity and benefits.

Applications for this job will be accepted at least until June 19, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs