Senior HPC AI Cluster Engineer

at Nvidia
USD 176,000-333,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 AWS @ 3 Ansible @ 4 Azure @ 3 Bash @ 4 CUDA @ 4 CentOS @ 8 Chef @ 4 Cloud Computing @ 3 Deep Learning GPU @ 4 HPC @ 4 InfiniBand @ 7 Jenkins @ 4 Kubernetes @ 4 Linux @ 8 Networking @ 4 Puppet @ 4 Python @ 4 Security @ 8 Slurm @ 4

Details

NVIDIA is seeking an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. The team builds supercomputers and AI clusters using advanced computing hardware and software. The role focuses on at-scale system design, performance tuning, accelerated computing, deep learning platforms, and large-scale performance environments. You will collaborate with HPC, operating system, GPU computing, and systems specialists to architect, develop, and bring up large-scale platforms.

Responsibilities

  • Design, implement, and maintain large-scale HPC and AI clusters with monitoring, logging, and alerting.
  • Manage Linux job and workload scheduling and orchestration tools.
  • Develop and maintain continuous integration and delivery pipelines.
  • Develop tooling to automate deployment and management of large-scale infrastructure environments.
  • Automate operational monitoring and alerting and enable self-service resource consumption.
  • Deploy monitoring solutions for servers, networks, and storage.
  • Troubleshoot from the bare-metal and operating-system levels through the software stack and application level.
  • Develop, redefine, and document standard methodologies for internal teams.
  • Support research and development activities and participate in proofs of concept and proofs of value for future improvements.

Requirements

  • Degree in Computer Science, Engineering, or a related field, or equivalent experience, plus 8 or more years of experience.
  • Knowledge of HPC and AI solution technologies, including CPUs, GPUs, high-speed interconnects, and supporting software.
  • Experience with job workload scheduling and orchestration tools such as Slurm and Kubernetes.
  • Excellent knowledge of Windows and Linux, including Red Hat/CentOS and Ubuntu; networking sockets, firewalld, iptables, Wireshark, networking internals, ACLs, operating-system security, TCP, DHCP, and DNS.
  • Experience with storage solutions such as Lustre, GPFS, and Weka.io, along with familiarity with emerging storage technologies.
  • Python programming and Bash scripting experience.
  • Experience with automation and configuration management tools such as Jenkins, Ansible, Puppet, and Chef.
  • Deep knowledge of networking protocols including InfiniBand and Ethernet.
  • Deep understanding of virtual systems such as VMware, Hyper-V, KVM, or Citrix.
  • Familiarity with cloud computing platforms such as AWS, Azure, and Google Cloud.

Preferred Qualifications

  • Knowledge of CPU and/or GPU architecture.
  • Knowledge of Kubernetes and container-related microservice technologies.
  • Experience with GPU-focused hardware and software, including DGX and CUDA.
  • Experience with RDMA fabrics, including InfiniBand or RoCE.

Compensation and Benefits

The base salary depends on location, experience, and compensation for employees in similar positions. The base salary range is USD 176,000–276,000 for Level 4 and USD 208,000–333,500 for Level 5. The position is also eligible for equity and benefits.

Applications will be accepted at least until August 24, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs