HPC Operations Engineer

at Nvidia
USD 124,000-241,500 per year
MIDDLE
✅ Hybrid

Tech Stack

Ansible @ 3 Bash @ 5 CentOS @ 5 Communication @ 3 Docker @ 3 HPC InfiniBand @ 2 Linux @ 5 Networking @ 3 Observability Perl @ 5 Python @ 5 SRE Slurm @ 2

Details

NVIDIA's Hardware Infrastructure Farm team designs and operates compute clusters that power silicon development. The role focuses on building and operating high-reliability, efficient, and high-performance HPC infrastructure while driving foundational improvements and automation to improve engineering productivity. The team applies SRE practices such as reducing reactive operational work, conducting blameless postmortems, and proactively identifying potential outages.

Responsibilities

  • Troubleshoot incoming support requests in a large-scale HPC environment.
  • Enhance deployment automation, configuration management, observability, operational monitoring, and day-to-day operations through automation.
  • Ensure compute servers are running the correct operating system and configuration.
  • Troubleshoot complex issues from bare metal through the application level to ensure system reliability and efficiency.
  • Collaborate with specialist teams to drive issues to closure.
  • Work with domain experts to improve how the chip development process uses infrastructure.
  • Contribute to overall quality and improve time to market for next-generation chips.

Requirements

  • Proficiency administering CentOS and RHEL Linux distributions.
  • Understanding of container technologies such as Docker.
  • Proficiency in Python and UNIX scripting languages such as Bash.
  • Excellent problem-solving skills, including the ability to analyze complex systems, identify bottlenecks, and implement scalable solutions.
  • Excellent communication and teamwork skills.
  • Bachelor's degree in Computer Science or a similar field, or equivalent experience, with 2 or more years of relevant post-degree experience.
  • Solid understanding of cluster configuration management tools such as Ansible.

Preferred Qualifications

  • Understanding of Linux technologies including NFS, automounter, LDAP, DNS, and TCP/IP networking in Red Hat Linux distributions.
  • Familiarity with job scheduler administration, such as IBM Spectrum LSF or SLURM, and experience building or operating large-scale compute infrastructure.
  • Knowledge of the FlexLM license management system.
  • Proficiency in Perl for maintaining legacy automation scripts.
  • Familiarity with high-speed networking technologies such as InfiniBand, RDMA, and RoCE.
  • Familiarity with fast, distributed storage systems such as Lustre and GPFS.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is committed to a diverse work environment and is an equal opportunity employer.

Compensation

  • Base salary range for Level 2: USD 124,000–195,500 per year.
  • Base salary range for Level 3: USD 152,000–241,500 per year.
  • The base salary is determined based on location, experience, and the pay of employees in similar positions.

The position is full time and hybrid. Applications will be accepted at least until April 12, 2026.

More jobs at Nvidia

Similar jobs