Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Ansible @ 3
Bash @ 5
CentOS @ 5
Communication @ 3
Docker @ 3
HPC
InfiniBand @ 2
Linux @ 5
Networking @ 3
Observability
Perl @ 5
Python @ 5
SRE
Slurm @ 2
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Hardware Infrastructure Farm team designs and operates compute clusters that power silicon development. The role focuses on building and operating high-reliability, efficient, and high-performance HPC infrastructure while driving foundational improvements and automation to improve engineering productivity. The team applies SRE practices such as reducing reactive operational work, conducting blameless postmortems, and proactively identifying potential outages.
Responsibilities
- Troubleshoot incoming support requests in a large-scale HPC environment.
- Enhance deployment automation, configuration management, observability, operational monitoring, and day-to-day operations through automation.
- Ensure compute servers are running the correct operating system and configuration.
- Troubleshoot complex issues from bare metal through the application level to ensure system reliability and efficiency.
- Collaborate with specialist teams to drive issues to closure.
- Work with domain experts to improve how the chip development process uses infrastructure.
- Contribute to overall quality and improve time to market for next-generation chips.
Requirements
- Proficiency administering CentOS and RHEL Linux distributions.
- Understanding of container technologies such as Docker.
- Proficiency in Python and UNIX scripting languages such as Bash.
- Excellent problem-solving skills, including the ability to analyze complex systems, identify bottlenecks, and implement scalable solutions.
- Excellent communication and teamwork skills.
- Bachelor's degree in Computer Science or a similar field, or equivalent experience, with 2 or more years of relevant post-degree experience.
- Solid understanding of cluster configuration management tools such as Ansible.
Preferred Qualifications
- Understanding of Linux technologies including NFS, automounter, LDAP, DNS, and TCP/IP networking in Red Hat Linux distributions.
- Familiarity with job scheduler administration, such as IBM Spectrum LSF or SLURM, and experience building or operating large-scale compute infrastructure.
- Knowledge of the FlexLM license management system.
- Proficiency in Perl for maintaining legacy automation scripts.
- Familiarity with high-speed networking technologies such as InfiniBand, RDMA, and RoCE.
- Familiarity with fast, distributed storage systems such as Lustre and GPFS.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to a diverse work environment and is an equal opportunity employer.
Compensation
- Base salary range for Level 2: USD 124,000–195,500 per year.
- Base salary range for Level 3: USD 152,000–241,500 per year.
- The base salary is determined based on location, experience, and the pay of employees in similar positions.
The position is full time and hybrid. Applications will be accepted at least until April 12, 2026.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior HPC Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Nvidia · Warsaw, Poland
PLN 221,200-507,000 per year
Senior HPC AI Cluster Engineer
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, Infrastructure Automation and Distributed Systems
Nvidia · United States
USD 224,000-431,200 per year
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year