Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Ansible @ 6
Bash @ 6
CUDA @ 6
CentOS @ 6
Communication @ 6
Debugging
Docker @ 4
GPU @ 4
Grafana @ 3
HPC @ 4
InfiniBand @ 3
Kubernetes @ 4
Linux @ 6
MPI @ 4
NCCL @ 4
Networking @ 3
Observability
Performance Analysis
Prometheus @ 3
Python @ 6
Slurm @ 4
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a highly skilled and experienced HPC Cluster Engineer to design, deploy, and operate GPU compute clusters for Electronic Design Automation (EDA) and high-performance computing workloads used across multiple teams and projects. The role involves collaborating with researchers and infrastructure teams to ensure GPU clusters are highly performant, scalable, and reliable.
Responsibilities
- Develop and enhance the ecosystem around GPU-accelerated computing, including scalable automation solutions.
- Continuously improve infrastructure provisioning, management, observability, and day-to-day operations through automation.
- Provide technical leadership and strategic guidance for managing large-scale HPC systems, including the deployment of compute, networking, and storage.
- Foster strong customer and cross-functional partnerships to ensure consistent cluster support and rapidly adapt to evolving user needs.
- Support researchers running EDA workloads, including performance analysis and optimization.
- Conduct root cause analysis, suggest corrective actions, and proactively identify and resolve issues.
- Build innovative tooling to accelerate researchers' productivity, debugging, and software performance at scale.
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- At least 5 years of experience building and operating large-scale compute infrastructure, including cluster configuration management tools such as BCM or Ansible.
- Experience with AI/HPC job schedulers and orchestrators such as Slurm, LSF, PBS, or Kubernetes.
- Applied experience with AI/HPC workflows using MPI and NCCL.
- Proficiency with Linux, including Rocky Linux, CentOS, RHEL, and/or Ubuntu distributions.
- Solid understanding of container technologies such as Enroot and Docker.
- Proficiency in Python and Bash.
- Experience analyzing and tuning performance for a variety of EDA workloads.
- Excellent problem-solving skills, including the ability to analyze complex systems, identify bottlenecks, and implement scalable solutions.
- Excellent communication and collaboration skills.
- Passion for continual learning and staying current with new technologies and effective approaches in HPC infrastructure.
Preferred Qualifications
- Background with NVIDIA GPUs, CUDA programming, NCCL, and MLPerf benchmarking.
- Experience supporting EDA workloads and tools.
- Familiarity with high-speed HPC networking, including InfiniBand, RDMA, and RoCE.
- Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC workloads.
- Familiarity with metrics collection and visualization at scale using Prometheus, OpenSearch, and Grafana.
Compensation and Benefits
The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. The role also includes eligibility for equity and benefits.
Applications for this job will be accepted at least until June 19, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior System Software Engineer - GPU Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior HPC Performance Engineer
Nvidia · Germany
PLN 221,200-507,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Nvidia · Warsaw, Poland
PLN 221,200-507,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year