Senior HPC Storage Engineer

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 7 Bash @ 4 CUDA @ 6 CentOS @ 6 Ceph @ 4 Deep Learning @ 4 Docker @ 4 GPU HPC @ 7 Linux @ 6 NCCL @ 6 Networking @ 7 Performance Analysis PyTorch @ 4 Python @ 4 TensorFlow @ 4

Details

NVIDIA's HW Infrastructure Storage Strategy team provides leadership in researching, designing, and implementing fast storage solutions for demanding high-performance computing and computationally intensive workloads. The role focuses on architectural changes across file, block, and object storage to support the scaling and performance requirements of expanding cloud infrastructure, including storage design for large-scale workloads, private and public cloud strategy, capacity modeling, and growth planning across a global computing environment.

Responsibilities

  • Research and analyze existing internal distributed storage services.
  • Research, design, and implement scalable next-generation distributed storage services for HPC workloads, optimizing performance and cost-effectiveness.
  • Develop tooling to automate the management of large-scale infrastructure environments, operational monitoring, alerting, and self-service resource consumption.
  • Define procedures and practices and perform technology evaluations related to distributed file systems.
  • Collaborate across teams to understand developers' workflows and capture infrastructure requirements.
  • Influence and guide methodologies for building, testing, and deploying applications to improve performance and resource utilization.
  • Support researchers running workloads on clusters, including performance analysis and optimization of deep learning workflows.
  • Perform root-cause analysis and recommend corrective actions for problems at various scales.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, a related field, or equivalent experience.
  • 8+ years of experience designing and/or operating large-scale storage infrastructure.
  • Experience analyzing and tuning storage performance for a variety of workloads.
  • Proficiency with CentOS/RHEL and/or Ubuntu Linux distributions.
  • Python programming and Bash scripting experience.
  • In-depth understanding of container technologies such as Docker and Enroot.

Preferred Qualifications

  • Extensive experience with parallel and distributed file systems such as Ceph, Weka.io, VAST, Lustre, and GPFS, as well as Linux storage kernel development.
  • Proficiency with NVIDIA GPUs, CUDA programming, and NCCL, including performance benchmarking with MLPerf.
  • Familiarity with storage hardware including HDDs, SSDs, NVMe, enclosures, and specialized appliances such as Network Appliance.
  • Strong background in software-defined networking and high-performance networking for AI/HPC clusters.
  • Practical experience with deep learning frameworks, specifically PyTorch and TensorFlow.

Benefits

NVIDIA offers competitive salaries, equity, and a comprehensive benefits package. The company is committed to an inclusive work environment and is an equal opportunity employer. Applications will be accepted at least until June 13, 2026.

More jobs at Nvidia

Similar jobs