Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
Bash @ 4
CUDA @ 6
CentOS @ 6
Ceph @ 4
Deep Learning @ 4
Docker @ 4
GPU
HPC @ 7
Linux @ 6
NCCL @ 6
Networking @ 7
Performance Analysis
PyTorch @ 4
Python @ 4
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's HW Infrastructure Storage Strategy team provides leadership in researching, designing, and implementing fast storage solutions for demanding high-performance computing and computationally intensive workloads. The role focuses on architectural changes across file, block, and object storage to support the scaling and performance requirements of expanding cloud infrastructure, including storage design for large-scale workloads, private and public cloud strategy, capacity modeling, and growth planning across a global computing environment.
Responsibilities
- Research and analyze existing internal distributed storage services.
- Research, design, and implement scalable next-generation distributed storage services for HPC workloads, optimizing performance and cost-effectiveness.
- Develop tooling to automate the management of large-scale infrastructure environments, operational monitoring, alerting, and self-service resource consumption.
- Define procedures and practices and perform technology evaluations related to distributed file systems.
- Collaborate across teams to understand developers' workflows and capture infrastructure requirements.
- Influence and guide methodologies for building, testing, and deploying applications to improve performance and resource utilization.
- Support researchers running workloads on clusters, including performance analysis and optimization of deep learning workflows.
- Perform root-cause analysis and recommend corrective actions for problems at various scales.
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, a related field, or equivalent experience.
- 8+ years of experience designing and/or operating large-scale storage infrastructure.
- Experience analyzing and tuning storage performance for a variety of workloads.
- Proficiency with CentOS/RHEL and/or Ubuntu Linux distributions.
- Python programming and Bash scripting experience.
- In-depth understanding of container technologies such as Docker and Enroot.
Preferred Qualifications
- Extensive experience with parallel and distributed file systems such as Ceph, Weka.io, VAST, Lustre, and GPFS, as well as Linux storage kernel development.
- Proficiency with NVIDIA GPUs, CUDA programming, and NCCL, including performance benchmarking with MLPerf.
- Familiarity with storage hardware including HDDs, SSDs, NVMe, enclosures, and specialized appliances such as Network Appliance.
- Strong background in software-defined networking and high-performance networking for AI/HPC clusters.
- Practical experience with deep learning frameworks, specifically PyTorch and TensorFlow.
Benefits
NVIDIA offers competitive salaries, equity, and a comprehensive benefits package. The company is committed to an inclusive work environment and is an equal opportunity employer. Applications will be accepted at least until June 13, 2026.