Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 7
Debugging
GPU @ 4
Go @ 7
HPC @ 4
InfiniBand @ 3
Kubernetes @ 6
Linux @ 3
Networking @ 3
Performance Analysis @ 4
Performance Optimization @ 4
Python @ 6
Rust @ 7
Security @ 6
Swift @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's DGX Cloud Storage team builds storage platforms that support tens of thousands of accelerators, maintain exabytes of data, and power large-scale AI workloads across cloud, neocloud, and on-premises environments. The role involves foundational storage engineering for AI infrastructure as a hands-on individual contributor and technical lead.
Responsibilities
- Contribute code, fixes, and features to open-source parallel and distributed file systems and distributed object storage, while engaging with upstream communities and maintainers.
- Write and review production code and examine kernel, NFS, NVMe-oF, or SPDK source code to diagnose bugs.
- Make technical decisions for storage deliveries against measurable performance and durability targets.
- Triage, troubleshoot, and determine the root cause of storage issues across GPU clusters with tens of thousands of GPUs, including I/O and metadata performance problems, data corruption, and recovery incidents.
- Validate storage architecture, capabilities, performance, and durability through scale tests, benchmarks, recovery drills, and qualification of new builds.
- Define configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure.
- Help operators and internal customers apply storage configuration and tuning standards.
- Collaborate with training, inference, accelerated-computing, site-reliability, operations, networking, and security teams, as well as cloud providers, neocloud operators, and storage vendors.
- Use modern AI coding and agentic tools for development, debugging, validation, and operations.
Requirements
- BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- More than 12 years of direct storage software engineering experience, including extensive experience with a high-performance parallel or distributed file system operating at multi-petabyte scale.
- Contributions to open-source projects involving distributed or parallel file systems.
- Hands-on experience writing and reviewing production code, examining file system, kernel, NVMe-oF, or SPDK source code, and conducting scale tests or recovery drills.
- Experience diagnosing and resolving storage issues in large GPU or HPC clusters, including I/O and metadata performance analysis.
- Strong proficiency in at least one systems language: C, C++, Rust, or Go.
- Proficiency in Python.
- Familiarity with Linux kernel storage and networking stacks, including the block layer, RDMA/RoCE/InfiniBand, NVMe, page cache, VFS, and multipath.
- Understanding of object storage such as S3 or Swift-class systems and block storage such as NVMe-oF and iSCSI.
- Strong written and verbal communication skills, with the ability to explain complex technical trade-offs to engineers, SREs, vendors, and internal customers.
- Comfort working in a 24/7 production environment where storage incidents directly affect GPU availability, with a security-first approach.
- 100% hands-on engineering, including personally writing and reviewing production code, investigating source code, and running scale tests or recovery drills.
Preferred Qualifications
- Maintainer status or sustained contributions to widely used public projects.
- Experience designing or operating storage for AI training or inference at very large GPU scale, with measurable improvements in GPU utilization or reductions in I/O bottlenecks.
- Kernel and file system development experience, including metadata scalability, data placement, failure recovery, or HSM or equivalent technologies.
- Kubernetes and CSI driver development for storage.
- Hands-on experience with SPDK, libfabric, or FUSE performance optimization.
Benefits
- Equity and benefits are provided.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Applications Software Engineer, Perception and Sensor Fusion
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Engineer, Security Architecture - DGX Cloud
Nvidia · United States
USD 272,000-431,200 per year
Senior System Software Engineer – Linux-Tegra Power Management and Performance Optimization
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
PhD Research Intern, Learning Embodied Skills from Human Data - 2027
Nvidia · Santa Clara, United States
USD 38-94 per hour
Senior Synthetic Data Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Distinguished Engineer, Storage – AI Cloud
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, Fleet Intelligence Agent Systems
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Principal Software Engineer - Rack-Scale Systems Infrastructure
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year