Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
Ceph @ 6
Communication @ 6
Debugging @ 6
GPU @ 4
GitHub
Go @ 7
HPC @ 4
InfiniBand @ 4
Linux @ 4
Networking @ 4
Observability
Python @ 7
Rust @ 7
Security @ 6
Swift @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Distinguished Engineer to lead its storage strategy for AI Cloud across the Neocloud Provider (NCP) and Cloud Service Provider (CSP) ecosystem. The role will define the architecture of high-performance parallel file systems, object stores, and block storage at exabyte scale for AI training and inference workloads across cloud, neocloud, and on-premises environments.
The position requires a hands-on technical leader who can collaborate with engineers, SREs, cloud providers, neocloud operators, storage vendors, and NVIDIA product organizations to establish the storage foundation for future AI infrastructure.
Responsibilities
- Lead the multi-year technical plan for AI Cloud Storage expansion across NCPs, including reference architecture, capabilities, performance and durability SLOs, qualification methodology, and roadmap for file, object, and block storage.
- Serve as chief storage architect with hands-on involvement in storage build reviews, production incident investigations, root-cause analysis, and prototype reference implementations.
- Define production-readiness standards for NCP storage, including availability and durability SLOs, efficiency per TiB, observability, blast-radius containment, and reduced operational toil.
- Establish GPU delivery gates requiring verification of storage-focused ancillary services before accepting GPU capacity.
- Guide architecture across training, inference, and accelerated-computing product lines while coordinating with site reliability, operations, networking, and security teams.
- Work with external cloud providers, neocloud operators, and storage vendors to align on a common architecture.
- Establish an open-source strategy for AI storage, including GitHub-first and security-first practices, upstream community engagement, and APIs, SDKs, and protocols for partners and the broader industry.
- Promote the use of AI coding and agentic tools across the storage organization by sharing patterns, prompts, and evaluation harnesses.
- Automate infrastructure management tasks, including live software upgrades, node and drive replacements, capacity rebalancing, cross-data-center data movement, and dataset lifecycle management.
- Design storage for future GPU generations and workloads such as disaggregated inference with storage-backed KV caching, write-once-read-many inference, exabyte-scale regional object stores, and cross-data-center dataset versioning and copy management.
- Mentor senior, principal, and distinguished engineers and represent NVIDIA in standards bodies, open-source communities, customer briefings, and industry forums including FAST, SC, OCP, SNIA, and the Linux Storage Summit.
Requirements
- BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- At least 18 years of practical engineering experience in storage technology.
- Extensive experience with a high-performance parallel file system such as Lustre, GPFS / Spectrum Scale, WEKA, VAST, BeeGFS, DAOS, or an equivalent technology at multi-petabyte scale.
- Broad expertise in object storage, including S3- or Swift-class systems, and block storage, including NVMe-oF, NVMesh-class systems, or iSCSI.
- Experience designing and managing storage platforms at exabyte scale for performance-critical AI training, HPC, video, or hyperscale data lake workloads.
- Direct responsibility for durability, availability, and performance SLOs measured in 9s.
- Demonstrated ability to establish technical strategy across business units and partner organizations, with measurable outcomes such as improved GPU utilization, reduced cost per petabyte, fewer incidents, or shorter bring-up times.
- Hands-on engineering experience, including writing and reviewing production code, reading Lustre, NFS, kernel, NVMe-oF, or SPDK source code, and personally running scale tests or recovery drills.
- Strong proficiency in at least one systems language: C, C++, Rust, or Go; proficiency in Python.
- Experience with Linux kernel storage and networking stacks, including the block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, and multipath.
- Daily use of advanced AI coding and autonomous tools for building, coding, debugging, validation, and operations.
- Excellent written and verbal communication skills, including the ability to communicate technical trade-offs to executives, SREs, vendors, and internal customers.
- Ability to operate in a 24/7 production environment where storage incidents directly affect GPU revenue, with a security-first approach.
Preferred Qualifications
- Experience designing or managing AI training or inference storage at 10,000 or more GPU scale.
- Open-source contributions or maintainership in Lustre, NFS, SPDK, NVMe / NVMe-oF, CSI, Ceph, MinIO, RocksDB, or related projects.
- Experience building disaggregated-inference or inference-time-compute storage architectures, including KV caching, WORM storage, storage-aware scheduling, or database-integrated inference.
- Public technical contributions such as patents, peer-reviewed papers, keynote talks, or RFCs related to AI infrastructure storage.
Benefits
NVIDIA offers equity, competitive benefits, and a comprehensive benefits package. The position is full time. Applications will be accepted at least until August 1, 2026.