Distinguished Engineer, Storage – AI Cloud

at Nvidia
USD 320,000-488,800 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API Ceph @ 6 Communication @ 6 Debugging @ 6 GPU @ 4 GitHub Go @ 7 HPC @ 4 InfiniBand @ 4 Linux @ 4 Networking @ 4 Observability Python @ 7 Rust @ 7 Security @ 6 Swift @ 6

Details

NVIDIA is seeking a Distinguished Engineer to lead its storage strategy for AI Cloud across the Neocloud Provider (NCP) and Cloud Service Provider (CSP) ecosystem. The role will define the architecture of high-performance parallel file systems, object stores, and block storage at exabyte scale for AI training and inference workloads across cloud, neocloud, and on-premises environments.

The position requires a hands-on technical leader who can collaborate with engineers, SREs, cloud providers, neocloud operators, storage vendors, and NVIDIA product organizations to establish the storage foundation for future AI infrastructure.

Responsibilities

  • Lead the multi-year technical plan for AI Cloud Storage expansion across NCPs, including reference architecture, capabilities, performance and durability SLOs, qualification methodology, and roadmap for file, object, and block storage.
  • Serve as chief storage architect with hands-on involvement in storage build reviews, production incident investigations, root-cause analysis, and prototype reference implementations.
  • Define production-readiness standards for NCP storage, including availability and durability SLOs, efficiency per TiB, observability, blast-radius containment, and reduced operational toil.
  • Establish GPU delivery gates requiring verification of storage-focused ancillary services before accepting GPU capacity.
  • Guide architecture across training, inference, and accelerated-computing product lines while coordinating with site reliability, operations, networking, and security teams.
  • Work with external cloud providers, neocloud operators, and storage vendors to align on a common architecture.
  • Establish an open-source strategy for AI storage, including GitHub-first and security-first practices, upstream community engagement, and APIs, SDKs, and protocols for partners and the broader industry.
  • Promote the use of AI coding and agentic tools across the storage organization by sharing patterns, prompts, and evaluation harnesses.
  • Automate infrastructure management tasks, including live software upgrades, node and drive replacements, capacity rebalancing, cross-data-center data movement, and dataset lifecycle management.
  • Design storage for future GPU generations and workloads such as disaggregated inference with storage-backed KV caching, write-once-read-many inference, exabyte-scale regional object stores, and cross-data-center dataset versioning and copy management.
  • Mentor senior, principal, and distinguished engineers and represent NVIDIA in standards bodies, open-source communities, customer briefings, and industry forums including FAST, SC, OCP, SNIA, and the Linux Storage Summit.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • At least 18 years of practical engineering experience in storage technology.
  • Extensive experience with a high-performance parallel file system such as Lustre, GPFS / Spectrum Scale, WEKA, VAST, BeeGFS, DAOS, or an equivalent technology at multi-petabyte scale.
  • Broad expertise in object storage, including S3- or Swift-class systems, and block storage, including NVMe-oF, NVMesh-class systems, or iSCSI.
  • Experience designing and managing storage platforms at exabyte scale for performance-critical AI training, HPC, video, or hyperscale data lake workloads.
  • Direct responsibility for durability, availability, and performance SLOs measured in 9s.
  • Demonstrated ability to establish technical strategy across business units and partner organizations, with measurable outcomes such as improved GPU utilization, reduced cost per petabyte, fewer incidents, or shorter bring-up times.
  • Hands-on engineering experience, including writing and reviewing production code, reading Lustre, NFS, kernel, NVMe-oF, or SPDK source code, and personally running scale tests or recovery drills.
  • Strong proficiency in at least one systems language: C, C++, Rust, or Go; proficiency in Python.
  • Experience with Linux kernel storage and networking stacks, including the block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, and multipath.
  • Daily use of advanced AI coding and autonomous tools for building, coding, debugging, validation, and operations.
  • Excellent written and verbal communication skills, including the ability to communicate technical trade-offs to executives, SREs, vendors, and internal customers.
  • Ability to operate in a 24/7 production environment where storage incidents directly affect GPU revenue, with a security-first approach.

Preferred Qualifications

  • Experience designing or managing AI training or inference storage at 10,000 or more GPU scale.
  • Open-source contributions or maintainership in Lustre, NFS, SPDK, NVMe / NVMe-oF, CSI, Ceph, MinIO, RocksDB, or related projects.
  • Experience building disaggregated-inference or inference-time-compute storage architectures, including KV caching, WORM storage, storage-aware scheduling, or database-integrated inference.
  • Public technical contributions such as patents, peer-reviewed papers, keynote talks, or RFCs related to AI infrastructure storage.

Benefits

NVIDIA offers equity, competitive benefits, and a comprehensive benefits package. The position is full time. Applications will be accepted at least until August 1, 2026.

More jobs at Nvidia

Similar jobs