Senior Storage Software Engineer - DGX Cloud

at Nvidia
USD 224,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Communication @ 7 Debugging GPU @ 4 Go @ 7 HPC @ 4 InfiniBand @ 3 Kubernetes @ 6 Linux @ 3 Networking @ 3 Performance Analysis @ 4 Performance Optimization @ 4 Python @ 6 Rust @ 7 Security @ 6 Swift @ 4

Details

NVIDIA's DGX Cloud Storage team builds storage platforms that support tens of thousands of accelerators, maintain exabytes of data, and power large-scale AI workloads across cloud, neocloud, and on-premises environments. The role involves foundational storage engineering for AI infrastructure as a hands-on individual contributor and technical lead.

Responsibilities

  • Contribute code, fixes, and features to open-source parallel and distributed file systems and distributed object storage, while engaging with upstream communities and maintainers.
  • Write and review production code and examine kernel, NFS, NVMe-oF, or SPDK source code to diagnose bugs.
  • Make technical decisions for storage deliveries against measurable performance and durability targets.
  • Triage, troubleshoot, and determine the root cause of storage issues across GPU clusters with tens of thousands of GPUs, including I/O and metadata performance problems, data corruption, and recovery incidents.
  • Validate storage architecture, capabilities, performance, and durability through scale tests, benchmarks, recovery drills, and qualification of new builds.
  • Define configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure.
  • Help operators and internal customers apply storage configuration and tuning standards.
  • Collaborate with training, inference, accelerated-computing, site-reliability, operations, networking, and security teams, as well as cloud providers, neocloud operators, and storage vendors.
  • Use modern AI coding and agentic tools for development, debugging, validation, and operations.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • More than 12 years of direct storage software engineering experience, including extensive experience with a high-performance parallel or distributed file system operating at multi-petabyte scale.
  • Contributions to open-source projects involving distributed or parallel file systems.
  • Hands-on experience writing and reviewing production code, examining file system, kernel, NVMe-oF, or SPDK source code, and conducting scale tests or recovery drills.
  • Experience diagnosing and resolving storage issues in large GPU or HPC clusters, including I/O and metadata performance analysis.
  • Strong proficiency in at least one systems language: C, C++, Rust, or Go.
  • Proficiency in Python.
  • Familiarity with Linux kernel storage and networking stacks, including the block layer, RDMA/RoCE/InfiniBand, NVMe, page cache, VFS, and multipath.
  • Understanding of object storage such as S3 or Swift-class systems and block storage such as NVMe-oF and iSCSI.
  • Strong written and verbal communication skills, with the ability to explain complex technical trade-offs to engineers, SREs, vendors, and internal customers.
  • Comfort working in a 24/7 production environment where storage incidents directly affect GPU availability, with a security-first approach.
  • 100% hands-on engineering, including personally writing and reviewing production code, investigating source code, and running scale tests or recovery drills.

Preferred Qualifications

  • Maintainer status or sustained contributions to widely used public projects.
  • Experience designing or operating storage for AI training or inference at very large GPU scale, with measurable improvements in GPU utilization or reductions in I/O bottlenecks.
  • Kernel and file system development experience, including metadata scalability, data placement, failure recovery, or HSM or equivalent technologies.
  • Kubernetes and CSI driver development for storage.
  • Hands-on experience with SPDK, libfabric, or FUSE performance optimization.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs