Senior Manager, Storage Engineering

at Nvidia
USD 248,000-396,800 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API @ 6 AWS @ 4 Ansible @ 4 Azure @ 4 Bash @ 6 Compliance GCP Grafana @ 6 HPC @ 4 Kubernetes @ 4 Leadership @ 4 Machine Learning Networking Observability People Management @ 4 Performance Optimization @ 4 Prometheus @ 6 Python @ 6 Reporting @ 6 Security Slurm @ 3 Stripe @ 4 Technical Leadership @ 4

Details

NVIDIA's IT Storage Engineering team architects, designs, deploys, and manages petabyte-scale storage infrastructure supporting demanding workloads across Hardware Engineering and Chip Design, Software Engineering, Manufacturing, AI/ML Research, and IT Operations. This role leads a team of storage engineers and collaborates across NVIDIA's global IT organization to scale storage infrastructure across on-premises data centers and cloud service providers.

Responsibilities

  • Lead petabyte-scale storage deployments across on-premises data centers and AWS, Azure, and GCP, owning the lifecycle from design and procurement through physical installation, configuration, and production handoff.
  • Engineer and maintain automation pipelines integrating storage systems with CMDB, configuration management, and observability platforms.
  • Drive data center efficiency initiatives by optimizing power consumption, PUE impact, drive density, power shelf utilization, rack footprint, and workload placement.
  • Develop self-service tools and dashboards for tracking storage capacity, consumption trends, and projected growth.
  • Partner with networking, compute, cloud, and data center infrastructure teams to resolve cross-domain bottlenecks and support infrastructure capacity planning.
  • Manage vendor relationships and lead hardware refresh and end-of-life planning across the storage portfolio.
  • Recruit, mentor, and develop storage deployment engineers; establish engineering standards, runbooks, and on-call practices.
  • Partner with architecture and security teams to evaluate storage technologies, lead proofs of concept, and develop production deployment standards.
  • Own the hardware and data lifecycle, including procurement, rack integration, provisioning, expansion, decommissioning, and secure data destruction, while ensuring compliance, auditability, and zero unplanned data loss.
  • Define data lifecycle management policies covering retention, tiering, archival, and deletion standards, improving storage efficiency and reducing costs associated with stale or redundant data.

Requirements

  • Bachelor's or master's degree in Computer Science, Storage Systems, or a related technical field, or equivalent experience.
  • 12 or more years of experience in large-scale storage architecture, operations, production engineering, or infrastructure.
  • 6 or more years of people management or technical leadership experience with storage, infrastructure, or site reliability teams.
  • Deep protocol-level knowledge of enterprise storage systems, including block storage and NVMe-oF, NFS, SMB/CIFS, S3-compatible object storage, GPFS/IBM Spectrum Scale, and Lustre.
  • Experience deploying and operating storage solutions from multiple major vendors, including NetApp ONTAP and StorageGRID, Pure Storage FlashArray and FlashBlade, Cloudian HyperStore, and DDN EXAScaler and A³I.
  • Strong understanding of storage hardware internals, including NVMe, SAS, SATA, QLC/TLC/MLC NAND, controller architectures, shelves and enclosures, cabling standards, and failure-domain planning.
  • Working knowledge of bare-metal server hardware, including rack units, HBA/NIC selection, BMC/iDRAC/iLO, firmware management, ToR switches, fiber and copper interconnects, and SFP/QSFP optics.
  • Expertise with Prometheus exporters, Grafana dashboards, and alerting rules for storage fleet health, performance SLOs, and capacity burn-rate tracking.
  • Configuration management experience with Ansible, including playbook authoring, role design, inventory management, and Ansible Tower/AAP.
  • Scripting and automation proficiency in Python and/or Bash, including REST API integrations for CMDB updates, provisioning workflows, and reporting pipelines.
  • Experience managing large-scale, multi-vendor storage environments of 100 PB or more in high-availability production settings.

Preferred Qualifications

  • Kubernetes storage integration experience, including CSI drivers, persistent volume lifecycle management, and StorageClass design.
  • HPC storage experience, including parallel file system performance optimization, stripe tuning, client-side caching, OST/MDT balancing, and collaboration with HPC compute and fabric teams.
  • Familiarity with NVMe, RDMA, DPUs, kernel-level I/O subsystems, and volumes.
  • Familiarity with IBM LSF or Slurm and storage-aware scheduling policies.
  • Experience building or scaling storage for AI/ML or HPC workloads across hybrid or multicloud environments, including AWS S3, Azure Blob, Google Cloud Storage, and on-premises infrastructure.

Benefits

  • Eligibility for equity and benefits.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.
  • Applications will be accepted at least until August 31, 2026.

More jobs at Nvidia

Similar jobs