Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 6
AWS @ 4
Ansible @ 4
Azure @ 4
Bash @ 6
Compliance
GCP
Grafana @ 6
HPC @ 4
Kubernetes @ 4
Leadership @ 4
Machine Learning
Networking
Observability
People Management @ 4
Performance Optimization @ 4
Prometheus @ 6
Python @ 6
Reporting @ 6
Security
Slurm @ 3
Stripe @ 4
Technical Leadership @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's IT Storage Engineering team architects, designs, deploys, and manages petabyte-scale storage infrastructure supporting demanding workloads across Hardware Engineering and Chip Design, Software Engineering, Manufacturing, AI/ML Research, and IT Operations. This role leads a team of storage engineers and collaborates across NVIDIA's global IT organization to scale storage infrastructure across on-premises data centers and cloud service providers.
Responsibilities
- Lead petabyte-scale storage deployments across on-premises data centers and AWS, Azure, and GCP, owning the lifecycle from design and procurement through physical installation, configuration, and production handoff.
- Engineer and maintain automation pipelines integrating storage systems with CMDB, configuration management, and observability platforms.
- Drive data center efficiency initiatives by optimizing power consumption, PUE impact, drive density, power shelf utilization, rack footprint, and workload placement.
- Develop self-service tools and dashboards for tracking storage capacity, consumption trends, and projected growth.
- Partner with networking, compute, cloud, and data center infrastructure teams to resolve cross-domain bottlenecks and support infrastructure capacity planning.
- Manage vendor relationships and lead hardware refresh and end-of-life planning across the storage portfolio.
- Recruit, mentor, and develop storage deployment engineers; establish engineering standards, runbooks, and on-call practices.
- Partner with architecture and security teams to evaluate storage technologies, lead proofs of concept, and develop production deployment standards.
- Own the hardware and data lifecycle, including procurement, rack integration, provisioning, expansion, decommissioning, and secure data destruction, while ensuring compliance, auditability, and zero unplanned data loss.
- Define data lifecycle management policies covering retention, tiering, archival, and deletion standards, improving storage efficiency and reducing costs associated with stale or redundant data.
Requirements
- Bachelor's or master's degree in Computer Science, Storage Systems, or a related technical field, or equivalent experience.
- 12 or more years of experience in large-scale storage architecture, operations, production engineering, or infrastructure.
- 6 or more years of people management or technical leadership experience with storage, infrastructure, or site reliability teams.
- Deep protocol-level knowledge of enterprise storage systems, including block storage and NVMe-oF, NFS, SMB/CIFS, S3-compatible object storage, GPFS/IBM Spectrum Scale, and Lustre.
- Experience deploying and operating storage solutions from multiple major vendors, including NetApp ONTAP and StorageGRID, Pure Storage FlashArray and FlashBlade, Cloudian HyperStore, and DDN EXAScaler and A³I.
- Strong understanding of storage hardware internals, including NVMe, SAS, SATA, QLC/TLC/MLC NAND, controller architectures, shelves and enclosures, cabling standards, and failure-domain planning.
- Working knowledge of bare-metal server hardware, including rack units, HBA/NIC selection, BMC/iDRAC/iLO, firmware management, ToR switches, fiber and copper interconnects, and SFP/QSFP optics.
- Expertise with Prometheus exporters, Grafana dashboards, and alerting rules for storage fleet health, performance SLOs, and capacity burn-rate tracking.
- Configuration management experience with Ansible, including playbook authoring, role design, inventory management, and Ansible Tower/AAP.
- Scripting and automation proficiency in Python and/or Bash, including REST API integrations for CMDB updates, provisioning workflows, and reporting pipelines.
- Experience managing large-scale, multi-vendor storage environments of 100 PB or more in high-availability production settings.
Preferred Qualifications
- Kubernetes storage integration experience, including CSI drivers, persistent volume lifecycle management, and StorageClass design.
- HPC storage experience, including parallel file system performance optimization, stripe tuning, client-side caching, OST/MDT balancing, and collaboration with HPC compute and fabric teams.
- Familiarity with NVMe, RDMA, DPUs, kernel-level I/O subsystems, and volumes.
- Familiarity with IBM LSF or Slurm and storage-aware scheduling policies.
- Experience building or scaling storage for AI/ML or HPC workloads across hybrid or multicloud environments, including AWS S3, Azure Blob, Google Cloud Storage, and on-premises infrastructure.
Benefits
- Eligibility for equity and benefits.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
- Applications will be accepted at least until August 31, 2026.
More jobs at Nvidia
Senior Systems Software Engineer, Low Latency Streaming Technology - Automotive
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Reinforcement Learning Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software and System Architect
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Customer Technical Program Manager
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Principal Technical Program Manager, Relational Deep Learning Platform
Nvidia · Santa Clara, United States
USD 240,000-379,500 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Infrastructure Security Engineer
SpaceXAI · Washington, United States, Austin, United States, New York City, United States, Palo Alto, United States
USD 100,000-258,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Engineer System Software, SDN Operations
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Manager, Storage Production Engineering
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Manager, Platform Operations
Collibra · United States
USD 168,000-210,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year