Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 4
Ansible @ 4
Azure @ 4
Ceph @ 4
Data Pipelines
DevOps
GPU
HPC @ 4
IaC
Kubernetes @ 4
Leadership @ 7
Machine Learning
Mentoring @ 6
Networking @ 4
Observability @ 7
People Management @ 7
Prometheus @ 7
Puppet @ 4
SRE @ 4
Technical Leadership @ 7
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is harnessing the possibilities of AI to build the next era of computing, with GPUs powering computers, robots, and self-driving cars. Its storage platforms support advanced AI and high-performance computing workloads. In this role, you will lead the team responsible for keeping storage systems fast, reliable, and ready to scale. You will work closely with engineering and AI teams, help shape how data moves through platforms, and guide decisions on new storage technologies.
Responsibilities
- Lead and coach a team of Storage Production Engineers, creating a collaborative, inclusive, and learning-focused environment.
- Design, deploy, and improve large-scale storage systems, including distributed storage, parallel file systems, and object storage.
- Use automation, monitoring, and analytics to make storage services more reliable, efficient, and easier to operate.
- Own capacity planning, data lifecycle management, cost awareness, and high-availability and disaster-recovery plans for storage.
- Evaluate and adopt modern storage approaches such as NVMe over Fabrics, RDMA, high-speed interconnects, and cloud-based storage.
- Guide incident response and root-cause analysis for storage issues, and implement changes to prevent recurring problems.
- Partner with engineering, DevOps, and AI/ML teams to improve data pipelines, access patterns, and workflow performance.
Requirements
- Bachelor's or master's degree in Computer Science, Storage Systems, or a related technical field, or equivalent experience.
- At least 12 years of experience in large-scale storage architecture, operations, production engineering, or infrastructure.
- At least 6 years of people management or technical leadership experience with storage, infrastructure, or site reliability teams.
- Direct experience managing infrastructure operations, including on-call rotations, incident response, ongoing maintenance, troubleshooting, production-system optimization, SLOs, and operational KPIs.
- Hands-on experience with parallel file systems such as Lustre or GPFS; distributed storage such as Ceph or MinIO; and enterprise object or NAS platforms such as S3-compatible systems, NetApp, or Pure Storage.
- Strong knowledge of block, file, and object storage, including performance tuning, data protection, and high-availability design.
- Experience with storage networking and protocols including NFS, SMB, iSCSI, Fibre Channel, RDMA, and NVMe-oF.
- Practical experience with automation and infrastructure as code using tools such as Terraform, Ansible, or Puppet.
- Strong knowledge of monitoring and observability tools such as Prometheus, InfluxDB, or the Elastic Stack, as well as logging and alerting.
Preferred Qualifications
- Experience managing and scaling SRE or Production Engineering teams in large-scale, mission-critical environments, with a focus on operational excellence, service availability, and performance.
- A track record of improving reliability, simplicity, and day-to-day operations for large, business-critical storage systems.
- Experience building or scaling storage for AI/ML or HPC workloads, including hybrid or multi-cloud environments such as AWS S3, Azure Blob, or Google Cloud Storage, as well as on-premises infrastructure.
- Experience with software-defined storage, cloud-native storage, and Kubernetes-based storage orchestration.
- A passion for mentoring, career development, and building a supportive, high-performing team culture.
Benefits
NVIDIA offers a comprehensive benefits package that may include medical, dental, and vision coverage; mental health resources; retirement and 401(k) plans; an employee stock purchase plan; paid time off and holidays; family and caregiving leave; and wellness and development programs. Specific benefits vary by location. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until August 2, 2026. NVIDIA is an equal opportunity employer and is committed to fostering an inclusive work environment.