Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 7
Azure @ 7
Bash @ 4
Cloud Computing
Communication @ 6
Docker @ 4
ElasticSearch @ 4
GCP @ 7
Go
Grafana @ 4
HPC @ 7
InfiniBand @ 4
Kibana @ 4
Kubernetes @ 4
Prometheus @ 4
Python @ 4
SRE
Slurm @ 4
Splunk @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Site Reliability Engineer focused on HPC storage. The role involves designing, implementing, and optimizing on-premises High-Performance Computing (HPC) storage solutions while leveraging cloud computing. The engineer will develop distributed storage solutions and automation tools, ensure efficient IT operations, collaborate with engineering teams, document best practices, and support infrastructure requirements for major projects.
Responsibilities
- Design and implement on-premises HPC infrastructure supplemented with cloud computing.
- Design and implement scalable, efficient storage solutions for data-intensive applications, optimizing performance and cost-effectiveness.
- Develop tooling to automate the deployment and management of large-scale infrastructure environments.
- Automate operational monitoring and alerting and enable self-service resource consumption.
- Document procedures and practices and perform technology evaluations related to distributed file systems.
- Collaborate across teams to understand developer workflows and gather infrastructure requirements.
- Guide methodologies for building, testing, and deploying applications to optimize performance and resource utilization.
Requirements
- Bachelor's degree in Computer Science or equivalent experience with 8+ years of relevant experience, a master's degree with 5+ years of experience, or a Ph.D. with 3 years of experience.
- 8+ years of experience developing technology solutions and resolving performance bottlenecks for HPC applications.
- Experience designing, deploying, and managing enterprise NAS solutions such as NetApp and Pure Storage.
- Experience with S3-based storage such as Cloudian and MinIO.
- Experience with one or more parallel or distributed filesystems, such as Lustre or GPFS.
- Python, Bash, or Golang programming and scripting experience.
- Strong experience operating services in leading cloud environments such as AWS, Azure, or GCP.
- Experience with monitoring stacks such as Prometheus and Grafana, Elasticsearch and Kibana, Splunk, or Zabbix.
- Excellent communication and collaboration skills.
Preferred Qualifications
- Experience with RDMA fabrics, including InfiniBand or RoCE.
- Experience with HPC cluster management tools such as Slurm, PBS, or LSF.
- Experience with containerization technologies such as Docker, Mesosphere DCOS, or Kubernetes.
Benefits
- Equity and benefits are provided.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior NPI Program Manager
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
GPU PCIe and Boot Architect - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior AI Engineer, High Performance AI
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Salesforce CPQ Developer
Nvidia · Santa Clara, United States
USD 176,000-276,000 per year
Senior Technical Program Manager - LLM Safety
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Principal Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Backend Engineer - Databases - Loki Query
Grafana Labs · Germany, Spain, United Kingdom, Ireland, Sweden
GBP 91,000-1,114,000 per year
Senior Backend Engineer - Databases - Loki Query
Grafana Labs · Germany, Spain, United Kingdom, Ireland, Sweden
EUR 97,000-121,000 per year
Senior Backend Engineer - Databases - Loki Query
Grafana Labs · Germany, Spain, United Kingdom, Ireland, Sweden
SEK 775,000-969,000 per year
Senior Backend Engineer - Databases - Loki Query
Grafana Labs · Germany, Spain, United Kingdom, Ireland, Sweden
EUR 83,000-104,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Manager, Software Engineering - Agentic IT Operations
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year