Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 7
Azure @ 7
Bash @ 4
Cloud Computing
Communication @ 6
Docker @ 4
ElasticSearch @ 4
GCP @ 7
Go
Grafana @ 4
HPC @ 7
InfiniBand @ 6
Kibana @ 4
Kubernetes @ 4
Prometheus @ 4
Python @ 4
SRE
Slurm @ 4
Splunk @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Site Reliability Engineer focused on HPC storage. The role involves designing, implementing, and optimizing on-premises High-Performance Computing (HPC) storage solutions supplemented by cloud computing. You will craft and deploy distributed storage solutions, build automation tools, ensure efficient IT operations, collaborate with engineering teams, document best practices, and support major artificial intelligence and HPC projects.
Responsibilities
- Design and implement on-premises HPC infrastructure supplemented by cloud computing.
- Design and implement scalable, efficient storage solutions for data-intensive applications, optimizing performance and cost-effectiveness.
- Develop tooling to automate the deployment and management of large-scale infrastructure environments.
- Automate operational monitoring and alerting and enable self-service consumption of resources.
- Document procedures and practices and perform technology evaluations related to distributed file systems.
- Collaborate across teams to understand developers' workflows and gather infrastructure requirements.
- Influence and guide methodologies for building, testing, and deploying applications to ensure optimal performance and resource utilization.
Requirements
- Bachelor's degree in Computer Science or equivalent experience with 8+ years of relevant experience; master's degree with 5+ years of experience; or Ph.D. with 3 years of experience.
- 8+ years of experience crafting technology solutions and resolving performance bottlenecks for HPC applications.
- Experience designing, deploying, and managing enterprise NAS solutions such as NetApp and Pure Storage.
- Experience with S3-based storage such as Cloudian or MinIO.
- Experience with one or more parallel or distributed file systems, such as Lustre or GPFS.
- Python, Bash, or Golang programming and scripting experience.
- Strong experience operating services in a leading cloud environment such as AWS, Azure, or GCP.
- Experience with monitoring stacks such as Prometheus and Grafana, Elasticsearch and Kibana, Splunk, or Zabbix.
- Familiarity with newer and emerging monitoring products.
- Excellent communication and collaboration skills.
Preferred Qualifications
- Background with RDMA fabrics, including InfiniBand or RoCE.
- Experience with HPC cluster management tools such as Slurm, PBS, or LSF.
- Experience with containerization technologies such as Docker, Mesosphere DCOS, or Kubernetes.
Compensation and Benefits
The base salary range is $168,000–$270,250 for Level 4 and $200,000–$322,000 for Level 5. Base salary is determined by location, experience, and compensation for employees in similar positions. The role also includes eligibility for equity and benefits. Applications will be accepted at least until August 14, 2026.