Senior Site Reliability Engineer - Storage

at Nvidia
USD 168,000-333,500 per year
SENIOR
✅ On-site

Tech Stack

AWS @ 7 Azure @ 7 Bash @ 4 Cloud Computing Communication @ 6 Docker @ 4 ElasticSearch @ 4 GCP @ 7 Go Grafana @ 4 HPC @ 7 InfiniBand @ 4 Kibana @ 4 Kubernetes @ 4 Prometheus @ 4 Python @ 4 SRE Slurm @ 4 Splunk @ 4

Details

NVIDIA is looking for a Senior Site Reliability Engineer focused on HPC storage. The role involves designing, implementing, and optimizing on-premises High-Performance Computing (HPC) storage solutions while leveraging cloud computing. The engineer will develop distributed storage solutions and automation tools, ensure efficient IT operations, collaborate with engineering teams, document best practices, and support infrastructure requirements for major projects.

Responsibilities

  • Design and implement on-premises HPC infrastructure supplemented with cloud computing.
  • Design and implement scalable, efficient storage solutions for data-intensive applications, optimizing performance and cost-effectiveness.
  • Develop tooling to automate the deployment and management of large-scale infrastructure environments.
  • Automate operational monitoring and alerting and enable self-service resource consumption.
  • Document procedures and practices and perform technology evaluations related to distributed file systems.
  • Collaborate across teams to understand developer workflows and gather infrastructure requirements.
  • Guide methodologies for building, testing, and deploying applications to optimize performance and resource utilization.

Requirements

  • Bachelor's degree in Computer Science or equivalent experience with 8+ years of relevant experience, a master's degree with 5+ years of experience, or a Ph.D. with 3 years of experience.
  • 8+ years of experience developing technology solutions and resolving performance bottlenecks for HPC applications.
  • Experience designing, deploying, and managing enterprise NAS solutions such as NetApp and Pure Storage.
  • Experience with S3-based storage such as Cloudian and MinIO.
  • Experience with one or more parallel or distributed filesystems, such as Lustre or GPFS.
  • Python, Bash, or Golang programming and scripting experience.
  • Strong experience operating services in leading cloud environments such as AWS, Azure, or GCP.
  • Experience with monitoring stacks such as Prometheus and Grafana, Elasticsearch and Kibana, Splunk, or Zabbix.
  • Excellent communication and collaboration skills.

Preferred Qualifications

  • Experience with RDMA fabrics, including InfiniBand or RoCE.
  • Experience with HPC cluster management tools such as Slurm, PBS, or LSF.
  • Experience with containerization technologies such as Docker, Mesosphere DCOS, or Kubernetes.

Benefits

  • Equity and benefits are provided.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs