Senior Site Reliability Engineer - HPC

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI AWS CI/CD @ 6 Communication @ 7 Debugging @ 7 GCP Go @ 6 HPC @ 4 IaC Kubernetes @ 4 Leadership @ 7 Mentoring @ 7 Observability @ 7 Perl @ 6 Python @ 6 Ruby @ 6 SRE @ 6 Slurm @ 4

Details

NVIDIA is looking for a Senior Site Reliability Engineer to join its Compute Farm team and help build the next generation of its global services platform. The role focuses on keeping critically important systems running and using AI to deliver reliable solutions for high-performance computing and accelerated computing environments.

Responsibilities

  • Own SRE solutions end to end, from design and implementation through operation and continuous improvement, ensuring integration with HPC schedulers, storage, and network fabrics.
  • Use Infrastructure as Code and configuration management to standardize and automate provisioning.
  • Deliver solutions in a globally distributed, multi-cloud hybrid environment spanning on-premises infrastructure, AWS, GCP, and OCI.
  • Design for failure using redundancy, failure domains, progressive delivery, and strict change control.
  • Ensure high uptime and Quality of Service for internal customers through operational excellence.
  • Conduct capacity management and planning to meet ongoing operational needs.
  • Detect performance issues and recommend solutions to maintain service quality.
  • Collaborate with multiple teams in a fast-paced environment to ensure seamless project completion.
  • Participate in on-call rotations and incident reviews, assist with root-cause identification, and produce high-quality root-cause analysis reports.

Requirements

  • Bachelor's degree in Computer Science or a related technical field, or equivalent experience, with 5+ years of professional experience building and supporting critical services.
  • Experience supporting large-scale HPC clusters using Slurm, LSF, or Kubernetes, including setup, tuning, and troubleshooting.
  • Proficiency in modern CI/CD techniques and Infrastructure as Code for managing services.
  • Strong experience building large-scale infrastructure platforms for automated host lifecycle management, fleet reliability and auto-healing, end-to-end observability, or data-driven operations using AIOps or machine-learning-driven signals.
  • Proficiency with monitoring, metrics, container management, and log collection tools.
  • 5+ years of coding or scripting experience in at least two high-level programming languages, such as Python, Go, Perl, or Ruby.
  • Experience mentoring engineers and influencing technical direction through design reviews, architecture documents, and strong partnerships with product and leadership teams.
  • Creative problem-solving abilities, excellent debugging skills, and strong communication and documentation skills.

Preferred Qualifications

  • Published technical write-ups or talks, such as conference presentations, meetups, or engineering blogs, covering real-world reliability, observability, large-scale HPC, or SRE problems and solutions.
  • Maintainer or co-maintainer experience for a production open-source component, such as a plugin, operator, exporter, controller, or SDK, used at large scale.

Benefits

  • Equity and benefits are provided.
  • NVIDIA offers a comprehensive benefits package.

Applications for this job will be accepted at least until June 19, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs