Service Reliability Engineer

at Nvidia
USD 168,000-333,500 per year
SENIOR
✅ Remote

Tech Stack

AI AWS @ 1 Ansible @ 7 Azure @ 1 Communication @ 7 GCP @ 1 GPU @ 3 Git @ 4 Grafana @ 6 HPC Jira @ 6 Kubernetes @ 7 Linux @ 6 Mathematics @ 4 Mentoring @ 7 Networking @ 7 OpenTelemetry @ 6 Python @ 7 SRE Security Slurm @ 7 System Administration @ 6

Details

NVIDIA's Compute Infrastructure Support team is seeking a driven Site Reliability Engineer to support on-premises and cloud products and services. The role focuses on maintaining near-100% availability across systems supporting AI and accelerated computing.

Responsibilities

  • Work within a 24/7 follow-the-sun support model across multiple continents, collaborating with a U.S.-based manager.
  • Work a four-day, 10-hour schedule, including either Saturday or Sunday, with flexible early or late shifts.
  • Monitor and manage large-scale production GPU and Kubernetes environments to maintain availability and performance.
  • Use advanced tools to proactively detect, prevent, and respond to incidents.
  • Analyze logs, metrics, and system behavior to diagnose issues and implement effective resolutions.
  • Develop predictive automated support routines to prevent production issues.
  • Improve automation by integrating incident analysis into auto-healing and automated break-fix solutions.
  • Perform systems administration, network administration, and security monitoring.
  • Coordinate with domain experts and service owners to resolve complex issues.
  • Continuously improve service quality and operational processes based on incident feedback.
  • Coordinate effectively across teams during incident resolution and provide strong customer-focused support.

Requirements

  • Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.
  • Familiarity with GPU hardware and high-performance computing environments.
  • Proficiency with Grafana, OpenTelemetry, PagerDuty, and JIRA.
  • Experience with AWS, Azure, GCP, or OCI is a plus; strong on-premises expertise is preferred.
  • At least 8 years of experience coordinating large-scale production systems, including more than 3 years in high-availability Internet, cloud, or data center environments.
  • Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or equivalent experience.
  • Expert-level Linux system administration skills.
  • Experience with Ansible and/or Python automation and strong shell scripting skills.
  • Strong knowledge of DNS, DHCP, storage systems, and core networking.
  • Experience troubleshooting and maintaining large-scale bare-metal infrastructure.
  • Strong partnership, documentation, communication, and mentoring skills.

Preferred Qualifications

  • Experience with scripting languages, particularly Python.
  • Experience running virtual machines under community-supported or commercial hypervisors.
  • Knowledge of application containers and container orchestration systems.
  • Basic understanding of Git.
  • Ability to master and maintain complex environments.

Compensation and Benefits

  • Level 4 base salary: USD 168,000–270,250 per year.
  • Level 5 base salary: USD 208,000–333,500 per year.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until August 14, 2026.

NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs