Senior Site Reliability Engineer, GeForce NOW

at Nvidia
USD 168,000-270,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 6 AWS @ 4 ArgoCD @ 4 Azure @ 4 Bash @ 6 Change Management Communication @ 6 Datadog @ 4 Debugging @ 4 GCP @ 4 GPU GitHub @ 4 GitHub Actions @ 4 Go @ 6 Kubernetes @ 7 LLM @ 4 Microservices @ 7 Observability Prometheus @ 4 Python @ 6 SRE @ 4

Details

NVIDIA is looking for a Senior Site Reliability Engineer to join its GeForce NOW team. The SRE team ensures that internal and external-facing GPU cloud gaming services meet reliability and uptime commitments while enabling developers to make carefully planned changes. The role focuses on service response and workflows, developing tools and services, and maintaining and improving service-level objectives (SLOs).

Responsibilities

  • Build tools to improve SRE observability.
  • Participate in the Kubernetes migration journey, including VMI setup and problem-solving.
  • Rapidly debug and triage incidents and user-reported issues.
  • Automate, script, and improve tooling to help achieve 100% automation of daily tasks.
  • Support services before launch through system design consulting, software platform and framework development, capacity management, and launch reviews.
  • Participate in an on-call rotation supporting production systems.
  • Partner with service owners to improve service reliability.
  • Lead production improvements involving change management, post-mortem reviews, workflow processes, and software automation.

Requirements

  • MS or BS in Computer Science, Engineering, a related field, or equivalent experience.
  • 8+ years of site reliability engineering experience working with large-scale distributed microservices in production, with a strong focus on automation and tooling.
  • Very strong Kubernetes experience, including complex, highly available VMI setups on Kubernetes.
  • Proven problem-solving and root-cause analysis skills, with a focus on optimization and efficiency.
  • Experience with Datadog, Prometheus, Alertmanager, or similar monitoring systems.
  • Experience managing multi-region cloud deployments on hyperscalers such as AWS, GCP, or Azure.
  • Experience designing and managing deployment pipelines using GitHub Actions, GitLab CI, ArgoCD, or similar tools.
  • Excellent communication, presentation, social, and analytical skills, including the ability to explain complex interactions clearly to different audiences.
  • Production-grade coding proficiency in Go, Python, or robust Bash scripting.
  • Primary production on-call experience responding to and mitigating high-severity infrastructure alerts and service degradations.

Preferred Qualifications

  • Experience with automated anomaly detection, log clustering tools, or LLM-assisted debugging platforms.
  • Comfort using AI on a day-to-day basis as an SRE.
  • Prior experience as an SRE or Service Engineer.

Compensation and Benefits

  • Base salary range: USD 168,000–270,250 per year.
  • Eligible for equity and benefits.

Applications will be accepted at least until August 15, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs