Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
AWS @ 4
ArgoCD @ 4
Azure @ 4
Bash @ 6
Change Management
Communication @ 6
Datadog @ 4
Debugging @ 4
GCP @ 4
GPU
GitHub @ 4
GitHub Actions @ 4
Go @ 6
Kubernetes @ 7
LLM @ 4
Microservices @ 7
Observability
Prometheus @ 4
Python @ 6
SRE @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Site Reliability Engineer to join its GeForce NOW team. The SRE team ensures that internal and external-facing GPU cloud gaming services meet reliability and uptime commitments while enabling developers to make carefully planned changes. The role focuses on service response and workflows, developing tools and services, and maintaining and improving service-level objectives (SLOs).
Responsibilities
- Build tools to improve SRE observability.
- Participate in the Kubernetes migration journey, including VMI setup and problem-solving.
- Rapidly debug and triage incidents and user-reported issues.
- Automate, script, and improve tooling to help achieve 100% automation of daily tasks.
- Support services before launch through system design consulting, software platform and framework development, capacity management, and launch reviews.
- Participate in an on-call rotation supporting production systems.
- Partner with service owners to improve service reliability.
- Lead production improvements involving change management, post-mortem reviews, workflow processes, and software automation.
Requirements
- MS or BS in Computer Science, Engineering, a related field, or equivalent experience.
- 8+ years of site reliability engineering experience working with large-scale distributed microservices in production, with a strong focus on automation and tooling.
- Very strong Kubernetes experience, including complex, highly available VMI setups on Kubernetes.
- Proven problem-solving and root-cause analysis skills, with a focus on optimization and efficiency.
- Experience with Datadog, Prometheus, Alertmanager, or similar monitoring systems.
- Experience managing multi-region cloud deployments on hyperscalers such as AWS, GCP, or Azure.
- Experience designing and managing deployment pipelines using GitHub Actions, GitLab CI, ArgoCD, or similar tools.
- Excellent communication, presentation, social, and analytical skills, including the ability to explain complex interactions clearly to different audiences.
- Production-grade coding proficiency in Go, Python, or robust Bash scripting.
- Primary production on-call experience responding to and mitigating high-severity infrastructure alerts and service degradations.
Preferred Qualifications
- Experience with automated anomaly detection, log clustering tools, or LLM-assisted debugging platforms.
- Comfort using AI on a day-to-day basis as an SRE.
- Prior experience as an SRE or Service Engineer.
Compensation and Benefits
- Base salary range: USD 168,000–270,250 per year.
- Eligible for equity and benefits.
Applications will be accepted at least until August 15, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
NVIDIA 2027 Internships: Ph.D. Research Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 38-94 per hour
NVIDIA 2027 Internships: Ph.D. Research Computer Vision and Deep Learning
Nvidia · Santa Clara, United States
USD 38-94 per hour
NVIDIA Spring 2027 Internships: Developer and Performance Technology
Nvidia · Santa Clara, United States
USD 20-71 per hour
Deep Learning Compiler Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
NVIDIA 2027 Internships: Ph.D. Research Computer Architecture and Systems
Nvidia · Santa Clara, United States
USD 38-94 per hour
Similar jobs
Staff Backend Engineer - Mimir Query, Databases | USA | Remote
Grafana Labs · United States
USD 175,000-210,000 per year
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Operations Engineer, BizTech
Airbnb · United States
USD 136,000-160,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Backend Engineer - Mimir Query, Databases
Grafana Labs · Canada
CAD 186,400-223,600 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Senior Engineer System Software, SDN Operations
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year