Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Change Management
Communication @ 4
Leadership @ 7
Mathematics @ 4
Observability @ 4
SRE @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build AI-powered systems that enhance reliability, speed, and employee experience. The position will lead Incident, Problem, and Change Management into an intelligent, automated operating model using observability, AI insights, and orchestration, moving from reactive processes to predictive and autonomous operations.
Responsibilities
- Manage the full lifecycle of Incident, Problem, and CM as a 24×7 operational function, ensuring high reliability and minimal business disruption.
- Transform incident response by bringing to bear AI detection, correlation, and guided remediation, reducing time to detect, respond, and resolve.
- Build and scale intelligent incident workflows that integrate monitoring, telemetry, and service context to enable faster and more consistent response.
- Evolve Problem Management into a data-driven field, using AI and analytics to identify patterns, eliminate recurring issues, and drive systemic fixes.
- Modernize CM by introducing risk-aware, data-driven decisioning, improving change success rates, and reducing blast radius.
- Drive the adoption of observability as a foundation, ensuring service-level visibility, signal quality, and actionable insights across the IT ecosystem.
- Lead the development of automation and orchestration platforms that reduce manual effort across the outage lifecycle, including detection, triage, communication, and RCA or equivalent experience.
- Partner closely with engineering, infrastructure, and business teams to align operations with service reliability goals and SLOs.
Requirements
- BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience).
- 5+ years of experience leading and managing global IT operations or service management teams, with growing scope and complexity.
- 12+ overall years of experience in Site Reliability Engineering, IT Service Management, with a focus on Incident Management, Problem Management, and Configuration Management.
- Proven proficiency in Incident, Problem, and CM with a consistent record of delivering measurable gains in reliability and efficiency.
- Demonstrated experience applying AI, automation, or advanced analytics to improve operational outcomes.
- Solid understanding of observability, monitoring ecosystems, and modern reliability practices (SRE principles, SLOs, error budgets).
- Demonstrated ability to move organizations from process-heavy to technology-focused operating models.
- Strong leadership capability with experience building and scaling engineering-focused teams (SRE, SWE, or equivalent).
- Ability to deliver executive-level communication and insights, translating operational signals into clear, actionable narratives for leadership.
- Ability to build and lead a high-performing team of SREs and engineers, encouraging a culture of ownership, innovation, and continuous improvement.
Benefits
- Eligible for equity and benefits.
Base salary range is 200,000 USD - 322,000 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
More jobs at Nvidia
Senior Software Engineer, Compute Sanitizer - GPU
Nvidia · United States
USD 184,000-356,500 per year
Senior Backend Platform Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Applied AI Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - AV Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Systems Software Engineer, Accelerated Kubernetes Performance And Scale - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Similar jobs
Senior Technical Program Manager
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Technical Program Manager
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior Site Reliability Engineer - HPC
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Data Analyst, Macro Analytics
Stripe · South San Francisco, United States, New York City, United States, Seattle, United States
USD 134,000-242,400 per year
Staff Full Stack Engineer, Identity
Stripe · South San Francisco, United States, New York City, United States, Seattle, United States
USD 224,000-336,000 per year
Technical Program Manager, Compute
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 290,000-365,000 per year
Staff Software Engineer, AI Reliability
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 325,000-485,000 per year