Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 7
Customer Support
DevOps @ 7
Distributed Systems @ 6
GenAI
Generative AI @ 4
Go @ 4
Grafana @ 4
LLM @ 4
Observability @ 4
Perl @ 4
Prometheus @ 4
Python @ 4
Ruby @ 4
SRE @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
At NVIDIA, Site Reliability Engineering provides an opportunity to define, develop, and support large-scale production systems with high efficiency and availability. This role combines software and systems engineering to ensure reliable service operation, consistent uptime, and efficient system performance. The SRE will work collaboratively to enable developers to make significant updates while maintaining reliable system operations.
Responsibilities
- Develop and support guidelines for incident management, planned maintenance, and blameless postmortems.
- Assist teams with high-severity incidents, drive root cause analysis, create high-quality postmortems, and develop post-incident corrective actions.
- Define reliability and supportability metrics, Service Level Objectives (SLOs), and error budgets.
- Develop and drive adoption of actionable, customer-centric monitoring and alerting.
- Apply automation and Generative AI/Agentic solutions to minimize manual and repetitive activities and improve customer support.
- Guide teams in establishing sustainable on-call and operational standards.
Requirements
- Degree in Computer Science or a related technical field involving coding, or equivalent experience.
- 8+ years of experience in SRE, DevOps, or Production Engineering.
- Strong understanding of SRE principles, including incident management, error budgets, SLOs, and Service Level Agreements (SLAs).
- Experience designing and deploying fault-tolerant, performant, and supportable systems.
- Background in infrastructure automation.
- Experience operating critical services in production.
- Experience with one or more of Python, Go, Perl, or Ruby.
- Hands-on experience with observability platforms such as Prometheus and Grafana.
- Strong communication skills and the ability to explain technical concepts effectively to diverse audiences.
- Flexibility and adaptability in a fast-paced environment with evolving requirements.
Preferred Qualifications
- Expertise in establishing incident management and postmortem processes.
- Experience driving adoption of common tools and processes across diverse groups.
- Experience working with LLM, Generative AI, or Agentic solutions to shorten mitigation time, reduce toil, and ensure SLOs are met.
- Hands-on expertise operating and scaling distributed systems with tight SLAs while ensuring high availability and performance.
Compensation and Benefits
- Base salary range for Level 4: USD 184,000–287,500 per year.
- Base salary range for Level 5: USD 224,000–356,500 per year.
- Base salary is determined based on location, experience, and the pay of employees in similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until June 19, 2026.
- NVIDIA uses AI tools in its recruiting processes.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AIOps and Observability
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Backend Engineer - Databases Pyroscope
Grafana Labs · Germany, Spain, Ireland, Sweden, United Kingdom
EUR 97,000-116,400 per year
Senior Backend Engineer - Databases Pyroscope
Grafana Labs · Sweden, Spain, Ireland, Germany, United Kingdom
SEK 775,400-930,500 per year
Principal Systems Software Engineer - Observability and Telemetry Platform
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior AI Infrastructure Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year