Site Reliability Engineer - Hardware Infrastructure

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Communication @ 7 Customer Support DevOps @ 7 Distributed Systems @ 6 GenAI Generative AI @ 4 Go @ 4 Grafana @ 4 LLM @ 4 Observability @ 4 Perl @ 4 Prometheus @ 4 Python @ 4 Ruby @ 4 SRE @ 7

Details

At NVIDIA, Site Reliability Engineering provides an opportunity to define, develop, and support large-scale production systems with high efficiency and availability. This role combines software and systems engineering to ensure reliable service operation, consistent uptime, and efficient system performance. The SRE will work collaboratively to enable developers to make significant updates while maintaining reliable system operations.

Responsibilities

  • Develop and support guidelines for incident management, planned maintenance, and blameless postmortems.
  • Assist teams with high-severity incidents, drive root cause analysis, create high-quality postmortems, and develop post-incident corrective actions.
  • Define reliability and supportability metrics, Service Level Objectives (SLOs), and error budgets.
  • Develop and drive adoption of actionable, customer-centric monitoring and alerting.
  • Apply automation and Generative AI/Agentic solutions to minimize manual and repetitive activities and improve customer support.
  • Guide teams in establishing sustainable on-call and operational standards.

Requirements

  • Degree in Computer Science or a related technical field involving coding, or equivalent experience.
  • 8+ years of experience in SRE, DevOps, or Production Engineering.
  • Strong understanding of SRE principles, including incident management, error budgets, SLOs, and Service Level Agreements (SLAs).
  • Experience designing and deploying fault-tolerant, performant, and supportable systems.
  • Background in infrastructure automation.
  • Experience operating critical services in production.
  • Experience with one or more of Python, Go, Perl, or Ruby.
  • Hands-on experience with observability platforms such as Prometheus and Grafana.
  • Strong communication skills and the ability to explain technical concepts effectively to diverse audiences.
  • Flexibility and adaptability in a fast-paced environment with evolving requirements.

Preferred Qualifications

  • Expertise in establishing incident management and postmortem processes.
  • Experience driving adoption of common tools and processes across diverse groups.
  • Experience working with LLM, Generative AI, or Agentic solutions to shorten mitigation time, reduce toil, and ensure SLOs are met.
  • Hands-on expertise operating and scaling distributed systems with tight SLAs while ensuring high availability and performance.

Compensation and Benefits

  • Base salary range for Level 4: USD 184,000–287,500 per year.
  • Base salary range for Level 5: USD 224,000–356,500 per year.
  • Base salary is determined based on location, experience, and the pay of employees in similar positions.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until June 19, 2026.
  • NVIDIA uses AI tools in its recruiting processes.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs