Staff Site Reliability Engineer - AI Platform Runtime

at Nvidia
USD 168,000-333,500 per year
SENIOR
✅ Hybrid

Tech Stack

AI AWS @ 4 Azure @ 6 CloudFormation @ 4 Communication @ 6 Distributed Systems GCP @ 6 Go @ 7 IaC JavaScript @ 7 Kubernetes @ 4 Machine Learning Mathematics @ 4 Networking @ 4 Observability @ 4 OpenTelemetry @ 4 Python @ 7 SRE Security Terraform @ 4 TypeScript @ 7

Details

Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and maintaining large-scale production systems with high efficiency and availability through software and systems engineering practices. The discipline requires knowledge of systems, networking, coding, databases, capacity management, continuous delivery and deployment, Kubernetes, and public cloud technologies.

SREs ensure that internal and external services operate with maximum reliability and uptime while enabling developers to make changes through careful preparation and planning. The work includes automation, performance tuning, capacity and latency management, observability, and improving production-system efficiency. NVIDIA promotes a culture of diversity, intellectual curiosity, problem-solving, openness, self-direction, collaboration, mentorship, and blameless improvement.

Responsibilities

  • Lead the technical strategy and roadmap for large-scale, cross-functional SRE initiatives that improve reliability, scalability, and developer productivity across enterprise systems.
  • Design and build resilient distributed systems that power NVIDIA's next-generation AI-driven enterprise products and services.
  • Architect and develop AI agents and AI skills to accelerate platform operations.
  • Drive automation and observability improvements, using metrics and analytics to enhance performance, reliability, and efficiency.
  • Collaborate across Cloud, Platform, Security, and AI/ML teams to implement modern SRE components that ensure high availability and secure operations.
  • Analyze and troubleshoot complex systems while championing best practices in system design, incident management, and postmortem analysis.
  • Mentor and influence engineers across teams, fostering technical excellence and a culture of reliability engineering.

Requirements

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
  • Bachelor's degree in Computer Science or a related technical field involving coding, such as physics or mathematics, or equivalent experience.
  • Strong proficiency in Python, TypeScript, JavaScript, or Go, with a focus on automation and infrastructure as code.
  • Experience with infrastructure-as-code tools such as AWS CDK, AWS CloudFormation, Terraform, or Crossplane.
  • Solid understanding of OpenTelemetry or other observability implementations at scale.
  • Deep expertise in systems architecture, networking, Kubernetes, and public cloud services including AWS, Azure, or GCP.
  • Outstanding problem-solving, communication, and teamwork skills, with the ability to influence across technical and interpersonal boundaries.

Preferred Qualifications

  • Passion for and experience with public cloud or large-scale automation systems.
  • Demonstrated ability to drive technical strategy and deliver measurable reliability outcomes in complex environments.
  • Strong ownership, curiosity, and innovation, with the ability to thrive in ambiguity and turn challenges into opportunities.

Benefits

The role includes eligibility for equity and NVIDIA benefits.

NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment. Applications will be accepted at least until September 12, 2026.

More jobs at Nvidia

Similar jobs