Senior Site Reliability Engineer, AIOps

at Nvidia
USD 148,000-276,000 per year
SENIOR
✅ On-site

Tech Stack

AI API @ 7 Bash @ 6 CI/CD @ 4 Change Management ClickHouse @ 4 Communication @ 7 Debugging @ 7 DevOps @ 6 Distributed Systems @ 6 ElasticSearch @ 4 Flink @ 4 GPU Grafana @ 4 Helm @ 4 IaC Kafka @ 4 Kubernetes @ 7 Linux @ 7 Microservices @ 7 Networking @ 7 Observability @ 4 Prometheus @ 4 Python @ 4 SRE @ 6 Spark @ 4 Terraform @ 4

Details

NVIDIA is building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. This role operates the platform itself rather than the compute cluster, with ownership of uptime, performance, data integrity, and safe change management.

You will own SLOs and SLIs, incident response, and postmortems for telemetry ingestion, processing, storage, APIs, and dashboards. You will partner with Software Engineering and Systems Engineering teams to translate platform signals into actionable, trustworthy alerts and automation.

Responsibilities

  • Continuously monitor platform health through dashboards, logs, and metrics; automate recurring checks and maintain reliability and resource efficiency.
  • Own Kubernetes deployments end to end, including runbooks, canary checks, post-deployment validation, rollbacks, and remediation.
  • Lead first-level incident triage by collecting diagnostics, identifying likely root causes, and providing clear, actionable findings to engineering.
  • Build and maintain runbooks, standard operating procedures, and checklists while driving continuous improvement through automation.
  • Manage deployment infrastructure and packaging using Helm and Terraform/IaC to keep environments scalable, consistent, and reproducible.
  • Contribute in adjacent functional areas to support the team.

Requirements

  • Bachelor’s or master’s degree in computer science, computer engineering, or equivalent experience.
  • At least 5 years of experience operating production distributed systems in SRE, DevOps, or Platform Operations roles.
  • Proven ownership of reliability for an observability or AIOps platform, including SLOs/SLIs, on-call responsibilities, incident response, and follow-up evaluations that drive measurable improvements.
  • Deep Kubernetes and container experience, including deploying, debugging, and scaling telemetry-heavy microservices covering ingestion, processing, storage, APIs, and user interfaces.
  • Solid scripting skills in Python and/or Bash.
  • Experience with CI/CD and infrastructure as code using Terraform and Helm.
  • Experience delivering safe rollouts, including canary releases and rollbacks, reproducible environments, and reduced operational toil.
  • Strong communication skills, including the ability to write excellent runbooks and documentation and translate ambiguous requirements into concrete operational practices.

Preferred Qualifications

  • Strong Linux and networking fundamentals, distributed-systems experience, and hands-on operations for Kubernetes, services, and streaming stacks.
  • Experience with observability platforms at scale.
  • Experience building trusted operational automation, including canary releases, automated rollback criteria, monitoring for monitoring lag/drop/error budgets, and replay or backfill pipelines with correctness checks.
  • Experience operating distributed and streaming systems such as Kafka, Pulsar, Flink, Spark, ClickHouse, Elasticsearch, time-series databases, and object storage.
  • Ability to reason about backpressure, hotspots, and failure domains end to end.
  • Programming experience building automation tools or services, ideally in Python or similar languages.
  • Experience running large-scale production deployments and multiple Kubernetes environments or clusters across teams or customers.
  • Hands-on experience with observability tools and dashboards, metrics, logs, and traces using platforms such as Prometheus and Grafana or similar tools.

Benefits

The position offers equity and benefits in addition to the base salary. NVIDIA states that it provides a diverse work environment and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs