Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 7
Bash @ 6
CI/CD @ 4
Change Management
ClickHouse @ 4
Communication @ 7
Debugging @ 7
DevOps @ 6
Distributed Systems @ 6
ElasticSearch @ 4
Flink @ 4
GPU
Grafana @ 4
Helm @ 4
IaC
Kafka @ 4
Kubernetes @ 7
Linux @ 7
Microservices @ 7
Networking @ 7
Observability @ 4
Prometheus @ 4
Python @ 4
SRE @ 6
Spark @ 4
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. This role operates the platform itself rather than the compute cluster, with ownership of uptime, performance, data integrity, and safe change management.
You will own SLOs and SLIs, incident response, and postmortems for telemetry ingestion, processing, storage, APIs, and dashboards. You will partner with Software Engineering and Systems Engineering teams to translate platform signals into actionable, trustworthy alerts and automation.
Responsibilities
- Continuously monitor platform health through dashboards, logs, and metrics; automate recurring checks and maintain reliability and resource efficiency.
- Own Kubernetes deployments end to end, including runbooks, canary checks, post-deployment validation, rollbacks, and remediation.
- Lead first-level incident triage by collecting diagnostics, identifying likely root causes, and providing clear, actionable findings to engineering.
- Build and maintain runbooks, standard operating procedures, and checklists while driving continuous improvement through automation.
- Manage deployment infrastructure and packaging using Helm and Terraform/IaC to keep environments scalable, consistent, and reproducible.
- Contribute in adjacent functional areas to support the team.
Requirements
- Bachelor’s or master’s degree in computer science, computer engineering, or equivalent experience.
- At least 5 years of experience operating production distributed systems in SRE, DevOps, or Platform Operations roles.
- Proven ownership of reliability for an observability or AIOps platform, including SLOs/SLIs, on-call responsibilities, incident response, and follow-up evaluations that drive measurable improvements.
- Deep Kubernetes and container experience, including deploying, debugging, and scaling telemetry-heavy microservices covering ingestion, processing, storage, APIs, and user interfaces.
- Solid scripting skills in Python and/or Bash.
- Experience with CI/CD and infrastructure as code using Terraform and Helm.
- Experience delivering safe rollouts, including canary releases and rollbacks, reproducible environments, and reduced operational toil.
- Strong communication skills, including the ability to write excellent runbooks and documentation and translate ambiguous requirements into concrete operational practices.
Preferred Qualifications
- Strong Linux and networking fundamentals, distributed-systems experience, and hands-on operations for Kubernetes, services, and streaming stacks.
- Experience with observability platforms at scale.
- Experience building trusted operational automation, including canary releases, automated rollback criteria, monitoring for monitoring lag/drop/error budgets, and replay or backfill pipelines with correctness checks.
- Experience operating distributed and streaming systems such as Kafka, Pulsar, Flink, Spark, ClickHouse, Elasticsearch, time-series databases, and object storage.
- Ability to reason about backpressure, hotspots, and failure domains end to end.
- Programming experience building automation tools or services, ideally in Python or similar languages.
- Experience running large-scale production deployments and multiple Kubernetes environments or clusters across teams or customers.
- Hands-on experience with observability tools and dashboards, metrics, logs, and traces using platforms such as Prometheus and Grafana or similar tools.
Benefits
The position offers equity and benefits in addition to the base salary. NVIDIA states that it provides a diverse work environment and is an equal opportunity employer.