Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 7
Bash @ 6
CI/CD @ 6
Change Management
ClickHouse @ 7
Debugging @ 7
DevOps @ 6
Distributed Systems @ 6
Flink @ 7
GPU
Grafana @ 4
Helm @ 6
IaC
Kafka @ 7
Kubernetes @ 7
Linux @ 1
Microservices @ 7
Networking @ 1
Observability @ 1
Prometheus @ 4
Python @ 4
SRE @ 6
Spark @ 7
Terraform @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. This role is for a DevOps Engineer to operate the platform itself (not the compute cluster): ensuring uptime, performance, data integrity, and safe change management. You will own SLOs/SLIs, incident response, and postmortems for telemetry ingestion, processing, storage, and the APIs/dashboards that operators depend on. You will partner with Software Engineering and Systems Engineering to translate platform signals into actionable, trustworthy alerts and automation.
Responsibilities
- Continuously monitor platform health via dashboards/logs/metrics, automate recurring checks, and keep reliability + resource efficiency on track.
- Own Kubernetes deployments end-to-end (runbooks, canary checks, post-deploy validation), and lead rollbacks/remediations when needed.
- Lead first-level incident triage: collect diagnostics, identify likely root causes, and hand off clear, actionable findings to engineering.
- Build and maintain runbooks/SOPs/checklists, pushing continuous improvement through automation.
- Manage deployment infrastructure and packaging (Helm + Terraform/IaC) to keep environments scalable, consistent, and reproducible.
- Contribute in adjacent functional areas to grow and help your team members!
Requirements
- BS/MS in CS/CE (or equivalent experience) and 5+ years operating production distributed systems as SRE/DevOps/Platform Ops.
- Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing incidents, and follow-up evaluations that drive measurable improvements.
- Deep Kubernetes + containers experience (deploying, debugging, scaling) for telemetry-heavy microservices—ingestion, processing, storage, APIs, and UI.
- Automation-first approach: solid scripting (Python/Bash), CI/CD, and infrastructure-as-code (Terraform + Helm) to deliver safe rollouts (canaries/rollbacks), reproducible environments, and minimal toil.
- Clear communicator who writes excellent runbooks/docs and can translate ambiguous requirements into concrete operational practices and dependable customer-facing reliability.
Ways to stand out from the crowd
- Strong Linux + networking fundamentals, distributed systems instincts, and hands-on ops for Kubernetes/services/streaming stacks are ideal; bonus for experience with observability platforms at scale.
- Experience building safe automation that operators trust: canary releases, automated rollback criteria, “monitoring for the monitoring” (lag/drop/error budgets), and replay/backfill pipelines with correctness checks.
- Strong in distributed/streaming systems operations (Kafka/Pulsar, Flink/Spark, ClickHouse/Elastic/TSDBs, object storage)—and can reason about backpressure, hotspots, and failure domains end-to-end.
- Proven programming experience building automation tools or services—ideally in Python, or similar languages—to simplify operations and scale recurring processes.
- Proven experience running large-scale production deployments and multiple Kubernetes environments or clusters across teams or customers, coordinating changes and rollouts with minimal disruption with hands-on experience with observability tools—you know your way around dashboards, metrics, logs, and traces using platforms like Prometheus, Grafana, or similar.
More jobs at Nvidia
HPC Performance Engineer
Nvidia · United States
USD 152,000-241,500 per year
Senior System Software Engineer - Halos Core And Robotics Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer – Dynamo Tools
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Systems Software Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Perception Engineer, Obstacle Foundation Models - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Systems Engineer, Storage - DGX Cloud
Nvidia · United States
USD 208,000-414,000 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior SRE Engineer
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Senior Site Reliability Engineer (In-Office Required)
Nebius · New York City, United States
USD 156,000-262,000 per year
Senior Software Engineer - Datacenter Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year