Senior Site Reliability Engineer, DGX Cloud

at Nvidia
USD 168,000-333,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 AWS Airflow @ 4 Ansible @ 4 Azure CUDA @ 4 Chef @ 4 Cloud Computing GCP GPU @ 4 GenAI Generative AI @ 4 Go @ 6 Grafana @ 4 HPC Kubernetes @ 4 LLM @ 4 Linux @ 4 Microservices @ 4 NCCL @ 4 Networking @ 4 Observability @ 4 OpenTelemetry @ 4 Performance Analysis @ 4 Prometheus @ 4 Puppet @ 4 PyTorch @ 4 Python @ 6 SGLang @ 4 SRE @ 4 Security @ 4 Splunk @ 4 TensorRT @ 4 Terraform @ 4 vLLM @ 4

Details

NVIDIA is driving AI and high-performance computing forward. DGX Cloud delivers a fully managed AI platform on major cloud providers, optimizing AI workloads using high-performance NVIDIA infrastructure. As a Senior Site Reliability Engineer on the DGX Cloud team, you will maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide.

This role offers the opportunity to work with innovative AI and cloud computing solutions as part of a team delivering ambitious projects and advancing large-scale infrastructure reliability.

Responsibilities

  • Build, implement, and support the operational and reliability aspects of large-scale Kubernetes clusters, with a focus on performance at scale, real-time monitoring, logging, and alerting.
  • Define SLOs and SLIs, monitor error allowances, and streamline reporting.
  • Support services before launch through system creation consulting, software tools, platforms and frameworks development, capacity management, and launch reviews.
  • Maintain live services by measuring and supervising availability, latency, and overall system health.
  • Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.
  • Scale systems sustainably through automation and drive changes that improve reliability and velocity.
  • Lead triage and root-cause analysis of high-severity incidents.
  • Practice balanced incident response and blameless postmortems.
  • Participate in an on-call rotation to support production services.

Requirements

  • Bachelor's degree in Computer Science or a related technical field, or equivalent experience.
  • 8+ years of experience operating production services.
  • Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.
  • Experience with infrastructure automation tools such as Terraform, Ansible, Chef, or Puppet.
  • Proficiency in at least one high-level programming language, such as Python or Go.
  • In-depth knowledge of Linux operating systems, TCP/IP networking fundamentals, and cloud security standards.
  • Solid understanding of SRE principles, including SLOs, SLIs, error budgets, and incident management.
  • Experience building and operating comprehensive observability stacks for monitoring, logging, and tracing using tools such as OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, or Splunk.

Preferred Qualifications

  • Experience operating GPU-accelerated clusters with KubeVirt in production.
  • Experience applying generative AI techniques to reduce operational toil.
  • Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.
  • Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis.

Compensation and Benefits

The base salary range is $168,000–$270,250 for Level 4 and $208,000–$333,500 for Level 5. Salary is determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.

More jobs at Nvidia

Similar jobs