Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS
Airflow @ 4
Ansible @ 4
Azure
CUDA @ 4
Chef @ 4
Cloud Computing
GCP
GPU @ 4
GenAI
Generative AI @ 4
Go @ 6
Grafana @ 4
HPC
Kubernetes @ 4
LLM @ 4
Linux @ 4
Microservices @ 4
NCCL @ 4
Networking @ 4
Observability @ 4
OpenTelemetry @ 4
Performance Analysis @ 4
Prometheus @ 4
Puppet @ 4
PyTorch @ 4
Python @ 6
SGLang @ 4
SRE @ 4
Security @ 4
Splunk @ 4
TensorRT @ 4
Terraform @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is driving AI and high-performance computing forward. DGX Cloud delivers a fully managed AI platform on major cloud providers, optimizing AI workloads using high-performance NVIDIA infrastructure. As a Senior Site Reliability Engineer on the DGX Cloud team, you will maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide.
This role offers the opportunity to work with innovative AI and cloud computing solutions as part of a team delivering ambitious projects and advancing large-scale infrastructure reliability.
Responsibilities
- Build, implement, and support the operational and reliability aspects of large-scale Kubernetes clusters, with a focus on performance at scale, real-time monitoring, logging, and alerting.
- Define SLOs and SLIs, monitor error allowances, and streamline reporting.
- Support services before launch through system creation consulting, software tools, platforms and frameworks development, capacity management, and launch reviews.
- Maintain live services by measuring and supervising availability, latency, and overall system health.
- Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.
- Scale systems sustainably through automation and drive changes that improve reliability and velocity.
- Lead triage and root-cause analysis of high-severity incidents.
- Practice balanced incident response and blameless postmortems.
- Participate in an on-call rotation to support production services.
Requirements
- Bachelor's degree in Computer Science or a related technical field, or equivalent experience.
- 8+ years of experience operating production services.
- Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.
- Experience with infrastructure automation tools such as Terraform, Ansible, Chef, or Puppet.
- Proficiency in at least one high-level programming language, such as Python or Go.
- In-depth knowledge of Linux operating systems, TCP/IP networking fundamentals, and cloud security standards.
- Solid understanding of SRE principles, including SLOs, SLIs, error budgets, and incident management.
- Experience building and operating comprehensive observability stacks for monitoring, logging, and tracing using tools such as OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, or Splunk.
Preferred Qualifications
- Experience operating GPU-accelerated clusters with KubeVirt in production.
- Experience applying generative AI techniques to reduce operational toil.
- Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.
- Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis.
Compensation and Benefits
The base salary range is $168,000–$270,250 for Level 4 and $208,000–$333,500 for Level 5. Salary is determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.