Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 1
Ansible @ 7
Azure @ 1
Communication @ 7
GCP @ 1
GPU @ 3
Git @ 4
Grafana @ 6
HPC
Jira @ 6
Kubernetes @ 7
Linux @ 6
Mathematics @ 4
Mentoring @ 7
Networking @ 7
OpenTelemetry @ 6
Python @ 7
SRE
Security
Slurm @ 7
System Administration @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Compute Infrastructure Support team is seeking a driven Site Reliability Engineer to support on-premises and cloud products and services. The role focuses on maintaining near-100% availability across systems supporting AI and accelerated computing.
Responsibilities
- Work within a 24/7 follow-the-sun support model across multiple continents, collaborating with a U.S.-based manager.
- Work a four-day, 10-hour schedule, including either Saturday or Sunday, with flexible early or late shifts.
- Monitor and manage large-scale production GPU and Kubernetes environments to maintain availability and performance.
- Use advanced tools to proactively detect, prevent, and respond to incidents.
- Analyze logs, metrics, and system behavior to diagnose issues and implement effective resolutions.
- Develop predictive automated support routines to prevent production issues.
- Improve automation by integrating incident analysis into auto-healing and automated break-fix solutions.
- Perform systems administration, network administration, and security monitoring.
- Coordinate with domain experts and service owners to resolve complex issues.
- Continuously improve service quality and operational processes based on incident feedback.
- Coordinate effectively across teams during incident resolution and provide strong customer-focused support.
Requirements
- Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.
- Familiarity with GPU hardware and high-performance computing environments.
- Proficiency with Grafana, OpenTelemetry, PagerDuty, and JIRA.
- Experience with AWS, Azure, GCP, or OCI is a plus; strong on-premises expertise is preferred.
- At least 8 years of experience coordinating large-scale production systems, including more than 3 years in high-availability Internet, cloud, or data center environments.
- Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or equivalent experience.
- Expert-level Linux system administration skills.
- Experience with Ansible and/or Python automation and strong shell scripting skills.
- Strong knowledge of DNS, DHCP, storage systems, and core networking.
- Experience troubleshooting and maintaining large-scale bare-metal infrastructure.
- Strong partnership, documentation, communication, and mentoring skills.
Preferred Qualifications
- Experience with scripting languages, particularly Python.
- Experience running virtual machines under community-supported or commercial hypervisors.
- Knowledge of application containers and container orchestration systems.
- Basic understanding of Git.
- Ability to master and maintain complex environments.
Compensation and Benefits
- Level 4 base salary: USD 168,000–270,250 per year.
- Level 5 base salary: USD 208,000–333,500 per year.
- Eligible for equity and benefits.
- Applications will be accepted at least until August 14, 2026.
NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
More jobs at Nvidia
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Director, Global Risk and Compliance
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior System Software Engineer - CPU SoC Boot Firmware
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Forward-Deployed Engineer, Enterprise AI and Automation
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Platform Engineer
Collibra · United States
USD 168,000-210,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Principal Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Staff Software Engineer - Databases SRE | Ireland | Remote
Grafana Labs · Ireland, Spain, Sweden, Germany, United Kingdom
EUR 117,600-141,100 per year
Staff Software Engineer - Databases SRE
Grafana Labs · Germany, Sweden, Spain, United Kingdom
EUR 109,700-131,700 per year