Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS
CI/CD @ 6
Communication @ 7
Debugging @ 7
GCP
Go @ 6
HPC @ 4
IaC
Kubernetes @ 4
Leadership @ 7
Mentoring @ 7
Observability @ 7
Perl @ 6
Python @ 6
Ruby @ 6
SRE @ 6
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Site Reliability Engineer to join its Compute Farm team and help build the next generation of its global services platform. The role focuses on keeping critically important systems running and using AI to deliver reliable solutions for high-performance computing and accelerated computing environments.
Responsibilities
- Own SRE solutions end to end, from design and implementation through operation and continuous improvement, ensuring integration with HPC schedulers, storage, and network fabrics.
- Use Infrastructure as Code and configuration management to standardize and automate provisioning.
- Deliver solutions in a globally distributed, multi-cloud hybrid environment spanning on-premises infrastructure, AWS, GCP, and OCI.
- Design for failure using redundancy, failure domains, progressive delivery, and strict change control.
- Ensure high uptime and Quality of Service for internal customers through operational excellence.
- Conduct capacity management and planning to meet ongoing operational needs.
- Detect performance issues and recommend solutions to maintain service quality.
- Collaborate with multiple teams in a fast-paced environment to ensure seamless project completion.
- Participate in on-call rotations and incident reviews, assist with root-cause identification, and produce high-quality root-cause analysis reports.
Requirements
- Bachelor's degree in Computer Science or a related technical field, or equivalent experience, with 5+ years of professional experience building and supporting critical services.
- Experience supporting large-scale HPC clusters using Slurm, LSF, or Kubernetes, including setup, tuning, and troubleshooting.
- Proficiency in modern CI/CD techniques and Infrastructure as Code for managing services.
- Strong experience building large-scale infrastructure platforms for automated host lifecycle management, fleet reliability and auto-healing, end-to-end observability, or data-driven operations using AIOps or machine-learning-driven signals.
- Proficiency with monitoring, metrics, container management, and log collection tools.
- 5+ years of coding or scripting experience in at least two high-level programming languages, such as Python, Go, Perl, or Ruby.
- Experience mentoring engineers and influencing technical direction through design reviews, architecture documents, and strong partnerships with product and leadership teams.
- Creative problem-solving abilities, excellent debugging skills, and strong communication and documentation skills.
Preferred Qualifications
- Published technical write-ups or talks, such as conference presentations, meetups, or engineering blogs, covering real-world reliability, observability, large-scale HPC, or SRE problems and solutions.
- Maintainer or co-maintainer experience for a production open-source component, such as a plugin, operator, exporter, controller, or SDK, used at large scale.
Benefits
- Equity and benefits are provided.
- NVIDIA offers a comprehensive benefits package.
Applications for this job will be accepted at least until June 19, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Deep Learning Algorithm Engineering Intern - 2026
Nvidia · Zurich, Switzerland
PLN 117,800-204,100 per year
Senior Project Delivery Manager - NVIS
Nvidia · United States
USD 168,000-322,000 per year
Principal Research Scientist, Synthetic Data Generation
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Staff Software Engineer - Agentic Automation
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Firmware Engineer, NIC Firmware
Nvidia · Seattle, United States
USD 152,000-287,500 per year
Similar jobs
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Senior Manager, Software Engineering - Agentic IT Operations
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms
GitLab · Canada, United Kingdom, United States
USD 126,400-314,400 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · Palo Alto, United States, San Francisco, United States
USD 220,000-405,000 per year