Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS
CI/CD @ 6
Communication @ 7
Debugging @ 7
GCP
Go @ 6
HPC @ 4
IaC
Kubernetes @ 4
Leadership @ 7
Mentoring @ 7
Observability @ 7
Perl @ 6
Python @ 6
Ruby @ 6
SRE @ 6
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Site Reliability Engineer to join its Compute Farm team and help build the next generation of its global services platform. The role focuses on keeping critically important systems running and using AI to deliver reliable solutions for high-performance computing and accelerated computing environments.
Responsibilities
- Own SRE solutions end to end, from design and implementation through operation and continuous improvement, ensuring integration with HPC schedulers, storage, and network fabrics.
- Use Infrastructure as Code and configuration management to standardize and automate provisioning.
- Deliver solutions in a globally distributed, multi-cloud hybrid environment spanning on-premises infrastructure, AWS, GCP, and OCI.
- Design for failure using redundancy, failure domains, progressive delivery, and strict change control.
- Ensure high uptime and Quality of Service for internal customers through operational excellence.
- Conduct capacity management and planning to meet ongoing operational needs.
- Detect performance issues and recommend solutions to maintain service quality.
- Collaborate with multiple teams in a fast-paced environment to ensure seamless project completion.
- Participate in on-call rotations and incident reviews, assist with root-cause identification, and produce high-quality root-cause analysis reports.
Requirements
- Bachelor's degree in Computer Science or a related technical field, or equivalent experience, with 5+ years of professional experience building and supporting critical services.
- Experience supporting large-scale HPC clusters using Slurm, LSF, or Kubernetes, including setup, tuning, and troubleshooting.
- Proficiency in modern CI/CD techniques and Infrastructure as Code for managing services.
- Strong experience building large-scale infrastructure platforms for automated host lifecycle management, fleet reliability and auto-healing, end-to-end observability, or data-driven operations using AIOps or machine-learning-driven signals.
- Proficiency with monitoring, metrics, container management, and log collection tools.
- 5+ years of coding or scripting experience in at least two high-level programming languages, such as Python, Go, Perl, or Ruby.
- Experience mentoring engineers and influencing technical direction through design reviews, architecture documents, and strong partnerships with product and leadership teams.
- Creative problem-solving abilities, excellent debugging skills, and strong communication and documentation skills.
Preferred Qualifications
- Published technical write-ups or talks, such as conference presentations, meetups, or engineering blogs, covering real-world reliability, observability, large-scale HPC, or SRE problems and solutions.
- Maintainer or co-maintainer experience for a production open-source component, such as a plugin, operator, exporter, controller, or SDK, used at large scale.
Benefits
- Equity and benefits are provided.
- NVIDIA offers a comprehensive benefits package.
Applications for this job will be accepted at least until June 19, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
AI Developer Technology Engineer
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Deep Learning Scientist, Multimodal Agentic RL
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Software Engineering Manager - Cloud Streaming
Nvidia · United States
USD 168,000-270,200 per year
Senior System Software Engineer, Agentic Retrieval
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Governance, Risk, and Compliance Certifications Engineer
Nvidia · United States, Santa Clara, United States
USD 168,000-270,200 per year
Similar jobs
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager – AI Platform & SRE
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms
GitLab · Canada, United Kingdom, United States
USD 126,400-314,400 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, Attestation Services – DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year