Senior Staff Site Reliability Operations

at Nvidia
USD 184,000-264,500 per year
SENIOR
✅ On-site

Tech Stack

AI Audit Bash @ 4 Communication @ 6 Compliance Leadership @ 6 Linux @ 7 Mentoring @ 4 Microsoft 365 @ 4 Networking @ 6 PowerShell @ 4 Python @ 4 Security @ 4 ServiceNow @ 4 Technical Leadership @ 6 macOS

Details

NVIDIA is seeking a Site Reliability Operations Technical Lead to serve as the senior technical individual contributor for reliability and support at its Seattle, Washington site. The role owns local technical service delivery, serves as the final support escalation point before regional and global platform teams, and provides technical leadership to site support engineers. Responsibilities include resolving complex issues across Active Directory, Exchange, database platforms, and compute infrastructure; leading major incidents; eliminating recurring problems; leading site-level projects; representing local requirements in global initiatives; and setting technical standards for the site support team.

Responsibilities

  • Own day-to-day site operations, including incidents, requests, critical issues, support coverage, queue health, SLA attainment, backlog, service quality, and site asset and inventory management across lifecycle, refresh, procurement, and compliance.
  • Serve as the Tier 3 escalation owner for the site and AMER across:
    • Identity: Active Directory, hybrid Entra ID, Group Policy, Kerberos/LDAP, SSO, and MFA.
    • Messaging: Exchange hybrid mail flow, mailboxes, and SMTP relay.
    • Compute: Windows, Linux, macOS, virtualization, storage, datacenter, and lab hardware.
    • Endpoints: Microsoft 365, Teams, Intune, Autopilot, and imaging through migrations.
  • Drive root-cause analysis and permanent fixes instead of repeat break-fix work.
  • Own endpoint compliance, vulnerability remediation, patch management, hardening, audit readiness, evidence, incident response, and privileged-access activities in partnership with InfoSec.
  • Drive critical issues with global platform teams and vendors using reproduction cases and diagnostic evidence through to a committed fix.
  • Lead site SRO engineers by setting standards, reviewing work, directing blocking issues, mentoring engineers, and maintaining the site knowledge base and runbook library.
  • Serve as the primary technical contact for site IT, partnering with employees, site and executive leadership, Facilities, Security, HR, and Procurement on incidents, planned changes, onboarding, moves, and office and lab expansions.
  • Build automation in PowerShell, Python, or Bash for diagnostics, remediation, health checks, and reporting.
  • Analyze ticket and reliability trends to eliminate recurring issues and champion AI-driven and agentic solutions that advance SRO strategy.
  • Represent site and AMER priorities in regional and global IT initiatives, standards, and architecture forums, and lead operational decisions in the manager's absence.

Requirements

  • 12+ years of experience in enterprise support engineering, infrastructure, or end-user services, including 5+ years in a senior, lead, or escalation-tier role in a multi-site environment.
  • Deep hands-on experience with Active Directory, hybrid Entra ID, Exchange hybrid, Windows and Linux servers, virtualization, enterprise storage, and datacenter hardware.
  • Experience with enterprise endpoint management, including Intune, Autopilot, MECM/SCCM, Jamf, or equivalent; Windows 11; the Microsoft 365 ecosystem; endpoint security; and vulnerability remediation.
  • Database operations support and networking fundamentals, including DNS, DHCP, VLANs, wireless, firewall policy, and switch-level troubleshooting.
  • Scripting and automation experience in Python, PowerShell, or Bash applied to real support problems.
  • Experience with ServiceNow or a similar ITSM platform.
  • Demonstrated technical leadership without formal authority, excellent executive-level communication during incidents, and a focus on root cause rather than symptom clearing.
  • Willingness to work on-site and hands-on, including lifting and moving equipment, joining an on-call rotation, and supporting after-hours maintenance windows and cutovers.
  • Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent experience.

Additional Qualifications

  • Experience supporting engineering, lab, R&D, or manufacturing environments with specialized equipment and non-standard availability requirements.
  • Experience serving as a local technical lead through a site buildout, relocation, or major migration.
  • Experience influencing global standards and tooling roadmaps for site and regional needs.
  • Experience with executive support programs, AV and hybrid conference room technologies, or build automation with measurable efficiency and experience benefits.

Compensation and Benefits

The base salary range is USD 184,000–264,500 per year, determined by location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

More jobs at Nvidia

Similar jobs