Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Audit
Bash @ 4
Communication @ 6
Compliance
GPU
Leadership @ 6
Linux @ 7
Microsoft 365 @ 4
Networking @ 6
PowerShell @ 4
Python @ 4
Security @ 4
ServiceNow @ 4
Technical Leadership @ 6
macOS
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
For over 25 years, NVIDIA has been at the forefront of transforming computer graphics, PC gaming, and accelerated computing. NVIDIA is applying the potential of AI to create the next era of computing, with GPUs powering computers, robots, and autonomous vehicles.
The Site Reliability Operations Technical Lead serves as the senior technical individual contributor for reliability and support at the site. This role owns local technical service delivery, acts as the final support escalation point before regional and global platform teams, and provides technical leadership to site support engineers.
The role is responsible for complex issues across Active Directory, Exchange, database platforms, and compute infrastructure; leading major incidents; and eliminating recurring problems. The successful candidate will lead site-level projects, represent local requirements in global initiatives, establish technical standards, conduct root cause analyses, support executives, and brief IT leadership on site risk.
Responsibilities
- Own day-to-day site operations, including incidents, requests, critical issues, support coverage, queue health, SLA attainment, backlog, service quality, and site asset and inventory management across lifecycle, refresh, procurement, and compliance.
- Serve as the Tier 3 escalation owner for the site and AMER across:
- Identity: Active Directory, hybrid Entra ID, Group Policy, Kerberos/LDAP, SSO, and MFA.
- Messaging: Exchange hybrid mail flow, mailboxes, and SMTP relay.
- Compute: Windows, Linux, macOS, virtualization, storage, datacenter hardware, and lab hardware.
- Endpoints: Microsoft 365, Teams, Intune, Autopilot, and imaging through migrations.
- Drive root-cause analysis and permanent fixes instead of repeated break-fix activities.
- Own endpoint compliance, vulnerability remediation, patch management, hardening, audit readiness, evidence collection, incident response partnership with InfoSec, and privileged access.
- Escalate critical issues to global platform teams and vendors using reproduction cases and diagnostic evidence, and follow them through to resolution.
- Act as technical lead for site SRO engineers by setting standards, reviewing work, directing blocking issues, providing mentorship, and maintaining the site knowledge base and runbook library.
- Serve as the primary technical contact for site IT and partner with employees, site and executive leadership, Facilities, Security, HR, and Procurement on incidents, planned changes, onboarding, moves, and office and lab expansions.
- Build automation in PowerShell, Python, or Bash for diagnostics, remediation, health checks, and reporting.
- Analyze ticket and reliability trends to eliminate recurring issues and champion AI-driven and agentic solutions supporting the SRO strategy.
- Represent site and AMER priorities in regional and global IT initiatives, standards, and architecture forums.
- Lead operational decisions in the manager's absence.
Requirements
- At least 8 years of experience in enterprise support engineering, infrastructure, or end-user services, including at least 5 years in a senior, lead, or escalation-tier role in a multi-site environment.
- Deep hands-on experience with Active Directory, hybrid Entra ID, Exchange hybrid, Windows and Linux servers, virtualization, enterprise storage, and datacenter hardware.
- Experience with enterprise endpoint management, including Intune, Autopilot, MECM/SCCM, Jamf, or equivalent; Windows 11; the Microsoft 365 ecosystem; endpoint security; and vulnerability remediation.
- Database operations support and networking fundamentals, including DNS, DHCP, VLANs, wireless, firewall policy, and switch-level troubleshooting.
- Scripting and automation experience with Python, PowerShell, or Bash applied to support problems.
- Experience with ServiceNow or a similar ITSM platform.
- Demonstrated technical leadership without formal authority, excellent executive-level communication during incidents, and a focus on root cause rather than symptom resolution.
- Willingness to work on-site and hands-on, including lifting and moving equipment, joining an on-call rotation, and supporting after-hours maintenance windows and cutovers.
- Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent experience.
Preferred Qualifications
- Experience supporting engineering, laboratory, research and development, or manufacturing environments with specialized equipment and non-standard availability requirements.
- Experience serving as a local technical lead during a site buildout, relocation, or major migration.
- Experience influencing global standards and tooling roadmaps for site and regional needs.
- Experience with executive support programs, AV and hybrid conference room technologies, or build automation with measurable efficiency and experience benefits.
Compensation and Benefits
The base salary range is USD 144,000–230,000 per year. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until October 5, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.