Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 7
Distributed Systems @ 7
Engineering Management @ 4
Go @ 4
HPC @ 4
Leadership @ 4
Linux @ 7
Networking @ 7
Performance Optimization @ 7
Python @ 4
Security
Technical Leadership @ 4
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Lead NVIDIA's Compute Core Engineering team and contribute to advancements in AI and computing technology. The role involves technical leadership, infrastructure strategy, engineering management, and modernization of globally distributed compute infrastructure.
Responsibilities
- Lead and expand NVIDIA's global Compute Core Engineering team; recruit, mentor, and develop technical talent while fostering ownership, collaboration, and innovation.
- Define the strategy and roadmap for infrastructure services across NVIDIA's data centers, cloud environments, and compute platforms.
- Transform traditional services into automated, self-service platforms designed for reliability and scalability, while modernizing legacy infrastructure with minimal disruption.
- Provide technical leadership for globally distributed services, including DNS, DHCP, NTP/PTP, LDAP, and Linux infrastructure.
- Establish engineering standards for performance, capacity, scalability, security, lifecycle management, and disaster recovery.
- Promote automation and operational excellence, including AI-assisted operations for diagnostics and routine infrastructure tasks.
- Partner with senior engineers to improve infrastructure performance and resiliency, including the development of automation platforms using Go, Python, Terraform, and GitOps approaches.
- Establish service-health metrics covering availability, latency, capacity utilization, and infrastructure efficiency.
- Lead incident reviews and reliability improvement programs.
- Lead forecasting activities and collaborate across teams to ensure infrastructure can meet future demand.
- Work with NVIDIA leadership and internal customers to build infrastructure products.
Requirements
- Bachelor's or master's degree in a related technical field, or equivalent experience.
- 12 or more years of experience in related fields, including significant technical leadership experience.
- 5 or more years of engineering management experience.
- Experience crafting and managing critically important infrastructure.
- Strong understanding of distributed systems architecture, networking technologies, and networking protocols.
- Deep experience with Linux operating systems and experience with container platforms and technologies.
- Experience with infrastructure automation and GitOps approaches.
- Software engineering experience with Go, Python, or similar programming languages.
- Experience using SLIs, SLOs, and performance metrics.
- Experience with capacity forecasting and lifecycle oversight.
- Ability to lead complex modernization projects.
- Strong communication skills, including the ability to explain technical topics to senior leadership.
- Ability to attract, support, and retain high-performing teams.
Preferred Qualifications
- Experience building and managing foundational infrastructure services.
- Experience supporting large AI, HPC, or accelerated-computing environments.
- Hands-on experience with advanced networking and infrastructure acceleration technologies.
- Deep understanding of Linux kernel internals and system performance optimization.
- Experience modernizing legacy infrastructure into self-service platforms.
- Experience building engineering platforms around Terraform and GitOps or equivalent technologies.
- Experience implementing AI-assisted operations for infrastructure management.
- Proven success reducing operational toil through engineering and automation.
- Experience leading globally distributed teams supporting critical infrastructure.
- Ability to remain technically engaged while providing organizational leadership and strategy.
Compensation and Benefits
The base salary range is USD 248,000–391,000 per year. Base salary is determined by location, experience, and the compensation of employees in similar positions. The role is also eligible for equity and benefits.
NVIDIA is committed to an inclusive work environment and is an equal opportunity employer. Applications will be accepted at least until September 3, 2026.
More jobs at Nvidia
Senior NPI Program Manager
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
GPU PCIe and Boot Architect - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior AI Engineer, High Performance AI
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Salesforce CPQ Developer
Nvidia · Santa Clara, United States
USD 176,000-276,000 per year
Senior Technical Program Manager - LLM Safety
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Principal Software Engineer – Infrastructure
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Staff Platform Engineer
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Principal Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Staff Software Engineer - Databases SRE | Ireland | Remote
Grafana Labs · Germany, Spain, United Kingdom, Ireland, Sweden
EUR 117,600-141,100 per year
Staff Software Engineer - Databases SRE
Grafana Labs · Germany, Spain, United Kingdom, Sweden
EUR 109,700-131,700 per year