Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 6
Communication @ 7
Distributed Systems @ 4
GPU @ 6
Kubernetes @ 6
LLM @ 6
Leadership @ 4
Linux @ 6
Machine Learning @ 4
SGLang @ 6
Technical Leadership @ 6
TensorRT @ 6
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a technology leader to define and drive the global strategy for scaled-out AI inferencing. The role involves architecting high-throughput, low-latency distributed pipelines and model-serving strategies for massive-scale production workloads. You will define the technical roadmap for the full model lifecycle, including deployment, versioning, automated scaling, and operations across enterprise and cloud environments.
Responsibilities
- Architect and drive the technical implementation of high-throughput, low-latency distributed inference systems for massive-scale AI workloads.
- Lead hardware-software co-optimization, including performance tuning at the kernel and driver level, GPU resource management, and hardware acceleration for production-grade model serving.
- Guide and influence open-source and ecosystem projects, including Dynamo, TensorRT-LLM, vLLM, SGLang, Linux, Kubernetes, and Ray.
- Lead the strategy for full-lifecycle model management, including automated deployment, versioning, and intelligent scaling across cloud and datacenter environments.
- Collaborate with customers, infrastructure providers, and partners to ensure NVIDIA solutions achieve industry-leading performance and availability.
- Lead all technical aspects of a large scope across ideation, architecture, design, development, deployment, operations, and continuous lifecycle management.
Requirements
- 16 or more years in technical roles, with a long-term focus on AI infrastructure and recent direct experience in large-scale inference orchestration.
- Proven experience building secure, highly available, and durable production distributed systems.
- 7–10 or more years of leadership experience.
- Bachelor's or master's degree, or higher, or equivalent experience in systems engineering, software engineering, or a related engineering field.
- Deep expertise in GPU architecture, hardware acceleration, low-level performance tuning, CUDA, kernels, and cloud-native architectures for multi-tenant model serving.
- Demonstrated success delivering technically complex, high-impact solutions with strong transparency into resource utilization, performance, and operational insights.
- Ability to build consensus and organizational alignment across technical leadership and senior corporate leadership, synthesize cross-functional needs into architecture and design, and guide execution across diverse teams.
- Strong collaboration, communication, and influencing skills, including the ability to work with peers, partners, engineering teams, and accelerated-computing customers.
Preferred Qualifications
- Real-world experience building systems that support artificial intelligence and machine learning workloads.
- Experience designing, developing, delivering, and operating secure, highly available, scaled-out systems in enterprise and cloud environments.
- A history of creating scalable processes and extensible systems that facilitate cross-functional collaboration and operations at scale.
- Familiarity with open-source ecosystems and projects such as Dynamo, TensorRT-LLM, vLLM, SGLang, and Ray, including the ability to influence open-source project governance and technical direction.
Benefits
The role includes competitive salaries, equity, and benefits. NVIDIA is an equal opportunity employer committed to an inclusive work environment.
The base salary range is USD 320,000–488,750. Applications will be accepted at least until August 15, 2026.
More jobs at Nvidia
Senior System Software Engineer, Platform - OpenBMC
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Security Engineer, RTOS and Virtualization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer - NVIDIA Warp
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software And System Architect
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Similar jobs
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Perplexity AI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 250,000-485,000 per year
Forward Deployment Engineering Manager
Nebius · United States
USD 225,800-281,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Machine Learning Engineer, LLM Inference Optimization
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
Principal Software Engineer, E2E Performance and Goodput — CSP Engagements
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Forward Deployed Engineer, Ecosystem
Nebius · United States
USD 208,800-261,000 per year