Distinguished Engineer, Scaled Out Inferencing

at Nvidia
USD 320,000-488,800 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 CUDA @ 6 Communication @ 7 Distributed Systems @ 4 GPU @ 6 Kubernetes @ 6 LLM @ 6 Leadership @ 4 Linux @ 6 Machine Learning @ 4 SGLang @ 6 Technical Leadership @ 6 TensorRT @ 6 vLLM @ 6

Details

NVIDIA is seeking a technology leader to define and drive the global strategy for scaled-out AI inferencing. The role involves architecting high-throughput, low-latency distributed pipelines and model-serving strategies for massive-scale production workloads. You will define the technical roadmap for the full model lifecycle, including deployment, versioning, automated scaling, and operations across enterprise and cloud environments.

Responsibilities

  • Architect and drive the technical implementation of high-throughput, low-latency distributed inference systems for massive-scale AI workloads.
  • Lead hardware-software co-optimization, including performance tuning at the kernel and driver level, GPU resource management, and hardware acceleration for production-grade model serving.
  • Guide and influence open-source and ecosystem projects, including Dynamo, TensorRT-LLM, vLLM, SGLang, Linux, Kubernetes, and Ray.
  • Lead the strategy for full-lifecycle model management, including automated deployment, versioning, and intelligent scaling across cloud and datacenter environments.
  • Collaborate with customers, infrastructure providers, and partners to ensure NVIDIA solutions achieve industry-leading performance and availability.
  • Lead all technical aspects of a large scope across ideation, architecture, design, development, deployment, operations, and continuous lifecycle management.

Requirements

  • 16 or more years in technical roles, with a long-term focus on AI infrastructure and recent direct experience in large-scale inference orchestration.
  • Proven experience building secure, highly available, and durable production distributed systems.
  • 7–10 or more years of leadership experience.
  • Bachelor's or master's degree, or higher, or equivalent experience in systems engineering, software engineering, or a related engineering field.
  • Deep expertise in GPU architecture, hardware acceleration, low-level performance tuning, CUDA, kernels, and cloud-native architectures for multi-tenant model serving.
  • Demonstrated success delivering technically complex, high-impact solutions with strong transparency into resource utilization, performance, and operational insights.
  • Ability to build consensus and organizational alignment across technical leadership and senior corporate leadership, synthesize cross-functional needs into architecture and design, and guide execution across diverse teams.
  • Strong collaboration, communication, and influencing skills, including the ability to work with peers, partners, engineering teams, and accelerated-computing customers.

Preferred Qualifications

  • Real-world experience building systems that support artificial intelligence and machine learning workloads.
  • Experience designing, developing, delivering, and operating secure, highly available, scaled-out systems in enterprise and cloud environments.
  • A history of creating scalable processes and extensible systems that facilitate cross-functional collaboration and operations at scale.
  • Familiarity with open-source ecosystems and projects such as Dynamo, TensorRT-LLM, vLLM, SGLang, and Ray, including the ability to influence open-source project governance and technical direction.

Benefits

The role includes competitive salaries, equity, and benefits. NVIDIA is an equal opportunity employer committed to an inclusive work environment.

The base salary range is USD 320,000–488,750. Applications will be accepted at least until August 15, 2026.

More jobs at Nvidia

Similar jobs