Senior Systems Software Engineer - Deep Learning Solutions

at Nvidia
USD 224,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 3 CUDA @ 4 Deep Learning @ 7 GPU @ 4 Linux @ 4 Machine Learning Parallel Programming @ 4 Performance Optimization @ 4 Robotics TensorRT @ 4

Details

NVIDIA is a global leader in physical AI, powering self-driving cars, humanoid robots, intelligent environments, and medical devices. Its software platforms help innovators build products that save lives, enhance working conditions, and improve living standards globally.

This role focuses on optimizing deep learning inference for autonomous vehicles and robotics on edge devices. The position requires a hands-on specialist who can examine model architectures at the operator level, identify performance issues through kernel trace analysis, and evaluate modern architectures—including transformers, vision-language models, diffusion and flow-matching models, and state-space models—on GPUs and system-on-chip platforms. The work spans model frameworks, compiler technology, and embedded hardware, with collaboration across automotive OEMs, robotics teams, compiler and runtime groups, and internal hardware teams.

Responsibilities

  • Engage directly with automotive OEMs and robotics partners to analyze, debug, and improve deep learning models on NVIDIA platforms.
  • Own performance benchmarking efforts for MLPerf Edge, industry benchmarks, and partner engagements. Define methodologies, ensure reproducibility, and translate results into optimization priorities.
  • Evaluate emerging deep learning architectures, including vision encoders, multimodal vision-language models, hybrid state-space model and Transformer backbones, diffusion and flow-matching decoders, and multi-camera tokenizers.
  • Assess compilation feasibility, memory footprint, and latency on target system-on-chip platforms.
  • Collaborate with compiler, runtime, and hardware teams to connect model-level insights with platform capabilities.
  • Contribute to build reviews and help define internal roadmap priorities based on customer workload patterns.
  • Represent NVIDIA's deep learning optimization expertise at conferences, webinars, and partner events.
  • Build and deploy inference solutions on Jetson, DRIVE, and GPU plus ARM platforms for autonomous vehicle and robotics workloads.
  • Develop Proofs of Readiness and collaborate with the compiler team on Torch-TRT, MLIR-TRT, and related frameworks.

Requirements

  • Master's degree or equivalent experience in Computer Science, Electrical Engineering, or a related field.
  • More than 12 years of industry experience, including at least 8 years specializing in deep learning model optimization, inference engineering, or neural network compilation.
  • Proficiency in understanding and reviewing model architectures at the operator and kernel level.
  • More than 5 years of expertise in embedded or edge software, including production inference solutions in power-limited and latency-sensitive environments.
  • Comprehensive knowledge of contemporary deep learning architectures, including Transformers, attention variants, vision encoders such as ViT, multimodal and vision-language model frameworks, diffusion models, and/or state-space models.
  • Expert knowledge of GPU architecture fundamentals, CUDA, and low-level performance optimization using heterogeneous computing.
  • Experience with TensorRT, compiler intermediate representations, or equivalent inference optimization toolchains.
  • Understanding of embedded operating system internals, including QNX and Linux, memory management, C/C++, and embedded and systems software concepts.
  • Experience with parallel programming, such as CUDA and OpenMP, and with memory hierarchies, data movement, and compute utilization.
  • Ability to collaborate directly with external partners and customers in a deeply technical role and solve workload issues within production constraints.

Preferred Qualifications

  • Experience with ML compiler frameworks such as TVM, MLIR, XLA, or Triton, or with inference runtime development.
  • Production deployment experience with autonomous vehicle perception or planning stacks, from sensor input through trajectory output.
  • Familiarity with physical AI models, including vision-language model and action-expert architectures, end-to-end driving models, or robot foundation models.
  • Contributions to MLPerf benchmarks or large-scale industry performance optimization efforts.
  • Experience with automotive safety standards such as ISO 26262 and SOTIF.

Compensation and Benefits

The base salary range is USD 224,000–356,500, determined by location, experience, and compensation for similar positions. The role also includes eligibility for equity and benefits.

Applications will be accepted at least until March 15, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs