Senior Systems Software Engineer, AI Stack and Performance - DGX Station
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 6
Deep Learning @ 4
GPU @ 4
JAX @ 7
LLM @ 4
Machine Learning
Microservices
NCCL @ 4
NVLink @ 7
Performance Analysis @ 4
Product Management @ 4
Profiling @ 4
PyTorch @ 7
Python @ 6
TensorFlow @ 7
TensorRT @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
DGX Station (Galaxy) is NVIDIA’s workstation-class AI computer, built on GB300 Blackwell GPUs with NVLink interconnect. It delivers data-center-grade AI compute in a deskside form factor and is shipped to OEM and OSV partners as a complete software and firmware GA release, including firmware bundles, DGX BaseOS, GPU drivers, the CUDA toolkit, DCGM, and DOCA/OFED.
This role owns AI stack readiness on DGX Station. The engineer will profile workloads, identify bottlenecks across GPU compute, NVLink, memory, and host interconnects, drive optimizations across the full stack from GPU kernels through frameworks to applications, and collaborate with framework, compiler, and GPU architecture teams to deliver production-ready performance for real AI workloads in multi-user and multi-GPU configurations.
Responsibilities
- Own production readiness of AI applications on DGX Station, including NemoClaw, Hermes agents, NIM microservices, and key customer workloads.
- Define “ready to ship” criteria, run validation, and close gaps across single-GPU and multi-GPU configurations.
- Profile and optimize LLM and deep learning workloads using PyTorch, TensorFlow, and JAX across training and inference on the GB300 Blackwell multi-GPU architecture.
- Characterize performance across model sizes, batch sizes, FP16, INT8, and FP8 precision modes, and single-GPU versus NVLink-connected multi-GPU scaling.
- Identify bottlenecks in GPU compute, NVLink bandwidth, host memory, PCIe, and CPU–GPU communication.
- Drive optimizations involving kernel tuning, memory placement, NVLink utilization, data pipeline efficiency, and scheduling.
- Collaborate with framework, compiler, and GPU architecture teams on TensorRT, NVCC, and Triton improvements, including kernel fusion, graph execution, operator scheduling, and memory management.
- Validate multi-user and concurrent workload scenarios, including simultaneous training jobs, inference serving, development workloads, MIG, and time-slicing.
- Validate the NVIDIA AI software stack, including CUDA toolkit, cuDNN, TensorRT, NCCL, Triton Inference Server, DCGM, and DOCA/OFED.
- Build and maintain automated performance benchmarking and regression tracking for models such as LLaMA, GPT, Stable Diffusion, and Whisper.
- Work with product management and OEM/OSV partners to understand target use cases and support customer deployment readiness and critical field issues.
Requirements
- Bachelor’s or master’s degree, or equivalent experience, in Computer Science, Electrical Engineering, or a related field.
- 12 or more years of systems software engineering experience with hands-on experience in AI/ML workload optimization, GPU performance analysis, or deep learning infrastructure.
- Strong proficiency with PyTorch, TensorFlow, or JAX, including graph execution, operator dispatch, memory management, and custom kernel integration.
- Experience profiling and optimizing GPU workloads using Nsight Systems, Nsight Compute, CUPTI, or equivalent tools.
- Ability to read GPU traces and translate findings into actionable optimizations.
- Strong understanding of GPU architecture, including compute units, memory hierarchy, NVLink, multi-GPU scaling, and their impact on AI workload performance.
- Experience with inference optimization, including INT8/FP8 quantization, TensorRT,
torch.compile, batching strategies, and serving frameworks. - Proficiency in C/C++, CUDA, and Python, with the ability to read and modify GPU kernels.
Preferred Qualifications
- Experience optimizing LLM training or inference on multi-GPU NVIDIA systems such as DGX, HGX, or multi-GPU workstations.
- Contributions to open-source AI frameworks, CUDA libraries, or inference engines.
- Experience with NCCL tuning, NVLink utilization, collective operations, and parallel training strategies.
- Experience collaborating with compiler and hardware architecture teams on kernel fusion, graph optimization, or hardware-specific performance improvements.
- Experience shipping AI-powered products where application performance on specific hardware was a hard shipping requirement.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.