Software Engineer, DGX Cloud AI Infrastructure

at Nvidia
USD 116,000-224,200 per year
MIDDLE
✅ On-site

Tech Stack

AI @ 5 CUDA @ 3 Communication @ 3 Debugging @ 3 Deep Learning Distributed Systems @ 3 GPU @ 3 HPC @ 5 InfiniBand @ 3 LLM NCCL @ 3 Profiling PyTorch Python @ 6 TensorRT

Details

NVIDIA is seeking a Software Engineer to support the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at large scale. The role focuses on deep learning systems, GPU performance, distributed computing, and large-scale operations.

Responsibilities

  • Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads.
  • Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo/Megatron, TensorRT-LLM, and related NVIDIA AI software stacks.
  • Perform root-cause analysis of failures in large distributed environments.
  • Contribute to resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across clusters.
  • Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows for new platforms.
  • Tune runtime settings, communication parameters, and deployment configurations in partnership with framework, systems, and platform teams.
  • Deliver data-driven recommendations based on profiling, benchmark results, and cluster characterization.

Requirements

  • Bachelor's or Master's degree in Computer Science or a related technical field, or equivalent experience.
  • At least 3 years of experience developing software for AI, HPC, or systems-level applications.
  • Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
  • Background in debugging and scaling distributed systems.
  • Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware.
  • Experience operating workloads in scheduled, containerized cluster environments.
  • Excellent analytical, debugging, and communication skills, with a collaborative approach across teams.
  • Strong Python and C/C++ programming skills.

Preferred Qualifications

  • Hands-on experience with NCCL and CUDA-aware distributed execution.
  • Deep familiarity with the RDMA software stack, including NCCL, IB verbs, UCX, and libfabric.
  • Experience debugging InfiniBand or RoCE congestion.
  • Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf.
  • Experience diagnosing performance jitter.
  • Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.

Compensation And Benefits

The base salary range is USD 116,000–189,750 for Level 2 and USD 140,000–224,250 for Level 3. Compensation is determined by location, experience, and pay for employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until August 15, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs