Software Engineer, DGX Cloud AI Infrastructure - New College Grad 2026

at Nvidia
USD 108,000-195,500 per year
JUNIOR
✅ On-site

Tech Stack

AI @ 3 CUDA @ 3 Communication @ 3 Debugging @ 3 Deep Learning Distributed Systems @ 3 GPU @ 3 HPC InfiniBand @ 2 LLM NCCL @ 3 Profiling PyTorch Python @ 6 TensorRT

Details

NVIDIA is building the software and systems that power advanced large language model workloads. This role focuses on bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at large scale.

The engineer will work on distributed LLM workloads across multi-GPU and multi-node deployments, while designing and implementing benchmarking tools, automation, and debugging workflows for deep learning systems, GPU performance, distributed computing, and large-scale operations.

Responsibilities

  • Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads.
  • Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks.
  • Perform root-cause analysis of failures in large distributed environments.
  • Contribute to resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across the cluster.
  • Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms.
  • Tune runtime settings, communication parameters, and deployment configurations in partnership with framework, systems, and platform teams.
  • Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization.

Requirements

  • Bachelor's or Master's degree in Computer Science or a related technical field, or equivalent experience.
  • Experience developing software for AI, high-performance computing, or systems-level applications.
  • Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
  • Background debugging and scaling distributed systems.
  • Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware.
  • Experience operating workloads in scheduled, containerized cluster environments.
  • Excellent analytical, debugging, and communication skills, with a collaborative approach across teams.
  • Strong Python and C/C++ programming skills.

Preferred Qualifications

  • Hands-on experience with NCCL and CUDA-aware distributed execution.
  • Deep familiarity with the RDMA software stack, including NCCL, InfiniBand verbs, UCX, and libfabric, as well as InfiniBand / RoCE congestion debugging.
  • Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf.
  • Experience diagnosing performance jitter.
  • Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.

Compensation and Benefits

The base salary range is USD 108,000–178,250 for Level 1 and USD 124,000–195,500 for Level 2. Salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until October 3, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs