Software Engineer, DGX Cloud AI Infrastructure - New College Grad 2026
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
CUDA @ 3
Communication @ 3
Debugging @ 3
Deep Learning
Distributed Systems @ 3
GPU @ 3
HPC
InfiniBand @ 2
LLM
NCCL @ 3
Profiling
PyTorch
Python @ 6
TensorRT
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is building the software and systems that power advanced large language model workloads. This role focuses on bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at large scale.
The engineer will work on distributed LLM workloads across multi-GPU and multi-node deployments, while designing and implementing benchmarking tools, automation, and debugging workflows for deep learning systems, GPU performance, distributed computing, and large-scale operations.
Responsibilities
- Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads.
- Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks.
- Perform root-cause analysis of failures in large distributed environments.
- Contribute to resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across the cluster.
- Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms.
- Tune runtime settings, communication parameters, and deployment configurations in partnership with framework, systems, and platform teams.
- Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization.
Requirements
- Bachelor's or Master's degree in Computer Science or a related technical field, or equivalent experience.
- Experience developing software for AI, high-performance computing, or systems-level applications.
- Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
- Background debugging and scaling distributed systems.
- Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware.
- Experience operating workloads in scheduled, containerized cluster environments.
- Excellent analytical, debugging, and communication skills, with a collaborative approach across teams.
- Strong Python and C/C++ programming skills.
Preferred Qualifications
- Hands-on experience with NCCL and CUDA-aware distributed execution.
- Deep familiarity with the RDMA software stack, including NCCL, InfiniBand verbs, UCX, and libfabric, as well as InfiniBand / RoCE congestion debugging.
- Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf.
- Experience diagnosing performance jitter.
- Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.
Compensation and Benefits
The base salary range is USD 108,000–178,250 for Level 1 and USD 124,000–195,500 for Level 2. Salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until October 3, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.