Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 5
CUDA @ 3
Communication @ 3
Debugging @ 3
Deep Learning
Distributed Systems @ 3
GPU @ 3
HPC @ 5
InfiniBand @ 3
LLM
NCCL @ 3
Profiling
PyTorch
Python @ 6
TensorRT
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Software Engineer to support the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at large scale. The role focuses on deep learning systems, GPU performance, distributed computing, and large-scale operations.
Responsibilities
- Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads.
- Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo/Megatron, TensorRT-LLM, and related NVIDIA AI software stacks.
- Perform root-cause analysis of failures in large distributed environments.
- Contribute to resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across clusters.
- Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows for new platforms.
- Tune runtime settings, communication parameters, and deployment configurations in partnership with framework, systems, and platform teams.
- Deliver data-driven recommendations based on profiling, benchmark results, and cluster characterization.
Requirements
- Bachelor's or Master's degree in Computer Science or a related technical field, or equivalent experience.
- At least 3 years of experience developing software for AI, HPC, or systems-level applications.
- Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
- Background in debugging and scaling distributed systems.
- Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware.
- Experience operating workloads in scheduled, containerized cluster environments.
- Excellent analytical, debugging, and communication skills, with a collaborative approach across teams.
- Strong Python and C/C++ programming skills.
Preferred Qualifications
- Hands-on experience with NCCL and CUDA-aware distributed execution.
- Deep familiarity with the RDMA software stack, including NCCL, IB verbs, UCX, and libfabric.
- Experience debugging InfiniBand or RoCE congestion.
- Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf.
- Experience diagnosing performance jitter.
- Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.
Compensation And Benefits
The base salary range is USD 116,000–189,750 for Level 2 and USD 140,000–224,250 for Level 3. Compensation is determined by location, experience, and pay for employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until August 15, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Senior Software Solutions Engineer
Nvidia · Poland
PLN 230,200-487,500 per year
Senior Offensive Security Engineer, Automotive
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Systems Operations and Administrator
Nvidia · Santa Clara, United States
USD 112,000-218,500 per year
Senior Software Engineer, Agent Simulation and Evaluation
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior AI Product Engineer
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Similar jobs
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
AI Inference Performance Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior AI Performance and Efficiency Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
DL Performance Software Engineer - LLM Inference
Nvidia · Toronto, Canada
CAD 135,000-220,000 per year
Member of Technical Staff (AI Inference Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States, New York City, United States
USD 220,000-485,000 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year