Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
CI/CD
CUDA @ 4
Debugging @ 4
Distributed Systems @ 7
GPU @ 4
HPC @ 4
JAX @ 3
MPI @ 4
NCCL @ 4
Parallel Programming @ 7
Profiling @ 4
PyTorch @ 3
Python @ 6
TensorFlow @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Software Engineer to lead the development of AI software resiliency for AI supercomputers operating at a scale of more than 100,000 GPUs. The role focuses on reducing cluster downtime and ensuring reliable AI training and inference workloads across cloud and HPC environments.
Responsibilities
- Develop and optimize AI software resiliency features, including fast checkpoint recovery, error detection, error isolation, and straggler or hang detection.
- Write high-quality, production-level C++ and Python code for large-scale distributed systems.
- Improve the performance of AI workloads running across thousands of GPUs.
- Implement error-handling techniques to detect silent data corruption and other failure scenarios.
- Develop monitoring tools for proactive failure mitigation.
- Collaborate with senior engineers, AI researchers, and hardware and software teams to integrate resiliency features into AI frameworks such as PyTorch and JAX/XLA.
- Develop tests to validate the robustness, scalability, and efficiency of resiliency mechanisms.
- Contribute to CI/CD pipelines that automate AI workload validation.
- Assist with debugging and performance tuning of large-scale AI workloads in cloud and HPC environments.
Requirements
- Bachelor's, master's, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- Proficiency in C++ and Python, including experience writing efficient, high-performance code.
- At least 6 years of relevant experience.
- Strong understanding of distributed systems, parallel programming, and fault tolerance in large-scale computing environments.
- Familiarity with AI frameworks such as PyTorch, JAX/XLA, TensorFlow, or similar technologies.
- Experience with debugging and profiling tools such as gdb, perf, Valgrind, or NVIDIA Nsight.
- Excellent problem-solving skills and the ability to work in a fast-paced, highly collaborative environment.
Preferred Qualifications
- Hands-on experience training models or working with model training teams.
- Experience with CUDA, NCCL, or MPI for GPU-accelerated computing, particularly at extreme scale.
- Knowledge of checkpointing strategies, error mitigation, or fault-tolerant computing in AI training.
- Experience with large-scale AI clusters, HPC environments, or cloud-based AI workloads.
- Strong systems programming skills and experience with low-level performance tuning.
Benefits
- Equity and employee benefits are provided.
- NVIDIA is an equal opportunity employer committed to a diverse work environment.
More jobs at Nvidia
Senior Applied Research Scientist – AI Native Numerical Methods
Nvidia · United States
USD 192,000-356,500 per year
Senior Applied Research Scientist, Multimodal Foundation Models – Healthcare
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Business Operations Technical Program Manager — DGX Cloud
Nvidia · Santa Clara, United States
USD 200,000-379,500 per year
Senior Applied Research Scientist – GPU Native Numerical Algorithms
Nvidia · United States
USD 192,000-356,500 per year
Senior Software Engineer, Networking
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
AI Inference Performance Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Training Performance Engineer
OpenAI · San Francisco, United States
USD 250,000-445,000 per year
Senior System Software Engineer - GPU Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior HPC Performance Engineer
Nvidia · Germany
PLN 221,200-507,000 per year
Member of Technical Staff (AI Inference Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States, New York City, United States
USD 220,000-485,000 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year