Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
CI/CD
CUDA @ 4
Debugging @ 4
Distributed Systems @ 7
GPU @ 4
HPC @ 4
JAX @ 3
MPI @ 4
NCCL @ 4
Parallel Programming @ 7
Profiling @ 4
PyTorch @ 3
Python @ 6
TensorFlow @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are now looking for a Senior Software Engineer for AI Resiliency.
At NVIDIA, we are pushing the boundaries of what’s possible in AI. We are currently seeking a Senior Software Engineer to lead the development of AI software resiliency for the most powerful AI supercomputers in the world. As a member of our AI Software Resiliency team, you will play a pivotal role in defining and implementing critical resiliency features for AI supercomputers at a scale of 100,000+ GPUs. Your expertise will be crucial in driving down cluster downtime towards zero, ensuring that our AI systems remain robust and reliable at all times.
Responsibilities
- Develop AI Software Resiliency Features: Implement and optimize software features that improve AI system reliability at a massive scale, such as fast checkpoint-recovery, error detection, error isolation, and straggler/hang detection.
- Hands-On Coding & Optimization: Contribute to large-scale distributed systems with high-quality, production-level C++ and Python code. Enhance performance for AI workloads running on thousands of GPUs.
- Fault Tolerance & Debugging: Work on AI system error handling, implementing techniques to detect silent data corruption (SDC) and other failure scenarios. Assist in developing monitoring tools for proactive failure mitigation.
- Collaborate Across Teams: Work closely with senior engineers, AI researchers, and hardware/software teams to integrate resiliency features into AI frameworks like PyTorch and JAX/XLA.
- Testing & Automation: Develop and implement tests to ensure robustness, scalability, and efficiency of resiliency mechanisms. Contribute to CI/CD pipelines to automate validation of AI workloads.
- Support Production Deployments: Assist in debugging and performance tuning large-scale AI workloads in cloud and HPC environments, ensuring seamless operation of AI training and inference workloads.
Requirements
- You’ve achieved a Bachelor’s, Master’s or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- Proficiency in C++ and Python, with experience in writing efficient, high-performance code.
- 6+ years of relevant experience.
- Strong understanding of distributed systems concepts, parallel programming, and fault tolerance in large-scale computing environments.
- Familiarity with AI frameworks such as PyTorch, JAX/XLA, TensorFlow, or similar.
- Experience with debugging and profiling tools (e.g., gdb, perf, valgrind, NVIDIA Nsight).
- Excellent problem-solving skills and ability to work in a fast-paced, highly collaborative environment.
Ways to Stand Out From the Crowd
- Hands-on experience in training models or working with model training teams.
- Hands-on experience with CUDA, NCCL, or MPI for GPU-accelerated computing, especially at extreme-scale.
- Knowledge of checkpointing strategies, error mitigation, or fault-tolerant computing in AI training.
- Experience working with large-scale AI clusters, HPC environments, or cloud-based AI workloads.
- Strong systems programming skills and experience with low-level performance tuning.
Base salary range: 184,000 USD - 287,500 USD.
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Training Performance Engineer
OpenAI · San Francisco, United States
USD 250,000-445,000 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Senior HPC Performance Engineer
Nvidia · Germany
PLN 221,200-507,000 per year
Member Of Technical Staff (Ai Inference Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States, New York City, United States
USD 220,000-485,000 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Senior Software Architect - Deep Learning And Hpc Communications
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year