Senior Software Engineer, AI Resiliency

at Nvidia
USD 184,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 3 CI/CD CUDA @ 4 Debugging @ 4 Distributed Systems @ 7 GPU @ 4 HPC @ 4 JAX @ 3 MPI @ 4 NCCL @ 4 Parallel Programming @ 7 Profiling @ 4 PyTorch @ 3 Python @ 6 TensorFlow @ 3

Details

NVIDIA is seeking a Senior Software Engineer to lead the development of AI software resiliency for AI supercomputers operating at a scale of more than 100,000 GPUs. The role focuses on reducing cluster downtime and ensuring reliable AI training and inference workloads across cloud and HPC environments.

Responsibilities

  • Develop and optimize AI software resiliency features, including fast checkpoint recovery, error detection, error isolation, and straggler or hang detection.
  • Write high-quality, production-level C++ and Python code for large-scale distributed systems.
  • Improve the performance of AI workloads running across thousands of GPUs.
  • Implement error-handling techniques to detect silent data corruption and other failure scenarios.
  • Develop monitoring tools for proactive failure mitigation.
  • Collaborate with senior engineers, AI researchers, and hardware and software teams to integrate resiliency features into AI frameworks such as PyTorch and JAX/XLA.
  • Develop tests to validate the robustness, scalability, and efficiency of resiliency mechanisms.
  • Contribute to CI/CD pipelines that automate AI workload validation.
  • Assist with debugging and performance tuning of large-scale AI workloads in cloud and HPC environments.

Requirements

  • Bachelor's, master's, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • Proficiency in C++ and Python, including experience writing efficient, high-performance code.
  • At least 6 years of relevant experience.
  • Strong understanding of distributed systems, parallel programming, and fault tolerance in large-scale computing environments.
  • Familiarity with AI frameworks such as PyTorch, JAX/XLA, TensorFlow, or similar technologies.
  • Experience with debugging and profiling tools such as gdb, perf, Valgrind, or NVIDIA Nsight.
  • Excellent problem-solving skills and the ability to work in a fast-paced, highly collaborative environment.

Preferred Qualifications

  • Hands-on experience training models or working with model training teams.
  • Experience with CUDA, NCCL, or MPI for GPU-accelerated computing, particularly at extreme scale.
  • Knowledge of checkpointing strategies, error mitigation, or fault-tolerant computing in AI training.
  • Experience with large-scale AI clusters, HPC environments, or cloud-based AI workloads.
  • Strong systems programming skills and experience with low-level performance tuning.

Benefits

  • Equity and employee benefits are provided.
  • NVIDIA is an equal opportunity employer committed to a diverse work environment.

More jobs at Nvidia

Similar jobs