Senior DGX Cloud AI Infrastructure Software Engineer

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 7 API Communication @ 6 Debugging @ 7 Distributed Systems @ 6 InfiniBand @ 4 JAX @ 4 NCCL @ 4 Observability @ 4 Prometheus @ 4 PyTorch @ 4 Python @ 6 TensorFlow @ 4

Details

Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers innovative AI research. The team develops tools for optimizing the efficiency and resiliency of AI workloads, including pre-training, post-training, and inference. The objective is to deliver a stable, scalable environment for AI researchers and provide the resources and scale needed to foster innovation.

As a Senior DGX Cloud AI Infrastructure Software Engineer, you will design, build, and maintain AI infrastructure that enables large-scale AI training and inferencing. You will apply software and systems engineering practices to ensure high efficiency and availability of AI systems. The role offers autonomy to work on meaningful projects within a dynamic, diverse, and supportive team that values learning, growth, blameless postmortems, iterative improvement, and risk-taking.

Responsibilities

  • Develop infrastructure software and tools for large-scale pre-training, post-training, and inference.
  • Develop and optimize tools and libraries to improve infrastructure efficiency and resiliency.
  • Co-design and implement APIs for integration with NVIDIA's resiliency stacks.
  • Enhance infrastructure and products underpinning NVIDIA's AI platforms.
  • Define meaningful and actionable reliability metrics to track and improve system and service reliability.
  • Apply problem-solving, root-cause analysis, and optimization skills.
  • Analyze root causes and triage failures from the application level to the hardware level.

Requirements

  • At least 8 years of experience developing software infrastructure for large-scale AI systems.
  • Bachelor's degree or higher in Computer Science or a related technical field, or equivalent experience.
  • Strong debugging skills and experience analyzing and triaging AI applications from the application level to the hardware level.
  • Experience with observability platforms for monitoring and logging, such as ELK, Prometheus, and Loki.
  • Proven track record building and scaling large-scale distributed systems.
  • Experience with AI training and inferencing infrastructure services.
  • Proficiency in Python, C/C++, and scripting languages.
  • Experience with quality software engineering practices, including test development, defensive programming, version control, and continuous integration.
  • Excellent communication and collaboration skills, along with intellectual curiosity, problem-solving ability, openness, and a commitment to diversity.

Preferred Qualifications

  • Experience working with large-scale clusters.
  • Experience defining and building observability and telemetry software stacks.
  • Experience with RDMA software stacks, including NCCL, InfiniBand verbs, UCX, and libfabric.
  • Experience performing root-cause analysis of failures at data-center scale.
  • Good understanding of the internals of deep-learning frameworks, including PyTorch, TensorFlow, JAX, and Ray.

Compensation and Benefits

  • Base salary range of $184,000–$287,500 for Level 4.
  • Base salary range of $224,000–$356,500 for Level 5.
  • Salary is determined based on location, experience, and compensation for employees in similar positions.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until April 6, 2026.
  • This posting is for an existing vacancy.
  • NVIDIA uses AI tools in its recruiting processes.
  • NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.

More jobs at Nvidia

Similar jobs