Senior GPU Supercomputer Scheduler Engineer

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Bash @ 4 Communication @ 6 Deep Learning @ 4 Docker @ 4 GPU Go @ 4 HPC Kubernetes @ 7 Linux @ 4 Machine Learning PyTorch @ 4 Python @ 4 Slurm @ 7 Software Development @ 6 TensorFlow @ 4

Details

NVIDIA's Managed AI Research Superclusters (MARS) team builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next generation of AI/ML systems. As a member of the Scheduling team, you will help design and implement GPU compute clusters for demanding deep learning, high-performance computing, and computationally intensive workloads. You will work on production-grade solutions for scheduling simultaneous, large-scale, multi-node GPU workloads with complex requirements and dependencies.

Responsibilities

  • Design and develop scheduling features and add-on services to improve GPU compute clusters, including resource fairness, GPU occupancy, GPU utilization, application resilience, application performance, and power usage.
  • Design and develop batch workload management and orchestration services.
  • Support staff and end users in resolving batch scheduler issues.
  • Build and improve the ecosystem around GPU-accelerated computing.
  • Analyze and optimize the performance of deep learning workflows.
  • Develop large-scale automation solutions.
  • Perform root-cause analysis and recommend corrective actions for problems of varying scale.
  • Identify and resolve problems proactively.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 5+ years of work experience.
  • Strong understanding of batch scheduling, preferably with schedulers such as SLURM or Kubernetes batch schedulers including Kueue and Volcano.
  • Significant experience with systems programming languages such as C/C++ and Go, as well as scripting languages such as Python and Bash.
  • Established experience with the Linux operating system, environment, and tools.
  • Experience analyzing and tuning performance for a variety of AI workloads.
  • In-depth understanding of container technologies such as Docker, Singularity, and Podman.
  • Flexibility and adaptability when working in a dynamic environment with different frameworks and requirements.
  • Excellent communication, interpersonal, and customer collaboration skills.

Preferred Qualifications

  • Knowledge of high-performance computing.
  • Open-source software contributions.
  • Experience with deep learning frameworks such as PyTorch and TensorFlow.
  • Passion for software development processes.

Compensation And Benefits

  • Base salary range: USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4, determined by location, experience, and compensation for employees in similar positions.
  • Eligible for equity and benefits.
  • NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.

More jobs at Nvidia

Similar jobs