Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Bash @ 4
Communication @ 6
Deep Learning @ 4
Docker @ 4
GPU
Go @ 4
HPC
Kubernetes @ 7
Linux @ 4
Machine Learning
PyTorch @ 4
Python @ 4
Slurm @ 7
Software Development @ 6
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Managed AI Research Superclusters (MARS) team builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next generation of AI/ML systems. As a member of the Scheduling team, you will help design and implement GPU compute clusters for demanding deep learning, high-performance computing, and computationally intensive workloads. You will work on production-grade solutions for scheduling simultaneous, large-scale, multi-node GPU workloads with complex requirements and dependencies.
Responsibilities
- Design and develop scheduling features and add-on services to improve GPU compute clusters, including resource fairness, GPU occupancy, GPU utilization, application resilience, application performance, and power usage.
- Design and develop batch workload management and orchestration services.
- Support staff and end users in resolving batch scheduler issues.
- Build and improve the ecosystem around GPU-accelerated computing.
- Analyze and optimize the performance of deep learning workflows.
- Develop large-scale automation solutions.
- Perform root-cause analysis and recommend corrective actions for problems of varying scale.
- Identify and resolve problems proactively.
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 5+ years of work experience.
- Strong understanding of batch scheduling, preferably with schedulers such as SLURM or Kubernetes batch schedulers including Kueue and Volcano.
- Significant experience with systems programming languages such as C/C++ and Go, as well as scripting languages such as Python and Bash.
- Established experience with the Linux operating system, environment, and tools.
- Experience analyzing and tuning performance for a variety of AI workloads.
- In-depth understanding of container technologies such as Docker, Singularity, and Podman.
- Flexibility and adaptability when working in a dynamic environment with different frameworks and requirements.
- Excellent communication, interpersonal, and customer collaboration skills.
Preferred Qualifications
- Knowledge of high-performance computing.
- Open-source software contributions.
- Experience with deep learning frameworks such as PyTorch and TensorFlow.
- Passion for software development processes.
Compensation And Benefits
- Base salary range: USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4, determined by location, experience, and compensation for employees in similar positions.
- Eligible for equity and benefits.
- NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software SDET Test Development Engineer
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior System Software Engineer - AI Performance And Efficiency Tools
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior AI Performance and Efficiency Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
AI/ML Specialist Solutions Architect
Nebius · United States, Canada
USD 250,000-320,000 per year
Solutions Architect
Nebius · United States, Canada
USD 250,000-320,000 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year