Senior System Software Engineer - AI Performance And Efficiency Tools
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 4
Communication @ 7
Debugging @ 7
Deep Learning @ 4
GPU @ 4
Kubernetes @ 4
Linux @ 4
NCCL @ 4
Networking @ 4
Profiling @ 4
PyTorch @ 4
Python @ 6
Slurm @ 4
Software Development @ 7
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a forward-thinking and creative software engineer to develop sophisticated analysis and debugging tools that help improve the performance and power efficiency of AI workloads and GPU clusters. The role involves collaborating with AI researchers, architecture teams, and software teams to provide intuitive, accurate insights into workloads and systems, identify software and hardware improvements, build high-level models, and debug complex performance and efficiency issues.
Responsibilities
- Build internal profiling and analysis tools for large-scale AI workloads.
- Build debugging tools for common problems involving memory and networking.
- Create benchmarking and simulation technologies for AI systems and GPU clusters.
- Partner with hardware architects to propose new features and improve existing features using real-world use cases.
- Collaborate concurrently with multiple global groups and users across different departments.
Requirements
- Bachelor's degree or higher in Computer Science or a related field, or equivalent experience.
- At least 6 years of software development experience.
- Strong software design, coding, analytical, and debugging skills.
- Proficiency in C++ and Python.
- Good understanding of deep learning frameworks such as PyTorch and TensorFlow, including distributed training and inference.
- Knowledge of GPU cluster job scheduling with Slurm or Kubernetes, as well as storage and networking.
- Experience with NVIDIA GPUs, CUDA programming, and NCCL.
- Strong problem-solving skills and customer-facing communication skills.
- Motivation for continuous learning and the ability to work with multiple global groups.
Preferred Qualifications
- Experience with GPU-cluster-scale continuous profiling and analysis tools or platforms.
- Experience analyzing the performance of large-scale AI training and inference workloads.
- Knowledge of Linux device drivers and/or compiler implementation.
- Knowledge of GPU and/or CPU architecture and general computer architecture principles.
Compensation And Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits. NVIDIA is an equal opportunity employer.