Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Bash
CUDA @ 6
Computer Vision @ 4
Deep Learning @ 6
Docker @ 6
GPU @ 4
LLM
Linux @ 6
NLP
Networking @ 4
Performance Analysis @ 7
Performance Monitoring @ 4
Profiling @ 4
PyTorch @ 6
Python @ 4
Slurm @ 6
System Architecture @ 7
TensorFlow @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Deep Learning Systems Engineer will analyze the performance and power consumption of deep learning applications on datacenter-class hardware and influence the design and optimization of datacenters. The role involves understanding how CPU, GPU, networking, and I/O relate to deep learning architectures for natural language processing, computer vision, autonomous driving, large language models, and other applications.
Responsibilities
- Develop software infrastructure to characterize and analyze a broad range of deep learning applications.
- Evolve cost-efficient datacenter architectures tailored to the needs of large language models (LLMs).
- Develop analysis and profiling tools using Python, Bash, and C++ to measure key performance metrics of deep learning workloads running on NVIDIA systems.
- Analyze system and software characteristics of deep learning applications.
- Develop analysis tools and methodologies to measure key performance metrics and estimate potential efficiency improvements.
Requirements
- Bachelor's degree in Electrical Engineering or Computer Science, or equivalent experience. A master's or PhD degree is preferred.
- Eight or more years of relevant experience.
- Experience in at least one of the following areas:
- System software, including Linux operating systems, compilers, GPU kernels using CUDA, or deep learning frameworks such as PyTorch and TensorFlow.
- Silicon architecture and performance modeling or analysis, including CPU, GPU, memory, or network architecture.
- Programming experience in C/C++ and Python.
- A deep understanding of computer system architecture and performance analysis, with demonstrated hands-on experience.
- Demonstrated ability to work in virtual environments and independently own tasks from beginning to end.
- Exposure to containerization platforms such as Docker and datacenter workload managers such as Slurm is a plus.
Preferred Qualifications
- Background in system software, operating system intrinsics, GPU kernels using CUDA, or deep learning frameworks such as PyTorch and TensorFlow.
- Experience with silicon performance monitoring or profiling tools such as perf, gprof, nvidia-smi, and DCGM.
- In-depth performance modeling experience in CPU, GPU, memory, or network architecture.
- Exposure to Docker and Slurm.
- Experience working with multisite or multifunctional teams.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to fostering a diverse work environment and is an equal opportunity employer.
Applications will be accepted at least until May 11, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Director, Global Risk and Compliance
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior System Software Engineer - CPU SoC Boot Firmware
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Forward-Deployed Engineer, Enterprise AI and Automation
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Senior HPC Storage Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Development Engineer in Test - Datacenter Server OS
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior System Software Engineer - AI Performance And Efficiency Tools
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year