Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 4
Communication @ 6
Deep Learning @ 4
GPU @ 4
JAX @ 4
LLM @ 4
Performance Analysis @ 4
Profiling @ 4
PyTorch @ 4
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Deep Learning model performance engineering team is hiring software engineers to build and optimize the libraries and tools that enable Deep Learning researchers and engineers to design, develop, and deploy efficient AI applications. The team builds optimizations directly into mainstream open-source Deep Learning frameworks, including PyTorch and JAX, to improve performance across NVIDIA's AI stack. The role involves collaboration across NVIDIA and with the broader open-source community to deliver state-of-the-art Deep Learning performance on NVIDIA platforms.
Responsibilities
- Build and support Transformer Engine, the open-source library for accelerating the training of Large Language Models.
- Collaborate on systems research that improves Deep Learning model performance, including extremely low-precision training and parallelism methods.
- Implement, benchmark, and optimize new Deep Learning models, including LLMs, based on groundbreaking research and scale them efficiently on NVIDIA GPUs and systems.
- Build and contribute to NVIDIA submissions for community benchmarks such as MLPerf.
- Engage with the open-source community and support enterprise customers and partners by delivering the benefits of NVIDIA's latest hardware and software innovations.
- Influence the design of new hardware generations and core platform software components for NVIDIA hardware and systems.
Requirements
- Bachelor's degree or equivalent experience in Computer Science, Electrical Engineering, or a related field.
- At least 3 years of experience with C++ and Python programming.
- Strong background, experience, or coursework in parallel systems programming, preferably on GPUs.
- Knowledge of computer architecture, code optimization, and/or operating systems.
- Proven experience developing large software projects.
- Excellent verbal and written communication skills.
Preferred Qualifications
- Experience with PyTorch, JAX, or another Deep Learning framework.
- Experience with performance analysis, profiling, and code optimization techniques, especially for multi-GPU or multi-node systems.
- Knowledge of modern LLM architectures, attention mechanisms, and/or low-level Deep Learning libraries such as cuBLAS, cuDNN, and cuSOLVER.
- Experience writing GPU kernels using CUDA, OpenAI Triton, CuTeDSL, Pallas, or similar libraries.
- Contributions to the open-source community and/or experience working with multidisciplinary teams.
Compensation and Benefits
The base salary is determined by location, experience, and the pay of employees in similar positions. The base salary ranges are USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until March 8, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is committed to fostering a diverse work environment and providing equal employment opportunities.