Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 1
CUDA @ 4
Debugging @ 6
Deep Learning @ 1
GPU @ 7
HPC
JAX @ 4
LLVM @ 1
Mentoring @ 1
OpenCL @ 4
Performance Analysis @ 6
PyTorch @ 4
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for versatile software engineers to join its XLA team and build high-performance, production-grade software at the core of next-generation AI systems.
Responsibilities
- Develop compiler optimization algorithms for deep learning workloads.
- Optimize inference and training performance for the JAX framework and the OpenXLA compiler on NVIDIA GPUs at scale.
- Craft and implement compiler optimization techniques for deep learning network graphs.
- Design graph partitioning and tensor sharding techniques for distributed training and inference.
- Perform performance tuning and analysis.
- Develop code generation for NVIDIA GPU backends using open-source compilers such as MLIR, LLVM, and OpenAI Triton.
- Design user-facing features in JAX and related libraries, along with other general software engineering work.
- Collaborate with deep learning framework teams and GPU hardware architecture teams to accelerate next-generation deep learning software.
- Work closely with GPU hardware engineering teams to design AI compiler software features for next-generation GPUs.
Requirements
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 4+ years of relevant work or research experience in performance analysis and compiler optimization.
- Ability to work independently, define project goals and scope, and lead development efforts using clean software engineering and testing practices.
- Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design.
- Strong foundation in CPU, GPU, or other high-performance hardware accelerator architecture.
- Knowledge of high-performance computing and distributed programming.
- CUDA or OpenCL programming experience is desired but not required.
- Experience with XLA, TVM, MLIR, LLVM, OpenAI Triton, deep learning models and algorithms, or deep learning framework design is a strong plus.
- Strong interpersonal skills and the ability to work in a dynamic, product-oriented team.
- Mentoring experience with junior engineers and interns is a bonus.
Preferred Experience
- Experience working with deep learning frameworks such as JAX, PyTorch, or TensorFlow.
- Extensive experience with CUDA or GPUs in general.
- Experience with open-source compilers such as XLA, LLVM, MLIR, or TVM.
Compensation and Benefits
The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. Compensation is determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until August 18, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Senior Deep Learning Algorithm Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Security Engineer, Detection Engineering
Nvidia · United States
USD 168,000-310,500 per year
Senior Agentic AI Software Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Software Engineer, Deep Learning Libraries - New College Graduate 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior AI Compiler Engineer, MLIR
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Similar jobs
Senior Compiler Engineer, AI Inference Platforms
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Deep Learning Systems Architect
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior AI Compiler Engineer, Algorithms and Code Generation
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Deep Learning Tools Engineer – CUDA Tile
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Deep Learning Compiler Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Research Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Distinguished Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year