Senior Deep Learning Compiler Engineer - XLA

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI Algorithms @ 1 CUDA @ 4 Debugging @ 6 Deep Learning @ 1 GPU @ 7 HPC JAX @ 4 LLVM @ 1 Mentoring @ 1 OpenCL @ 4 Performance Analysis @ 6 PyTorch @ 4 TensorFlow @ 4

Details

NVIDIA is looking for versatile software engineers to join its XLA team and build high-performance, production-grade software at the core of next-generation AI systems.

Responsibilities

  • Develop compiler optimization algorithms for deep learning workloads.
  • Optimize inference and training performance for the JAX framework and the OpenXLA compiler on NVIDIA GPUs at scale.
  • Craft and implement compiler optimization techniques for deep learning network graphs.
  • Design graph partitioning and tensor sharding techniques for distributed training and inference.
  • Perform performance tuning and analysis.
  • Develop code generation for NVIDIA GPU backends using open-source compilers such as MLIR, LLVM, and OpenAI Triton.
  • Design user-facing features in JAX and related libraries, along with other general software engineering work.
  • Collaborate with deep learning framework teams and GPU hardware architecture teams to accelerate next-generation deep learning software.
  • Work closely with GPU hardware engineering teams to design AI compiler software features for next-generation GPUs.

Requirements

  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 4+ years of relevant work or research experience in performance analysis and compiler optimization.
  • Ability to work independently, define project goals and scope, and lead development efforts using clean software engineering and testing practices.
  • Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design.
  • Strong foundation in CPU, GPU, or other high-performance hardware accelerator architecture.
  • Knowledge of high-performance computing and distributed programming.
  • CUDA or OpenCL programming experience is desired but not required.
  • Experience with XLA, TVM, MLIR, LLVM, OpenAI Triton, deep learning models and algorithms, or deep learning framework design is a strong plus.
  • Strong interpersonal skills and the ability to work in a dynamic, product-oriented team.
  • Mentoring experience with junior engineers and interns is a bonus.

Preferred Experience

  • Experience working with deep learning frameworks such as JAX, PyTorch, or TensorFlow.
  • Extensive experience with CUDA or GPUs in general.
  • Experience with open-source compilers such as XLA, LLVM, MLIR, or TVM.

Compensation and Benefits

The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. Compensation is determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.

Applications will be accepted at least until August 18, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs