Principal High-Performance LLM Training Engineer

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 7 CUDA Communication @ 4 Deep Learning @ 7 Distributed Systems @ 6 GPU @ 7 HPC JAX LLM @ 7 Leadership @ 6 Machine Learning Networking Performance Analysis Performance Optimization @ 6 Profiling @ 7 PyTorch Reinforcement Learning @ 7 Technical Leadership @ 6

Details

NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA's full hardware and software stack. This role sits at the intersection of distributed training, GPU architecture, systems software, deep learning frameworks, and performance engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements across PyTorch, JAX, NeMo, and NeMo RL, and use insights from real workloads to help shape future NVIDIA GPU, system, and software roadmaps.

We are looking for a deeply technical leader who can operate across abstraction layers, from application-level training behavior to framework and runtime internals, CUDA libraries, communication collectives, memory systems, networking, and GPU architecture. Success in this role includes directly improving performance, setting technical direction, raising organizational standards, and influencing multifunctional decisions across NVIDIA.

Responsibilities

  • Lead end-to-end performance analysis and optimization of innovative LLM pre-training and post-training workloads on the latest NVIDIA hardware and software platforms.
  • Drive workloads closer to speed-of-light performance by identifying and removing bottlenecks across compute, memory, communication, scheduling, parallelism strategy, kernel efficiency, framework overhead, and system-level scaling.
  • Develop production-quality software, tools, models, benchmarks, and analysis infrastructure that improve training performance, efficiency, and developer velocity across NVIDIA's AI software stack.
  • Build and refine performance models, workload characterizations, and simulation methodologies to guide future GPU, networking, system, and software architecture decisions.
  • Serve as a technical authority for AI training performance, partnering with teams across GPU architecture, systems, CUDA libraries, compilers, networking, frameworks, product management, and applied AI.
  • Translate workload insights into concrete hardware and software recommendations and advocate for changes that improve performance and efficiency across the AI ecosystem.
  • Mentor engineers and provide technical leadership across the organization, helping establish best practices for large-scale AI performance analysis and optimization.

Requirements

  • Master's degree or PhD, or equivalent experience, in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • At least 12 years of relevant work or research experience.
  • Demonstrated principal-level technical impact in one or more of the following areas: large-scale AI training systems, GPU performance optimization, distributed systems, high-performance computing, ML frameworks, compilers and runtimes, or hardware/software co-design.
  • Deep hands-on experience analyzing and optimizing large-scale deep learning workloads, especially transformer-based models, LLM pre-training, reinforcement learning, fine-tuning, or other post-training workloads.
  • Strong understanding of GPU and AI accelerator architecture, from individual accelerators to data-center-scale systems.
  • Experience with distributed training techniques including data parallelism, tensor parallelism, pipeline parallelism, expert parallelism, sequence parallelism, activation checkpointing, mixed-precision training, and communication/computation overlap.
  • Strong track record using profiling, tracing, benchmarking, and performance-modeling tools to diagnose complex bottlenecks and drive measurable improvements.
  • Excellent communication and technical leadership skills, with the ability to influence architecture and software decisions across multiple teams without relying on direct authority.

Additional Information

GPU computing is a productive and pervasive platform for deep learning and AI. NVIDIA integrates and optimizes deep learning frameworks, works with systems companies and cloud service providers to make GPUs available in data centers and the cloud, and develops computers and software for edge devices such as self-driving cars and autonomous robots.

This opportunity offers collaboration with forward-thinking teams in a creative and autonomous work environment that encourages innovation. The role involves working across the full hardware and software stack, from GPU architecture to application code, to achieve optimal performance.

Applications for this job will be accepted at least until May 2, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.

Benefits

  • Equity and benefits are provided in addition to the base salary.

More jobs at Nvidia

Similar jobs