Principal Software Engineer, E2E Performance and Goodput — CSP Engagements

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI CUDA Communication @ 4 Data Analysis @ 7 GPU @ 8 HPC @ 8 LLM @ 4 Leadership @ 6 Machine Learning NCCL NVLink @ 3 Pandas @ 7 Performance Analysis @ 4 Profiling @ 4 Python @ 7 SGLang @ 4 TensorRT @ 4 vLLM @ 4

Details

We're looking for a Principal Engineer to join NVIDIA's CSP Engagements team as the technical focal point for end-to-end performance. You will work directly with engineering teams at key CSP and hyperscale customers to help them achieve performance targets on NVIDIA platforms. The role augments NVIDIA's performance and benchmark teams with a dedicated CSP-facing focus, driving shared understanding of platform performance characteristics, incorporating workload-specific feedback into NVIDIA's optimization priorities, and validating performance targets in customer-representative configurations.

Cross-CSP visibility will enable you to identify patterns and drive systemic improvements in documentation, configuration guidance, and tooling.

Responsibilities

  • Drive performance characterization work streams with engineering teams at key CSP and hyperscale customers, ensuring they understand platform performance expectations, profiling methodology, and tuning options for their workloads.
  • Gather and synthesize CSP performance feedback, identify gaps between expected and actual throughput, and champion optimization priorities with NVIDIA's CUDA, NCCL, driver, and firmware teams.
  • Ensure open-source performance and stress tools such as STREAM, GPU Burn, and GPU BLAST are updated and validated for the latest NVIDIA rack-scale systems, GPU architectures, and CPU platforms.
  • Work with CSPs to ensure their performance and validation tooling reflects the latest GPU capabilities, memory hierarchy changes, and platform-specific tuning parameters.
  • Conduct cross-CSP performance comparisons and pattern analysis to identify configuration, software, or workload differences that explain performance gaps between deployments.
  • Collaborate with CSPs to ensure performance-related integration work, including profiling infrastructure, benchmark harnesses, and configuration validation, is ready ahead of deployment milestones.
  • Define test strategies and tooling requirements for performance validation, both for NVIDIA internal certification and customer acceptance.

Requirements

  • 15+ years of experience in systems performance engineering, ideally in GPU, HPC, or ML infrastructure.
  • Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • Proficiency in GPU workload profiling using Nsight Systems, Nsight Compute, DCGM metrics, or equivalent instrumentation.
  • Understanding of distributed training performance dynamics, including computation and communication overlap, pipeline bubbles, memory bandwidth utilization, and collective efficiency.
  • Knowledge of statistical methods for performance analysis, including regression detection, confidence intervals, and large-scale A/B comparison.
  • Understanding of how the full software stack affects performance, including driver overhead, collective algorithm selection, memory allocation, scheduling, and firmware power management.
  • Strong data analysis and visualization skills using Python, pandas, and dashboards.
  • Customer-focused approach and passion for understanding and resolving performance issues.
  • Ability to communicate performance findings to both highly technical audiences and executive leadership.
  • Demonstrated success influencing multiple engineering teams to prioritize performance improvements.

Preferred Qualifications

  • Experience profiling and optimizing distributed training at 1,000+ GPU scale using Megatron-LM, DeepSpeed, or FSDP.
  • Background in ML infrastructure performance at a CSP or hyperscaler.
  • Familiarity with NVIDIA platforms, including DGX, HGX, and NVLink topology, as well as NVIDIA profiling tools.
  • Experience building automated performance regression detection systems for production environments.
  • Understanding of inference workload performance dynamics, including vLLM, TensorRT-LLM, SGLang, and continuous batching.

About NVIDIA

NVIDIA is leading developments in artificial intelligence, high-performance computing, and visualization. The GPU, NVIDIA's invention, is at the heart of its products and services. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

Compensation and Additional Information

The base salary range is USD 272,000–431,250. Salary is determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.

Applications will be accepted at least until August 1, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs