Senior Data Center Performance Engineer - Benchmarking and Optimization

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 CUDA @ 3 Docker @ 3 GPU @ 3 HPC @ 4 InfiniBand @ 4 JAX @ 4 Kubernetes @ 3 Linux @ 7 MPI @ 4 Machine Learning NCCL @ 4 NVLink @ 4 Networking @ 4 Parallel Programming @ 3 Performance Monitoring Performance Optimization @ 6 Profiling @ 7 PyTorch @ 4 Python @ 6 Slurm @ 3 System Architecture @ 7 TensorFlow @ 4

Details

NVIDIA is expanding its ecosystem of data center platform designs, from single-node HGX/DGX systems to large multi-node NVLink domain rack architectures. These platforms combine NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and an optimized NVIDIA AI and HPC software stack. The engineer will lead performance benchmarking and optimization efforts for data center products and help ensure industry-leading performance for accelerated computing workloads.

Responsibilities

  • Design and execute comprehensive performance benchmarking strategies for data center platforms and products.
  • Characterize real-world AI training, inference, and HPC workloads at scale.
  • Define, track, and report key performance indicators, including throughput, latency, efficiency, and scaling.
  • Build automation tools and frameworks for performance monitoring and analysis.
  • Identify and analyze performance bottlenecks across compute, memory, network, and storage subsystems.
  • Work closely with architecture, hardware, software, networking, storage, and customer teams to resolve performance issues.
  • Drive performance improvements through system tuning, configuration optimization, and architectural recommendations for future-generation systems.

Requirements

  • M.S. or Ph.D. in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 8+ years of experience in performance engineering or system architecture.
  • Deep understanding of computer architecture, hardware-software interaction, and computing at scale.
  • Strong proficiency with performance profiling tools such as Linux perf and NVIDIA Nsight Systems.
  • Familiarity with GPU computing and parallel programming using CUDA.
  • Experience with HPC networking technologies, including InfiniBand, RoCE, and NVLink.
  • Programming skills in Python, C++, and shell scripting.
  • Excellent analytical and problem-solving abilities.
  • Adaptability and passion for learning new technologies.
  • Ability to communicate effectively and work with cross-functional global teams.

Preferred Qualifications

  • Experience with AI/ML frameworks such as PyTorch, TensorFlow, and JAX.
  • Knowledge of MPI, collective communications, NCCL, distributed training, and distributed inference.
  • Familiarity with NVIDIA DGX, HGX, and other data center solutions.
  • Familiarity with containers, cloud provisioning, and scheduling tools such as Docker, Kubernetes, and SLURM.
  • Understanding of storage systems and I/O performance.
  • Track record of performance optimization in production environments.
  • Experience with AI code generation tools.

Compensation And Benefits

The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Base salary will be determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications for this job will be accepted at least until July 30, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs