Senior Research Engineer, Foundation Model Training Infrastructure

at Nvidia
USD 224,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 8 CUDA @ 7 Debugging GPU @ 7 HPC @ 7 JAX @ 4 Kubernetes @ 7 LLM @ 7 MLOps @ 8 PyTorch @ 4 Python @ 7 Robotics @ 6 Slurm @ 7 TensorFlow @ 4

Details

NVIDIA is seeking a senior or principal engineer specializing in cutting-edge infrastructure for large-scale foundation model training in the Generalist Embodied Agent Research (GEAR) group. The team is leading Project GR00T, NVIDIA's initiative to build foundation models and full-stack technology for humanoid robots.

The role involves working with a collaborative research team focused on multimodal foundation models, large-scale robot learning, embodied AI, and physics simulation. Contributions will impact research projects and product roadmaps.

Responsibilities

  • Design and maintain large-scale distributed training systems supporting multimodal foundation models for robotics.
  • Optimize GPU and cluster utilization for efficient model training and fine-tuning on massive datasets.
  • Implement scalable data loaders and preprocessors for multimodal datasets, including videos, text, and sensor data.
  • Develop robust monitoring and debugging tools to ensure the reliability and performance of training workflows on large GPU clusters.
  • Collaborate with researchers to integrate cutting-edge model architectures into scalable training pipelines.

Requirements

  • Bachelor's degree in Computer Science, Robotics, Engineering, or a related field.
  • 10+ years of full-time industry experience in large-scale MLOps and AI infrastructure.
  • Proven experience designing and optimizing distributed training systems with frameworks such as PyTorch, JAX, or TensorFlow.
  • Deep understanding of GPU acceleration, CUDA programming, and cluster management tools such as Kubernetes.
  • Strong programming skills in Python and a high-performance language such as C++ for efficient system development.
  • Strong experience with large-scale GPU clusters, HPC environments, and job scheduling or orchestration tools such as SLURM and Kubernetes.

Preferred Qualifications

  • Master's or PhD degree in Computer Science, Robotics, Engineering, or a related field.
  • Demonstrated Tech Lead experience, coordinating a team of engineers and driving projects from conception to deployment.
  • Strong experience building large-scale LLM and multimodal LLM training infrastructure.
  • Contributions to popular open-source AI frameworks or research publications in top-tier AI conferences such as NeurIPS, ICRA, ICLR, or CoRL.

Compensation and Benefits

  • Base salary range: $224,000–$356,500 USD, determined by location, experience, and compensation for similar positions.
  • Eligible for equity and benefits.

Applications will be accepted at least until January 13, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs