Senior Research Engineer, Foundation Model Training Infrastructure
at Nvidia
USD 224,000-356,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 8
CUDA @ 7
Debugging
GPU @ 7
HPC @ 7
JAX @ 4
Kubernetes @ 7
LLM @ 7
MLOps @ 8
PyTorch @ 4
Python @ 7
Robotics @ 6
Slurm @ 7
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a senior or principal engineer specializing in cutting-edge infrastructure for large-scale foundation model training in the Generalist Embodied Agent Research (GEAR) group. The team is leading Project GR00T, NVIDIA's initiative to build foundation models and full-stack technology for humanoid robots.
The role involves working with a collaborative research team focused on multimodal foundation models, large-scale robot learning, embodied AI, and physics simulation. Contributions will impact research projects and product roadmaps.
Responsibilities
- Design and maintain large-scale distributed training systems supporting multimodal foundation models for robotics.
- Optimize GPU and cluster utilization for efficient model training and fine-tuning on massive datasets.
- Implement scalable data loaders and preprocessors for multimodal datasets, including videos, text, and sensor data.
- Develop robust monitoring and debugging tools to ensure the reliability and performance of training workflows on large GPU clusters.
- Collaborate with researchers to integrate cutting-edge model architectures into scalable training pipelines.
Requirements
- Bachelor's degree in Computer Science, Robotics, Engineering, or a related field.
- 10+ years of full-time industry experience in large-scale MLOps and AI infrastructure.
- Proven experience designing and optimizing distributed training systems with frameworks such as PyTorch, JAX, or TensorFlow.
- Deep understanding of GPU acceleration, CUDA programming, and cluster management tools such as Kubernetes.
- Strong programming skills in Python and a high-performance language such as C++ for efficient system development.
- Strong experience with large-scale GPU clusters, HPC environments, and job scheduling or orchestration tools such as SLURM and Kubernetes.
Preferred Qualifications
- Master's or PhD degree in Computer Science, Robotics, Engineering, or a related field.
- Demonstrated Tech Lead experience, coordinating a team of engineers and driving projects from conception to deployment.
- Strong experience building large-scale LLM and multimodal LLM training infrastructure.
- Contributions to popular open-source AI frameworks or research publications in top-tier AI conferences such as NeurIPS, ICRA, ICLR, or CoRL.
Compensation and Benefits
- Base salary range: $224,000–$356,500 USD, determined by location, experience, and compensation for similar positions.
- Eligible for equity and benefits.
Applications will be accepted at least until January 13, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
User Interface - User Experience Designer
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior Application Engineer, HPC and AI for Physics
Nvidia · United States
USD 140,000-270,200 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior System Software Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
Senior Research Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · Palo Alto, United States, San Francisco, United States
USD 220,000-405,000 per year
NVIDIA Spring 2027 Internships: Developer and Performance Technology
Nvidia · Santa Clara, United States
USD 20-71 per hour
Senior Software Development Engineer in Test - Datacenter Server OS
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
NVIDIA 2027 Internships: Deep Learning
Nvidia · Santa Clara, United States
USD 20-71 per hour
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Software Development Engineer in Test - Datacenter Server OS
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior AI Performance and Efficiency Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year