Senior Research Engineer, Foundation Model Training Infrastructure
at Nvidia
USD 224,000-356,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 8
CUDA @ 7
Debugging
GPU @ 7
HPC @ 7
JAX @ 4
Kubernetes @ 7
LLM @ 7
MLOps @ 8
PyTorch @ 4
Python @ 7
Robotics @ 6
Slurm @ 7
TensorFlow @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a senior or principal engineer specializing in cutting-edge infrastructure for large-scale foundation model training in the Generalist Embodied Agent Research (GEAR) group. The team is leading Project GR00T, NVIDIA's initiative to build foundation models and full-stack technology for humanoid robots.
The role involves working with a collaborative research team focused on multimodal foundation models, large-scale robot learning, embodied AI, and physics simulation. Contributions will impact research projects and product roadmaps.
Responsibilities
- Design and maintain large-scale distributed training systems supporting multimodal foundation models for robotics.
- Optimize GPU and cluster utilization for efficient model training and fine-tuning on massive datasets.
- Implement scalable data loaders and preprocessors for multimodal datasets, including videos, text, and sensor data.
- Develop robust monitoring and debugging tools to ensure the reliability and performance of training workflows on large GPU clusters.
- Collaborate with researchers to integrate cutting-edge model architectures into scalable training pipelines.
Requirements
- Bachelor's degree in Computer Science, Robotics, Engineering, or a related field.
- 10+ years of full-time industry experience in large-scale MLOps and AI infrastructure.
- Proven experience designing and optimizing distributed training systems with frameworks such as PyTorch, JAX, or TensorFlow.
- Deep understanding of GPU acceleration, CUDA programming, and cluster management tools such as Kubernetes.
- Strong programming skills in Python and a high-performance language such as C++ for efficient system development.
- Strong experience with large-scale GPU clusters, HPC environments, and job scheduling or orchestration tools such as SLURM and Kubernetes.
Preferred Qualifications
- Master's or PhD degree in Computer Science, Robotics, Engineering, or a related field.
- Demonstrated Tech Lead experience, coordinating a team of engineers and driving projects from conception to deployment.
- Strong experience building large-scale LLM and multimodal LLM training infrastructure.
- Contributions to popular open-source AI frameworks or research publications in top-tier AI conferences such as NeurIPS, ICRA, ICLR, or CoRL.
Compensation and Benefits
- Base salary range: $224,000–$356,500 USD, determined by location, experience, and compensation for similar positions.
- Eligible for equity and benefits.
Applications will be accepted at least until January 13, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Research Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Software SDET Test Development Engineer
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
AI Inference Performance Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior AI Performance and Efficiency Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Solutions Architect
Nebius · United States, Canada
USD 250,000-320,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year