Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
Computer Vision @ 7
Machine Learning @ 7
Robotics @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
At NVIDIA, the world model team is advancing multimodal AI, robotics, and world foundation models for Physical AI. The Senior Research Manager will lead world-model evaluation and benchmarking across NVIDIA's Physical AI model portfolio.
The role will build a team and research agenda for evaluating world models through closed-system evaluations, where the model under test is pluggable, and open-system evaluations, where access to model internals enables deeper diagnostics, causal analysis, and mechanistic evaluation. The goal is to define what makes a world model useful for Physical AI, discover model failures, and turn findings into improved data, training recipes, model roadmaps, and downstream systems.
Responsibilities
- Lead a team of Research Scientists focused on world-model evaluation, benchmarking, and diagnostics for NVIDIA Physical AI models, including world foundation models, world-action models, synthetic data generation systems, robotics, simulation, and embodied AI workflows.
- Define the scientific roadmap for closed-system and open-system evaluation, including open-loop and closed-loop benchmarks, metrics, failure taxonomies, model comparison, and evaluation-to-training feedback loops.
- Develop benchmarks for physical plausibility, temporal consistency, scene dynamics, object permanence, spatial reasoning, action conditioning, affordances, controllability, long-horizon coherence, synthetic data generation quality, and world-action model usefulness.
- Develop open-system and mechanistic evaluation methods using model internals, including representation probing, causal interventions, activation analysis, ablations, sparse autoencoders, attention and feature analysis, and circuit-style diagnostics.
- Drive evaluation-to-model-improvement loops with training, post-training, data curation, simulation, robotics, synthetic data generation, world-action model, and applied research teams. This includes failure discovery, data generation, post-training priorities, model roadmap feedback, and re-evaluation.
- Publish high-quality papers, technical reports, benchmarks, and open-source evaluation artifacts while establishing rigorous standards for validity, reproducibility, dataset hygiene, leakage prevention, and model comparison.
Requirements
- Strong research background in machine learning, computer vision, multimodal AI, robotics, world models, representation learning, model evaluation, or mechanistic interpretability.
- Experience leading research teams, research programs, or cross-functional technical initiatives with measurable scientific and product impact.
- Deep understanding of modern foundation models, including video models, vision-language-action models, diffusion or flow models, self-supervised learning, or world-model architectures.
- Experience designing serious benchmarks, evaluation datasets, metrics, diagnostic tools, or model analysis frameworks for complex machine learning systems.
- Familiarity with world-model evaluation and open-system analysis techniques, such as physical plausibility, temporal consistency, action conditioning, counterfactual reasoning, representation probing, activation patching, causal interventions, sparse autoencoders, or feature attribution.
- PhD or equivalent experience in Computer Science, Electrical Engineering, Robotics, Machine Learning, AI, or a related field.
- 12 or more years of overall relevant research or engineering experience and 5 or more years of management experience.
- Ability to work onsite at NVIDIA's Santa Clara headquarters; this is not a remote position.
Preferred Qualifications
- Experience building influential benchmarks, evaluation suites, model diagnostics, or interpretability tools used by research or production teams.
- Publications in areas such as world models, video generation, physical AI, embodied AI, robotics, representation learning, mechanistic interpretability, self-supervised learning, or model evaluation.
- Experience evaluating generative video models, action-conditioned world models, robotics foundation models, world-action models, synthetic data generation systems, simulation systems, or vision-language-action models.
- A strong point of view on what current benchmarks miss and enthusiasm for building the next generation of evaluation science for Physical AI.
Benefits
The role includes eligibility for equity and NVIDIA benefits. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.
Applications for this job will be accepted at least until June 11, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.