Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 7
Debugging @ 4
GPU @ 4
Profiling @ 7
PyTorch @ 7
Python @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is building the next generation of AI systems that can perceive, reason about, and generate dynamic worlds. The team advances world foundation models for high-fidelity, temporally stable video and world generation in physical AI, simulation, and interactive experiences. This role operates at the applied-research boundary, developing and validating model improvements and hardening them into production-grade checkpoints and recipes. The technical focus is human appearance, motion, and action understanding, with work delivered in partnership with data, platform, and product engineering teams.
Responsibilities
- Research, implement, and validate model architecture and algorithm changes that improve video-generation fidelity, with an emphasis on human-centric quality.
- Prototype improvements across spatial multimodal modeling, modality alignment, flow-based or diffusion-based video generation, and neural-rendering-inspired representations to improve controllability and long-horizon consistency.
- Improve training and inference efficiency through architectural and post-training techniques, including compute and memory optimization, distillation, pruning, and compression.
- Define model-training objectives that improve sim-to-real and real-to-sim generalization, especially for human motion, contact, and interaction dynamics across real-world and synthetic or simulation data.
- Develop domain-specific benchmarks for evaluating world foundation models, including models that generate and understand video, simulations, and physical environments.
- Translate research results into robust implementations, including training code, production-grade checkpoints, model integrations, and demonstrations that showcase capability gains across teams.
Requirements
- PhD in Computer Science, Graphics, Computer Engineering, or a closely related field, or equivalent experience.
- 8+ years of applied research and/or industry experience in vision, graphics, adjacent machine-learning domains, or a similar area.
- 3+ years of direct experience designing, training, and evaluating generative models for image, video, or audio, with strong modern deep-learning fundamentals.
- Hands-on experience improving generative models with a focus on perceptual quality and temporal stability, especially for human generation.
- Advanced proficiency in Python, PyTorch, C++, and CUDA, along with strong research-engineering practices such as reproducibility, testing, profiling, and experiment tracking.
- Experience training and debugging large models in multi-GPU and/or multi-node environments and distributed-training workflows.
- Practical knowledge of inference and runtime bottlenecks and optimization techniques.
- Strong visual-quality evaluation skills and an interest in diagnosing artifacts such as sharpness, texture detail, and temporal stability using perceptual metrics, human-preference signals, or learned evaluators.
Preferred Qualifications
- A proven research track record, including publications at top conferences such as NeurIPS, CVPR, or ICLR, with clear evidence of impact on model quality or robustness.
- Experience using agentic workflows and AI coding companions to accelerate research and production development, including code generation, debugging, test creation, experiment automation, benchmark development, documentation, and large-codebase navigation.
Benefits
- Competitive salary.
- Equity and comprehensive benefits package.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
Applications will be accepted at least until June 27, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.