Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
Agentic AI @ 4
CI/CD @ 4
Deep Learning
Distributed Systems @ 7
Experimentation @ 7
GPU @ 7
Git @ 6
GitHub @ 6
Jira @ 6
Kubernetes @ 4
Performance Analysis @ 4
QA @ 4
Reinforcement Learning @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA’s Deep Learning Software team is looking for a Senior Technical Program Manager to lead programs across model pre-training, production reinforcement learning runs, evaluation, and agentic AI infrastructure. The team builds software foundations that help research and engineering teams train, evaluate, and deliver sophisticated AI models, partnering across research, platform engineering, distributed computing, evaluation, and open-source development.
Responsibilities
- Lead multifunctional programs across training frameworks, evaluation environments, agent and model runtimes, datasets, verifiers, and distributed training infrastructure.
- Partner with AI researchers, engineering leaders, product, infrastructure, and QA teams to define roadmaps, achievements, release plans, and measurable success criteria.
- Coordinate large-scale reinforcement learning training and evaluation experiments, including GPU resources, dependency tracking, run scheduling, results reporting, release readiness, technical decisions, integration plans, and program updates.
Requirements
- Bachelor’s degree in computer science, engineering, or a related technical field, or equivalent experience.
- 10+ years of technical program management, engineering program management, or related experience delivering sophisticated software platforms.
- Experience leading global, matrixed programs across research, software engineering, infrastructure, QA, release teams, and partner groups.
- Strong understanding of the AI model lifecycle, including training, post-training, evaluation, experimentation, production readiness, reinforcement learning concepts, and GPU-accelerated distributed systems.
- Experience running software releases across repositories, dependencies, test configurations, quality gates, collaborator approvals, open-source workflows, and CI/CD systems.
- Experience with tools such as GitHub, Git, Jira, Linear, Aha!, or Confluence.
Preferred Qualifications
- Experience supporting reinforcement learning, post-training, agentic AI, or large-scale model evaluation programs.
- Familiarity with PPO, GRPO, asynchronous reinforcement learning, distributed inference, rollout generation, policy optimization, evaluation harnesses, verifiers, benchmark development, or reproducible experimentation.
- Knowledge of GPU infrastructure, distributed training, Kubernetes, workload schedulers, cluster capacity management, performance analysis, open-source contributor workflows, release readiness, or operational metrics.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until August 1, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Technical Program Manager, Deep Learning Initiatives
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, DL Libraries Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Engineer, Cloud Site Reliability Engineering
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Technical Marketing Engineer, Enterprise AI Software
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Solutions Architect, IPP
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Software Engineer, RL Post-Training Frameworks
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year