Senior Manager, Software Engineering - RL Post-Training Frameworks
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 8
Algorithms @ 4
CUDA @ 3
Distributed Systems @ 8
GPU
HPC
Hiring @ 4
Kubernetes @ 4
LLM @ 3
Machine Learning
NCCL @ 3
Networking
Reinforcement Learning
SGLang
Slurm @ 4
TensorRT @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Lead globally distributed teams and the systems they build into a production-quality reinforcement learning ecosystem for researchers and model builders. Reinforcement learning post-training enables modern AI systems to reason, use tools, follow detailed instructions, and act as agents. At scale, RL creates demanding systems challenges spanning inference, rollout, reward and critic evaluation, training, GPUs, CPUs, networking, storage, and open-source runtimes.
NVIDIA is building an RL Frameworks engineering team responsible for the open-source tools and infrastructure used by researchers, model builders, and external partners. The role covers RL frameworks such as VeRL, Miles, Slime, SkyRL, TorchTitan, and related post-training stacks, along with the systems they build on and integrate with, including Megatron-Core, Ray, Monarch, NIXL, SGLang, Kubernetes, and NVIDIA platform libraries.
Responsibilities
- Own NVIDIA's RL post-training frameworks strategy, including direct investment, upstream partnerships, and prioritization based on customer impact, ecosystem leverage, technical feasibility, and opportunity cost.
- Evaluate architecture and performance across training, inference, rollout, orchestration, and the NVIDIA platform.
- Drive integrations that improve RL framework quality and user value, and convert decisions into measurable execution plans.
- Establish benchmarking and reproducibility criteria across open-source frameworks and distributed runtimes.
- Partner closely with product management, research, Developer Relations, customer-facing teams, hardware, CUDA, networking, math libraries, compilers, and external open-source collaborators.
- Recruit and develop managers and senior individual contributors.
- Create an effective US/APAC operating model, review capacity against commitments, and establish clear ownership and decision rights.
- Coach engineers to contribute effectively in open-source ecosystems and deliver high-quality upstream work.
- Turn technical and partner questions into concrete, measurable actions, set delivery goals, and maintain engineering quality standards.
- Build durable open-source improvements that enable RL workloads to run effectively on NVIDIA systems.
Requirements
- Master's or PhD degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 10+ years of software engineering experience in distributed systems, AI frameworks, ML infrastructure, high-performance computing, or systems software.
- 4+ years of experience as an engineering manager for software teams.
- Strong technical background in distributed AI systems, including training, inference, orchestration, and end-to-end performance.
- Experience defining domain-level technical strategy, making build-versus-buy or upstream-versus-internal investment decisions, and creating multi-team execution plans.
- Ability to drive engineering work across organizational boundaries, influence without direct authority, and communicate tradeoffs clearly to senior leaders and executives.
- Experience hiring and leading engineering teams, developing technical leaders or new managers, and creating staffing plans for evolving technical domains.
- Experience establishing workflows, success criteria, metrics, or decision gates that improve engineering execution across teams.
- Background collaborating with open-source communities, research teams, external partners, or customer-facing teams.
Preferred Qualifications
- Hands-on experience with RL post-training frameworks or algorithms such as RLHF, PPO, GRPO, DPO, reward modeling, VeRL, Miles, Slime, SkyRL, OpenRLHF, NeMo-Aligner, or TorchTitan.
- Experience with runtime and orchestration systems such as Ray, Monarch, Kubernetes, Slurm, or comparable actor- and task-based systems.
- Experience scaling workloads across thousands of GPUs or heterogeneous systems, including fault tolerance, elastic recovery, stragglers, resource contention, or benchmark reproducibility.
- Familiarity with NVIDIA platform components such as CUDA, NCCL, cuDNN, TensorRT-LLM, Transformer Engine, Nsight, NeMo, or Megatron-Core.
- Ability to turn customer or partner needs into reusable upstream improvements rather than one-off support.
Compensation and Benefits
- Base salary: USD 272,000–431,250 per year.
- Eligible for equity and benefits.
- Applications accepted at least until August 29, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.