Senior Software Engineer, RL Post-Training Frameworks

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Communication @ 7 Deep Learning @ 6 Distributed Systems @ 6 Experimentation GPU HPC InfiniBand Kubernetes LLM @ 4 Machine Learning NCCL NVLink Networking PyTorch Python @ 7 Reinforcement Learning SGLang TensorRT vLLM

Details

Reinforcement learning post-training teaches models to reason through difficult problems, follow complex instructions, and act as autonomous agents. NVIDIA is building an RL Frameworks engineering team to develop the open-source tools and infrastructure used by AI researchers and post-training teams.

The team works across the software stack, including RL frameworks such as VeRL, Miles, and TorchTitan, as well as distributed runtimes including Ray and Monarch. The work involves collaborating with researchers and AI labs, optimizing deep learning frameworks, and building distributed infrastructure for next-generation AI systems.

Responsibilities

  • Architect and build RL post-training infrastructure that scales from experimentation on a single GPU to production deployments across thousands of nodes.
  • Tune RL training, inference, and rollout loops on GPUs, CPUs, and LPUs for performance.
  • Contribute to and improve the performance and usability of open-source RL frameworks, and partner with the teams that maintain them.
  • Build systems supporting fault tolerance, elastic scaling, and fast restarts so long-running distributed training jobs can survive failures, stragglers, and resource contention.
  • Support CPU-driven rollout workloads, including tool use, code execution, and agentic environments, alongside GPU- or LPU-accelerated generation and GPU-accelerated training.
  • Advocate for researcher and partner requirements with NVIDIA networking, math library, and compiler teams.
  • Work with hardware teams to take advantage of next-generation hardware capabilities in post-training workloads.

Requirements

  • MS or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 5+ years of professional experience in distributed systems, high-performance computing, deep learning infrastructure, or ML systems engineering.
  • Strong proficiency in Python and C/C++.
  • Demonstrated experience building or contributing to large-scale distributed systems or runtime frameworks in production at a frontier AI lab, hyperscaler, or major technology company.
  • Strong verbal and written communication skills, with the ability to collaborate across organizational and geographic boundaries.

Preferred Technical Depth

  • Reinforcement learning for LLM post-training, including RLHF, PPO, GRPO, DPO, and reward modeling, as well as the distributed execution challenges they create.
  • PyTorch internals, including FSDP, tensor parallelism, pipeline parallelism, and their composition.
  • Kubernetes runtime internals, including container lifecycle, pod scheduling, resource quotas, and GPU allocation.
  • End-to-end distributed systems design, including service boundaries, data flows, consistency models, failure modes, and recovery approaches.

Additional Experience

  • Networking technologies such as NCCL, NVLink, and InfiniBand; advanced multidimensional parallelism such as Megatron-LM, FSDP2, TP/DP/PP, and MoE; or memory optimizations such as quantization-aware training and mixed precision.
  • Integration of high-performance inference engines such as vLLM, SGLang, or TensorRT-LLM into RL training loops for GPU-accelerated rollout.
  • Actor- and task-based distributed programming with Ray, Monarch, or comparable systems.
  • Multi-turn training, multi-agent co-evolution, or VLM post-training.
  • Open-source contributions to RL post-training or distributed training projects such as VeRL, Miles, TorchTitan, OpenRLHF, NeMo-Aligner, or DeepSpeed-Chat.
  • Kubernetes work beyond routine operations, including custom operators, GPU device plugins, or scheduling contributions.
  • Direct experience operating frontier-scale training, including RL post-training at thousands of GPUs or large-scale LLM or multimodal pre-training.
  • Hands-on experience with production distributed failures at scale, including stragglers, resource contention, and hardware faults.

Benefits

NVIDIA offers competitive salaries, a comprehensive benefits package, equity, and benefits for employees and their families. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Salary is determined based on location, experience, and compensation for employees in similar positions.

NVIDIA is an equal opportunity employer committed to fostering a diverse work environment. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs