ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 3
Communication @ 4
Debugging @ 4
Distributed Systems @ 4
GPU @ 4
InfiniBand @ 3
Kubernetes @ 4
Leadership @ 7
Machine Learning
NCCL @ 3
Networking
PyTorch @ 7
Python @ 7
R
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About Nebius
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel.
The role
Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient.
The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering.
A Senior ML Systems Engineer owns substantial training or RL infrastructure components end to end. They are deeply hands-on, can debug difficult distributed training failures independently, and can deliver measurable improvements in experiment throughput, stability, and GPU utilization.
Your responsibilities
- Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads.
- Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems.
- Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism.
- Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training.
- Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput.
- Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers.
- Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users.
- Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems.
- Write clear design docs, incident reports, benchmark reports, and operating guides.
Must-haves
- Strong Python and PyTorch engineering skills.
- Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads.
- Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing.
- Experience debugging production or research training jobs across multiple GPUs or nodes.
- Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity.
- Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership.
Nice-to-have
- Experience with Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms.
- Experience with RL infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO/GRPO/RLHF systems.
- Familiarity with NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100/H200/B200 clusters, or storage/network bottlenecks.
- Experience supporting SFT, DPO, PPO, GRPO, RLAIF, reward model serving, rollout generation, or agent training workloads.
- Open-source contributions to distributed training, RL infrastructure, PyTorch, Ray, Megatron, DeepSpeed, or related systems.
Key employee benefits in the US
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan: Up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: Up to $85/month for mobile and internet.
- Disability & life insurance: Company-paid short-term, long-term and life insurance coverage.
Pay Transparency
We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.
Base Compensation Range: $195,200 — $262,200 USD
Benefits & Perks
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams