Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
CUDA @ 1
Communication @ 2
Distributed Systems @ 3
GPU @ 3
HPC @ 3
JAX @ 3
MPI @ 2
NCCL @ 2
Observability
Performance Analysis
Profiling @ 3
PyTorch @ 3
Python @ 1
Rust @ 1
TensorFlow @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Training Runtime designs the core distributed machine-learning training runtime that powers early research experiments through frontier-scale model runs. The team is building a unified, modular runtime focused on high-performance data movement, fault-tolerant training frameworks, deterministic orchestration, observability, and distributed process management.
The role drives efficiency improvements across the distributed training stack by analyzing large-scale training runs, identifying utilization gaps, and designing optimizations that improve throughput and uptime. This includes GPU kernel performance analysis, collective communication throughput, I/O bottleneck investigation, and model sharding for massive-scale training. The role is based in San Francisco, California, with a hybrid work model requiring three days in the office per week.
Responsibilities
- Profile end-to-end training runs to identify performance bottlenecks across compute, communication, and storage.
- Optimize GPU utilization and throughput for large-scale distributed model training.
- Collaborate with runtime and systems engineers to improve kernel efficiency, scheduling, and collective communication performance.
- Implement model graph transformations to improve end-to-end throughput.
- Build tooling to monitor and visualize MFU, throughput, and uptime across clusters.
- Partner with researchers to ensure new model architectures scale efficiently during pre-training.
- Contribute to infrastructure decisions that improve the reliability and efficiency of large training jobs.
Requirements
- Strong programming skills in Python and C++; Rust or CUDA experience is a plus.
- Experience running distributed training jobs on multi-GPU systems or HPC clusters.
- Ability to debug complex distributed systems and measure efficiency rigorously.
- Exposure to frameworks such as PyTorch, JAX, or TensorFlow.
- Understanding of how large-scale training loops are built.
- Ability to collaborate across teams and translate profiling data into practical engineering improvements.
Nice to Have
- Familiarity with NCCL, MPI, or UCX communication libraries.
- Experience with large-scale data loading and checkpointing systems.
- Prior work on training runtimes, distributed scheduling, or machine-learning compiler optimization.
Benefits
- Base salary range of $250,000–$445,000 per year, plus equity.
- Medical, dental, and vision insurance, with employer contributions to Health Savings Accounts.
- Pre-tax accounts for health, dependent care, and commuter expenses.
- 401(k) retirement plan with employer match.
- Paid parental, medical, and caregiver leave.
- Paid time off, paid company holidays, office closures, and sick or safe time.
- Mental health and wellness support.
- Employer-paid basic life and disability coverage.
- Annual learning and development stipend.
- Daily office meals and eligible meal delivery credits.
- Relocation support for eligible employees.
- Additional benefits may include charitable donation matching and wellness stipends.