Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Algorithms @ 7
CUDA @ 4
Communication @ 6
GPU
LLM @ 4
Machine Learning @ 6
Networking
PyTorch @ 7
Python @ 7
SGLang @ 4
TensorRT @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About Nebius
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
Nebius Token Factory needs scientists who can turn frontier inference bottlenecks into research problems, publish credible work, and then help ship the results into production. This is not a papers-only research role. The Applied Scientist is expected to design rigorous experiments, write strong code, collaborate with engineers, and convert research into deployed inference capabilities.
A Senior Applied Scientist owns well-scoped research and production optimization projects. They can publish or prepare high-quality technical work while also producing code, experiments, and prototypes that engineers can use.
Responsibilities
- Own focused research projects from hypothesis through experiment, ablation, prototype, and production handoff.
- Prepare internal reports, technical blogs, or papers when the work is externally credible.
- Partner directly with MLEs to ensure research prototypes become usable production components.
- Define and execute research programs in efficient LLM and VLM inference with measurable production impact.
- Invent, evaluate, and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse, KV-cache compression, long-context inference, MoE routing, and model/runtime co-optimization.
- Build high-quality prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks, then work with MLEs and platform engineers to productionize them.
- Design rigorous evaluation methodology covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.
- Publish papers, technical reports, blog posts, and open-source artifacts that build external credibility for Nebius Token Factory.
- Collaborate with MLE, GPU kernel, backend infrastructure, product, and customer teams to choose high-leverage research bets.
- Mentor engineers and scientists on experimental design, scientific rigor, and model/system tradeoffs.
Requirements
- PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field.
- Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas.
- Strong hands-on coding ability in Python and PyTorch; ability to move from idea to experiment to prototype quickly.
- Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs.
- Strong experimental design skills, including ablations, baselines, metrics, statistical reasoning, and failure analysis.
- Excellent written and verbal communication.
Nice-to-have
- First-author publications in NeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, HPCA, or comparable venues.
- Experience deploying ML models or inference optimizations in production.
- Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals.
- Experience with post-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation when connected to inference quality or efficiency.
- Open-source research artifacts, widely used benchmarks, high-quality technical blogs, or invited talks in efficient AI systems.
Benefits
Key employee benefits in the US:
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan: Up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: Up to $85/month for mobile and internet.
- Disability & life insurance: Company-paid short-term, long-term and life insurance coverage.
Benefits & Perks:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
Pay Transparency
We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.
Base Compensation Range: $195,200 — $262,200 USD