Senior Applied Scientist, Efficient LLM Inference & Model Optimization
at Nebius
USD 195,200-262,200 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Algorithms @ 7
CUDA @ 4
Communication @ 6
Experimentation @ 7
GPU
LLM @ 4
Machine Learning @ 4
Mathematics @ 6
PyTorch @ 7
Python @ 7
SGLang @ 4
TensorRT @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform supporting developers and enterprises from data and model training through production deployment. The Nebius Token Factory team is looking for an Applied Scientist to turn frontier inference bottlenecks into research problems, publish credible work, and help ship results into production. This role combines rigorous research, strong coding, experimentation, prototyping, and production optimization.
Responsibilities
- Own focused research projects from hypothesis through experimentation, ablation, prototyping, and production handoff.
- Prepare internal reports, technical blogs, or papers when the work is externally credible.
- Partner with machine learning engineers to turn research prototypes into production components.
- Define and execute research programs in efficient LLM and VLM inference with measurable production impact.
- Invent, evaluate, and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse and compression, long-context inference, MoE routing, and model/runtime co-optimization.
- Build prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks, then work with machine learning and platform engineers to productionize them.
- Design evaluation methodologies covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.
- Publish papers, technical reports, blog posts, and open-source artifacts.
- Collaborate with machine learning, GPU kernel, backend infrastructure, product, and customer teams to select high-leverage research opportunities.
- Mentor engineers and scientists on experimental design, scientific rigor, and model/system trade-offs.
Requirements
- PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied mathematics, or a closely related field.
- Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas.
- Strong hands-on coding ability in Python and PyTorch, with the ability to move quickly from ideas to experiments and prototypes.
- Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving trade-offs.
- Strong experimental design skills, including ablations, baselines, metrics, statistical reasoning, and failure analysis.
- Excellent written and verbal communication.
Nice-to-Have Qualifications
- First-author publications in NeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, HPCA, or comparable venues.
- Experience deploying machine learning models or inference optimizations in production.
- Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals.
- Experience with post-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation connected to inference quality or efficiency.
- Open-source research artifacts, widely used benchmarks, high-quality technical blogs, or invited talks in efficient AI systems.
Benefits
- Base compensation range of $195,200–$262,200 USD.
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet.
- Company-paid short-term disability, long-term disability, and life insurance coverage.
- Career growth and learning opportunities, flexibility and ownership, a collaborative and innovative culture, and the opportunity to work on impactful AI projects.
Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer and provides accommodations during the application process when needed.
More jobs at Nebius
Data Center GM
Nebius · United States
USD 200,000-250,000 per year
Field Network Technician
Nebius · United States
USD 83,200-99,800 per year
Senior Technical Project Manager – Applied AI
Nebius · Palo Alto, United States
USD 147,200-224,000 per year
Director, Forward Deployed Engineering
Nebius · United States
USD 270,800-310,000 per year
Forward Deployment Engineering Manager
Nebius · United States
USD 225,800-281,000 per year
Similar jobs
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Inference Performance Engineer, AI Inference Configuration Optimization
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Machine Learning Engineer, LLM Inference Optimization
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior Deep Learning Frameworks CUDA Software Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Machine Learning Inference
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year