Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
API @ 4
AWS
Azure
Communication @ 6
DevOps
Docker @ 6
GCP
GenAI
Generative AI @ 6
Git
Kubernetes @ 6
LLM @ 7
MLOps
Machine Learning
Prompt Engineering
Python @ 7
RAG
Reinforcement Learning @ 7
SGLang @ 4
TensorRT @ 4
Vertex AI
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius Token Factory is a serverless platform for running and customizing open-source large language models in production. It supports serverless inference and fine-tuning, including LoRA, full fine-tuning, and reinforcement fine-tuning, with optimizations such as custom speculative decoding, quantization, cache-aware routing, and dedicated endpoints.
The Senior ML Solutions Architect will support customers using Nebius Token Factory's serverless inference and fine-tuning platforms for open-source LLMs across multiple modalities. The role involves designing optimized inference workflows, building customized LLM-based solutions, architecting scalable AI applications, and collaborating with the backend team to improve the platform.
Responsibilities
- Optimize LLM inference across various modalities to drive business value and support customer goals.
- Support supervised and reinforcement learning fine-tuning to maximize model quality for customers.
- Design and implement LLM-based solutions using Nebius Token Factory inference services.
- Build production-ready applications leveraging serverless LLM APIs, including multimodal text, vision, audio, and domain-specific models.
- Provide technical expertise in prompt engineering, retrieval-augmented generation (RAG) architectures, and model selection.
- Collaborate with product and engineering teams to surface customer feedback and shape the platform roadmap.
- Guide customers in scaling from proof of concept to production, with a focus on performance, reliability, and cost efficiency.
Requirements
- 5+ years of experience in ML/AI systems, including at least 2 years focused on LLMs and generative AI.
- Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches.
- Hands-on experience running LLMs in production, including deploying and operating inference workloads.
- Experience with LLM fine-tuning, including supervised fine-tuning, SFT, LoRA, and data preparation and curation. Reinforcement-learning-based fine-tuning is a strong plus.
- Experience building LLM evaluation benchmarks and offline/online evaluation pipelines, including LLM-as-a-judge setups.
- Experience with inference frameworks and libraries such as vLLM, SGLang, TensorRT-LLM, and Transformers.
- Experience deploying LLM-powered applications using OpenAI, Anthropic, or open-source model APIs.
- Strong Python programming skills.
- Excellent communication skills and the ability to explain technical concepts to diverse audiences.
Preferred Qualifications
- Experience with multimodal AI models, including vision-language and speech models.
- Proficiency with Docker and Kubernetes.
- Contributions to open-source ML/AI projects.
Preferred Technical Stack
- Programming: Python
- ML frameworks and libraries: vLLM, TensorRT-LLM, SGLang, Transformers, OpenAI SDKs, Anthropic SDKs
- MLOps and DevOps: Kubernetes, Docker, Git
- Cloud platforms: AWS, SageMaker, Bedrock, GCP, Vertex AI, Azure, Azure ML
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet.
- Company-paid short-term disability, long-term disability, and life insurance.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.