Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
API @ 4
AWS
Azure
Communication @ 6
Debugging @ 4
DevOps
Docker @ 6
GCP
GenAI
Generative AI @ 7
Git
Hiring @ 4
Kubernetes @ 6
LLM @ 4
Leadership @ 6
MLOps
Machine Learning
Mentoring @ 6
Prompt Engineering
Python @ 7
RAG
SGLang @ 4
Technical Leadership @ 6
TensorRT @ 4
Vertex AI
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform for developers and enterprises, supporting workloads from data and model training through production deployment. This role is part of Nebius Token Factory, a serverless platform for running and customizing open-source large language models in production through serverless inference and fine-tuning, including LoRA, full fine-tuning, and reinforcement fine-tuning. The platform includes custom speculative decoding, quantization, cache-aware routing, and dedicated endpoints.
The Principal ML Solutions Architect will serve as the most senior technical authority for customers using Token Factory's serverless inference and fine-tuning platforms. The role involves designing and implementing optimized inference and fine-tuning workflows, setting technical direction for strategic accounts, solving complex performance and quality problems, mentoring Solutions Architects, and influencing the platform roadmap in partnership with backend, product, and research teams.
The position can be performed remotely from the United States.
Responsibilities
- Own complex, high-stakes customer engagements from architecture through production across multiple modalities, driving measurable business value.
- Optimize LLM inference at the framework and hardware levels and codify best practices into reusable team playbooks.
- Lead supervised and reinforcement fine-tuning efforts to maximize model quality.
- Design and implement production-ready LLM solutions using Token Factory inference services.
- Provide technical expertise in prompt engineering, RAG architectures, model selection, and cost/performance trade-offs at scale.
- Partner with product, engineering, and research teams to identify customer needs, prototype platform features, and influence the roadmap.
- Guide customers from proof of concept to production, focusing on performance, reliability, and cost efficiency.
- Mentor Senior and mid-level Solutions Architects through technical reviews, enablement, and knowledge sharing.
- Represent Token Factory externally through talks, blog posts, and conferences.
Requirements
- 8+ years of experience in ML/AI systems, including at least 4 years focused on LLMs and generative AI.
- Demonstrated technical leadership, including ownership of ambiguous, high-impact problems and the ability to influence decisions across teams and customers.
- Expert knowledge of the LLM ecosystem, including model architectures, fine-tuning approaches, and inference internals.
- Deep hands-on experience with inference optimization, including quantization, KV-cache management, batching, and routing.
- Experience running LLMs in production at scale, including deploying, operating, and debugging inference workloads at the framework level.
- Experience with LLM fine-tuning, including SFT, LoRA, data preparation and curation, and RL-based fine-tuning.
- Experience building LLM evaluation systems, including task-specific benchmarks, offline and online evaluation pipelines, and LLM-as-a-judge setups.
- Experience with inference frameworks and libraries such as vLLM, SGLang, and TensorRT-LLM, including the ability to read, modify, and contribute to their internals.
- Experience deploying LLM-powered applications using OpenAI, Anthropic, or open-source model APIs.
- Strong Python programming skills.
- Excellent communication skills, with the ability to explain technical concepts to engineers, executives, and other audiences.
Preferred Qualifications
- Contributions to or maintainership of major open-source inference or ML projects, including vLLM, SGLang, or TensorRT-LLM.
- Published research, conference talks, or widely read technical writing in LLM or model-serving topics.
- Experience with multimodal AI models, including vision-language and speech models.
- Proficiency with Docker, Kubernetes, and infrastructure-as-code.
- Experience building or owning internal tooling and automation for ML workflows at scale.
Technical Stack
- Programming: Python
- ML frameworks and libraries: vLLM, TensorRT-LLM, SGLang, Transformers, OpenAI SDKs, Anthropic SDKs
- MLOps and DevOps: Kubernetes, Docker, Git
- Cloud platforms: AWS, SageMaker, Bedrock, GCP, Vertex AI, Azure, Azure ML
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet.
- Company-paid short-term disability, long-term disability, and life insurance.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
Compensation
The base compensation range is $208,000–$261,000 USD. Actual compensation depends on job-related factors, including experience, skills, qualifications, hiring level, and geographic location.
Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.