Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
AWS
Azure
Communication @ 3
Debugging @ 3
DevOps @ 5
Docker @ 5
GCP
Git
Kubernetes @ 5
LLM @ 6
MLOps
Machine Learning
Python @ 5
Reporting @ 3
SGLang @ 3
TensorRT @ 3
Vertex AI
vLLM @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform for developers and enterprises, supporting data and model training through production deployment. This role sits within Nebius Token Factory, a serverless platform for running and customizing open-source large language models in production. Token Factory provides serverless inference and fine-tuning with optimizations including custom speculative decoding, quantization, cache-aware routing, and dedicated endpoints.
The Manager, ML Solutions Architecture will lead Nebius's US regional Solutions Architecture teams. The team owns the technical delivery of customer engagements, including deploying open-source models, tuning serving stacks, benchmarking against customer success criteria, and maintaining the technical relationship through production. The role focuses on proof-of-concept delivery and post-sales technical support and reports to the Head of Solutions Architecture.
Responsibilities
Lead the Team
- Manage a team of four Solutions Architects, with continued growth planned.
- Conduct 1:1s, set goals, complete performance reviews, prepare promotion cases, and create individual growth plans.
- Assess each Solutions Architect's strengths, gaps, and preferences and allocate accounts and engagements accordingly.
- Establish a cadence for surfacing blockers and coordinate with development, product, and business teams to resolve them.
- Hold team members accountable for outcomes and reflect performance in ratings and compensation decisions.
- Onboard new team members through their first independently delivered engagement.
- Coach Solutions Architects to become stronger engineers and communicators and determine when they are ready for additional scope.
Own Delivery
- Own delivery outcomes, including time from proof-of-concept kickoff to the first optimized dedicated endpoint, success-criteria hit rate, and the quality of the technical relationship after production launch.
- Review benchmarking methodology, serving configurations, results, and closure documents before they reach customers.
- Ensure staffing and escalation coverage across accounts and time zones, including post-sales requests outside sprint boundaries.
- Identify infeasible requirements early and support decisions with evidence.
Own Team Operations
- Maintain and extend documentation covering responsibilities, runbooks, guides, onboarding, definitions of done, and engagement closure templates.
- Ensure documentation is accurate, accessible, and used by the team.
- Keep the ticket tracker as the system of record for proof-of-concept and production status.
- Define and report metrics showing whether delivery is becoming faster and more reliable.
Collaborate Across Teams
- Work with development and research teams to convert recurring customer problems into prioritized platform work and represent customer technical needs in roadmap discussions.
- Partner with pre-sales to maintain the boundary between scoping and execution, challenge under-scoped engagements, and provide feasibility feedback.
- Coordinate with account management to support production handoffs and keep post-sales technical requests moving.
- Provide business and leadership teams with clear assessments of account health, technical feasibility, and capacity needs.
Requirements
- At least three years of experience managing technical teams, including performance management and difficult conversations.
- Experience managing a customer-facing team operating on customer timelines and handling customer escalations.
- Strong ML knowledge, including LLM architectures, fine-tuning approaches such as SFT, LoRA, and reinforcement-learning-based methods, evaluation design, and inference internals.
- Understanding of quantization, KV-cache management, batching, routing, and speculative decoding.
- Experience with or understanding of vLLM, SGLang, and TensorRT-LLM.
- Sufficient Python proficiency to read and review the team's code.
- Technical judgment to review benchmarks and identify flaws in methodology.
- Excellent communication skills and the ability to explain technical concepts to engineers, executives, and enterprise customers under pressure.
- Willingness to handle operational work including documentation, process design, reporting, and follow-through.
- Comfort working with ambiguity across distributed teams and time zones, with a bias toward documenting decisions and processes.
Preferred Qualifications
- References from former managers and direct reports.
- Experience scaling a team through rapid growth while maintaining delivery quality.
- Customer-facing technical experience at a cloud, inference, or AI infrastructure provider.
- Experience defining processes and documentation for a team starting without them.
- Hands-on experience running LLMs in production and debugging inference workloads at the framework level.
- Experience with multimodal AI models, including vision-language and speech models.
- Proficiency with DevOps tooling such as Docker and Kubernetes and with infrastructure-as-code.
Preferred Technical Stack
- Programming: Python
- ML frameworks and libraries: vLLM, TensorRT-LLM, SGLang, Transformers, OpenAI SDKs, Anthropic SDKs
- MLOps and DevOps: Kubernetes, Docker, Git
- Cloud platforms: AWS, SageMaker, Bedrock, GCP, Vertex AI, Azure, Azure ML
Benefits
- Competitive compensation and benefits.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
The position is remote within the United States. Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility. Nebius is an equal opportunity employer and provides accommodations during the application process upon request.