Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Distributed Systems
GPU @ 3
LLM
Machine Learning
Networking
React
SRE
TensorRT @ 3
vLLM @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About Nebius
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The Role
We’re looking for a Technical Account Manager (TAM) who will join our Token Factory team to help our customers successfully transition from proof-of-concept to production and scale their AI workloads on Nebius infrastructure.
This role sits at the intersection of engineering, delivery, and customer success — ensuring that what was promised during pre-sales actually works reliably in production. You will work closely with customer engineering teams, Solution Architects, and Product/Infrastructure teams to drive stable, performant, and cost-efficient deployments.
This role is NOT:
- A sales role (though you will support expansion through value)
- A pure support role (you won’t just react to tickets)
- A solution architect role (you won’t design systems from scratch)
You’re welcome to work remotely from the United States.
Responsibilities
- Own the production journey
- Lead the transition from PoC to production
- Ensure customer workloads are deployed, stable, and scalable
- Drive time-to-production and time-to-value
- Ensure technical success in production
- Understand customer architectures and use cases
- Monitor and improve:
- performance (latency, throughput)
- cost efficiency
- reliability
- Identify and resolve bottlenecks proactively
- Act as a trusted technical partner
- Work directly with customer engineering teams
- Provide guidance on best practices and optimisation
- Translate technical challenges into actionable solutions
- Manage risks and incidents
- Act as a primary technical contact for production issues
- Coordinate with internal teams to resolve incidents
- Communicate clearly during high-pressure situations
- Drive continuous improvement
- Identify opportunities to optimise and expand usage
- Provide structured feedback to Product and Infrastructure teams
- Help shape better solutions based on real customer needs
Requirements
Technical background
- Practical knowledge of inference frameworks (e.g. vLLM, TensorRT, or similar)
- Solid understanding of:
- cloud or infrastructure systems
- distributed systems or high-load applications
- AI/ML workloads (LLMs, inference, etc.) is a strong plus
- Ability to troubleshoot and reason about system performance
Customer-facing experience
- Experience working directly with technical customers (e.g. engineers, ML teams)
- Ability to communicate complex topics clearly and effectively
Ownership & execution
- Strong sense of ownership — you drive outcomes, not just tasks
- Ability to manage multiple customers and priorities
- Structured, proactive, and solution-oriented mindset
It will be an added bonus if you have
- Experience with GPU workloads or AI infrastructure
- Background in solutions engineering, SRE, or technical support in B2B environments
Key employee benefits
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan: Up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: Up to $85/month for mobile and internet.
- Disability & life insurance: Company-paid short-term, long-term and life insurance coverage.
Benefits & Perks
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams