Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API
AWS
Azure
Communication @ 6
DevOps
Docker
FastAPI
Flask
GCP
GPU
GenAI
Generative AI @ 3
Git
Kubernetes
LLM @ 3
LangChain
MLOps
Machine Learning @ 3
Networking
Prompt Engineering
PyTorch @ 3
Python @ 6
SGLang @ 3
TensorRT @ 3
Vertex AI
vLLM @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About Nebius
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel.
Summary
- Location: Remote from USA
- Duration: 3 months
- Compensation: Paid
- Eligibility: Current University student (Computer Science or related field), Recent Graduate or Early Career specialist
- Work authorization: permitted to work in the job’s location
The Role
We're looking for an ML Solutions Architect (Early Career) to join the team behind Nebius Token Factory's serverless inference and fine-tuning platform for open-source LLMs. Working alongside senior Solutions Architects, you'll take on real technical work — building and testing LLM-based solutions, benchmarking, and inference optimization — and learn how scalable AI applications are built and tuned on our platform, in close collaboration with our backend team.
This is a hands-on learning role with close mentorship from senior SAs. Strong performers will be considered for a full-time Solutions Architect position at the end of the program.
This is a paid temporary contract, open to students and recent graduates. You're welcome to work remotely from any timezone.
Responsibilities
- Help build and test LLM-based solutions and applications using Token Factory's inference services, including multimodal models (text, vision, audio).
- Assist senior SAs with prompt engineering, model selection, benchmarking, and inference optimization.
- Run performance and quality experiments to support proof-of-concept work.
- Contribute to internal tooling and automation that improves how the SA team delivers.
Requirements
- Currently pursuing or recently completed a BSc/MSc/PhD in Computer Science, Machine Learning, or a related field.
- Strong Python programming skills.
- Hands-on generative AI experience, including with common ML frameworks (e.g., PyTorch, Transformers).
- Strong communication skills, with a willingness to explain technical concepts to diverse audiences.
Nice-to-haves
- Experience deploying/serving LLMs with vLLM, SGLang, or TensorRT-LLM.
- Familiarity with inference optimization techniques such as quantization, batching, caching, and routing.
- Knowledge of model architectures and fine-tuning approaches.
- Contributions to open-source ML/AI projects.
Preferred Technical Stack
- Programming Languages: Python.
- ML Frameworks and Libraries: vLLM, SGLang, TensorRT-LLM, Transformers, OpenAI/Anthropic SDKs.
- Frameworks for Agentic Pipelines: Langchain / Langsmith / smolagents / equivalent.
- API and Web Frameworks: FastAPI, Flask.
- MLOps and DevOps tools: Kubernetes (K8s), Docker, Git.
- Cloud Platforms: AWS (SageMaker, Bedrock), GCP (Vertex AI), Azure (Azure ML).
Key employee benefits in the US
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan: Up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: Up to $85/month for mobile and internet.
- Disability & life insurance: Company-paid short-term, long-term and life insurance coverage.
Benefits & Perks
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
Pay Transparency
We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.