AI/ML Specialist Solutions Architect

at Nebius
📍 Canada
📍 United States
USD 250,000-320,000 per year
MIDDLE
✅ Remote
Featured

Tech Stack

AI @ 3 Ansible Communication @ 3 DevOps Docker GPU @ 3 Git Go Hadoop Helm Hiring @ 3 IaC JAX @ 3 Java Kafka Kubernetes MLOps @ 5 Machine Learning @ 3 NoSQL PyTorch @ 3 Python SQL Slurm Spark TensorFlow Terraform Vector Databases scikit-learn

Details

Nebius is building a full-stack AI cloud platform for developers and enterprises, supporting workloads from data and model training through production deployment. The Customer Experience Team solves real-world AI and machine learning challenges at massive GPU cloud scale, working with GPUs such as H200, B200, and GB200 and modern machine learning frameworks.

This role supports AI-focused customers using Nebius services. You will act as a trusted advisor, collaborate with clients to design scalable AI solutions, resolve technical challenges, and manage large-scale AI deployments involving hundreds to thousands of GPUs. The position can be performed remotely from the United States or Canada.

Responsibilities

  • Design customer-centric solutions that maximize business value and align with strategic goals.
  • Build and maintain long-term relationships to foster trust and ensure customer satisfaction.
  • Deliver technical presentations, produce whitepapers, create manuals, and host webinars for audiences with varying technical expertise.
  • Collaborate with engineering and product teams to prioritize and relay customer feedback.

Requirements

  • 3+ years of experience with cloud technologies in MLOps engineering, machine learning engineering, or similar roles.
  • Strong understanding of machine learning ecosystems, including models, use cases, and tooling.
  • Proven experience setting up and optimizing distributed training pipelines across multi-node and multi-GPU environments.
  • Expertise deploying inference infrastructure for production workloads.
  • Ability to transition machine learning pipelines from proof of concept to scalable production systems.
  • Hands-on knowledge of PyTorch or JAX.
  • Excellent verbal and written communication skills.

Preferred Tooling

  • Programming languages: Python, Go, Java, C++
  • Infrastructure as Code: Terraform, Ansible
  • Orchestration: Kubernetes, Slurm
  • DevOps tools: Git, Docker, Helm
  • Big data frameworks: Spark, Kafka, Hadoop
  • Databases: SQL, NoSQL, including vector databases
  • Machine learning frameworks and libraries: PyTorch, TensorFlow, JAX, HuggingFace, Scikit-learn

Compensation

The on-target earnings range is $250,000–$320,000 USD. Actual compensation depends on experience, skills, qualifications, hiring level, and geographic location.

Benefits

  • 100% company-paid medical, dental, and vision coverage for employees and families.
  • 401(k) plan with up to a 4% company match and immediate vesting.
  • 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
  • Remote work reimbursement of up to $85 per month for mobile and internet.
  • Company-paid short-term disability, long-term disability, and life insurance.
  • Career growth and learning opportunities.
  • Flexibility and ownership.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment and talented teams.

Nebius is an equal opportunity employer committed to an inclusive and diverse workplace. Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire.

More jobs at Nebius

Similar jobs