Forward Deployed Engineer - Physical AI Cloud Platform

at Nebius
USD 179,500-224,300 per year
SENIOR
✅ Remote

Tech Stack

AI @ 6 API AWS @ 4 Airflow @ 4 Azure @ 4 CI/CD @ 4 CUDA @ 3 Claude Code @ 6 Codex @ 6 Communication @ 7 Data Pipelines @ 4 Distributed Systems @ 4 GCP @ 4 GPU @ 3 Go @ 7 HPC @ 3 IaC Kubernetes @ 4 Machine Learning Networking @ 4 Observability @ 4 Python @ 7 SRE @ 7 Security Slurm @ 4

Details

Nebius is building a full-stack AI cloud platform supporting developers and enterprises from data and model training through production deployment. The Forward Deployed Engineer, Cloud Platform is a senior, high-autonomy individual contributor responsible for infrastructure that makes the physical AI platform fast, reliable, scalable, secure, and cost-effective. The role is embedded with strategic customers and ISV partners and owns end-to-end technical execution from discovery and infrastructure design through production rollout.

The position is remote within the United States, with the San Francisco Bay Area, California, or Austin, Texas preferred.

Responsibilities

  • Own discovery, technical scoping, infrastructure design, implementation, and production rollout for design partners and ISV engagements.
  • Build and operate cloud infrastructure for simulation, training, evaluation, inference, and batch workloads.
  • Build platform services for job execution, scheduling, retries, observability, logging, secrets, access control, and cost tracking.
  • Integrate Nebius cloud services into product experiences to abstract infrastructure complexity from customers.
  • Build onboarding infrastructure for pilots, including sandbox environments, dataset storage, workflow execution, and deployment.
  • Optimize cloud cost, utilization, performance, security, and reliability across workloads.
  • Debug infrastructure issues across application, network, storage, compute, and orchestration layers.
  • Partner with Physical AI Systems and Platform & Product FDEs to support GPU-heavy workloads and expose infrastructure capabilities through APIs, SDKs, and product workflows.
  • Help define infrastructure architecture for multi-tenant SaaS, enterprise deployments, and high-throughput physical AI workloads.
  • Turn repeated customer infrastructure challenges into reusable platform capabilities and incorporate them into the core platform.
  • Use AI coding tools such as Claude Code, Codex, and Cursor to accelerate production software development.
  • Co-author reference architectures, solution templates, and technical blogs, and maintain feedback loops with Field CTO, Product, and Engineering teams.

Requirements

  • 6+ years of hands-on engineering experience in backend engineering, cloud infrastructure, platform engineering, or SRE.
  • At least 2 years in a customer-facing or deployment-oriented technical role, such as Forward Deployed Engineer, founding engineer, technical co-founder, or technical lead embedded with strategic customers.
  • Experience building distributed systems, job orchestration, compute platforms, internal developer platforms, or ML infrastructure.
  • Strong Python, Go, or similar systems and backend programming skills.
  • Fluency with AI coding tools including Claude Code, Codex, and Cursor.
  • Experience with Kubernetes, containers, CI/CD, observability, cloud networking, storage, IAM/RBAC, and infrastructure as code.
  • Familiarity with GPU workloads, batch jobs, training pipelines, inference workloads, or HPC-style compute environments.
  • Proven ability to debug infrastructure across application, network, storage, compute, and orchestration layers.
  • Strong instincts for workload isolation, RBAC, uptime, and traceability.
  • Ability to navigate ambiguity with a bias toward simple, composable infrastructure serving real customer workflows.
  • Strong written and verbal communication skills, including communication with customer CTOs and internal technical leaders.

Additional Qualifications

  • Prior experience as a Forward Deployed Engineer or in an equivalent customer-embedded engineering function.
  • Experience with Nebius, AWS, GCP, Azure, Lambda Labs, or other AI cloud infrastructure.
  • Experience with Slurm, Soperator, Kubernetes GPU scheduling, Ray, Argo, Airflow, Metaflow, or similar orchestration tools.
  • Experience with ML training infrastructure, model serving, simulation workloads, or large-scale data pipelines.
  • Experience supporting enterprise customers, design partners, or production pilots.
  • Familiarity with NVIDIA GPU infrastructure, CUDA workloads, Isaac Sim, Omniverse, or simulation at scale.

Benefits

  • 100% company-paid medical, dental, and vision coverage for employees and families.
  • 401(k) plan with up to 4% company match and immediate vesting.
  • 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
  • Remote work reimbursement of up to $85 per month for mobile and internet.
  • Company-paid short-term disability, long-term disability, and life insurance.
  • Career growth and learning opportunities.
  • Flexibility and ownership.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment and talented teams.

Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.

More jobs at Nebius

Similar jobs