Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
API
AWS @ 4
Airflow @ 4
Azure @ 4
CI/CD @ 4
CUDA @ 3
Claude Code @ 6
Codex @ 6
Communication @ 7
Data Pipelines @ 4
Distributed Systems @ 4
GCP @ 4
GPU @ 3
Go @ 7
HPC @ 3
IaC
Kubernetes @ 4
Machine Learning
Networking @ 4
Observability @ 4
Python @ 7
SRE @ 7
Security
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform supporting developers and enterprises from data and model training through production deployment. The Forward Deployed Engineer, Cloud Platform is a senior, high-autonomy individual contributor responsible for infrastructure that makes the physical AI platform fast, reliable, scalable, secure, and cost-effective. The role is embedded with strategic customers and ISV partners and owns end-to-end technical execution from discovery and infrastructure design through production rollout.
The position is remote within the United States, with the San Francisco Bay Area, California, or Austin, Texas preferred.
Responsibilities
- Own discovery, technical scoping, infrastructure design, implementation, and production rollout for design partners and ISV engagements.
- Build and operate cloud infrastructure for simulation, training, evaluation, inference, and batch workloads.
- Build platform services for job execution, scheduling, retries, observability, logging, secrets, access control, and cost tracking.
- Integrate Nebius cloud services into product experiences to abstract infrastructure complexity from customers.
- Build onboarding infrastructure for pilots, including sandbox environments, dataset storage, workflow execution, and deployment.
- Optimize cloud cost, utilization, performance, security, and reliability across workloads.
- Debug infrastructure issues across application, network, storage, compute, and orchestration layers.
- Partner with Physical AI Systems and Platform & Product FDEs to support GPU-heavy workloads and expose infrastructure capabilities through APIs, SDKs, and product workflows.
- Help define infrastructure architecture for multi-tenant SaaS, enterprise deployments, and high-throughput physical AI workloads.
- Turn repeated customer infrastructure challenges into reusable platform capabilities and incorporate them into the core platform.
- Use AI coding tools such as Claude Code, Codex, and Cursor to accelerate production software development.
- Co-author reference architectures, solution templates, and technical blogs, and maintain feedback loops with Field CTO, Product, and Engineering teams.
Requirements
- 6+ years of hands-on engineering experience in backend engineering, cloud infrastructure, platform engineering, or SRE.
- At least 2 years in a customer-facing or deployment-oriented technical role, such as Forward Deployed Engineer, founding engineer, technical co-founder, or technical lead embedded with strategic customers.
- Experience building distributed systems, job orchestration, compute platforms, internal developer platforms, or ML infrastructure.
- Strong Python, Go, or similar systems and backend programming skills.
- Fluency with AI coding tools including Claude Code, Codex, and Cursor.
- Experience with Kubernetes, containers, CI/CD, observability, cloud networking, storage, IAM/RBAC, and infrastructure as code.
- Familiarity with GPU workloads, batch jobs, training pipelines, inference workloads, or HPC-style compute environments.
- Proven ability to debug infrastructure across application, network, storage, compute, and orchestration layers.
- Strong instincts for workload isolation, RBAC, uptime, and traceability.
- Ability to navigate ambiguity with a bias toward simple, composable infrastructure serving real customer workflows.
- Strong written and verbal communication skills, including communication with customer CTOs and internal technical leaders.
Additional Qualifications
- Prior experience as a Forward Deployed Engineer or in an equivalent customer-embedded engineering function.
- Experience with Nebius, AWS, GCP, Azure, Lambda Labs, or other AI cloud infrastructure.
- Experience with Slurm, Soperator, Kubernetes GPU scheduling, Ray, Argo, Airflow, Metaflow, or similar orchestration tools.
- Experience with ML training infrastructure, model serving, simulation workloads, or large-scale data pipelines.
- Experience supporting enterprise customers, design partners, or production pilots.
- Familiarity with NVIDIA GPU infrastructure, CUDA workloads, Isaac Sim, Omniverse, or simulation at scale.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet.
- Company-paid short-term disability, long-term disability, and life insurance.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.