Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 6
GPU @ 4
Machine Learning
Networking
Prioritization @ 7
Project Management @ 6
SRE
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform supporting developers and enterprises from data and model training through production deployment. The company operates across compute, storage, networking, and applied AI.
Responsibilities
- Lead infrastructure and capacity delivery projects, including new region deployments, capacity expansion, and maintenance framework initiatives.
- Coordinate model onboarding and production delivery from infrastructure readiness to customer availability.
- Build and maintain execution plans, identify risks early, and resolve blockers before they affect delivery.
- Work with engineering managers, technical leads, product managers, SRE, infrastructure, networking, security, and other platform teams.
- Facilitate technical decision-making and ensure ownership and accountability across complex, cross-team projects.
- Improve engineering delivery processes to make execution more predictable while maintaining a sustainable pace for engineering teams.
Requirements
- Excellent project management and delivery skills, including the ability to break down ambiguous initiatives into executable plans.
- Ability to identify critical paths, manage complex dependency graphs, and coordinate multiple parallel workstreams.
- Strong risk management and prioritization skills in fast-changing environments.
- Excellent written and verbal communication skills in English.
- Experience leading incident coordination, facilitating discussions, documenting decisions, and driving follow-up actions.
- Strong technical background with the ability to understand engineering discussions, infrastructure dependencies, and architectural trade-offs without being the primary implementer.
Bonus Qualifications
- Experience with GPU infrastructure, AI/ML platforms, or large-scale inference systems.
- Experience delivering cloud infrastructure or data center deployment projects.
- Familiarity with capacity planning, production operations, and reliability engineering.
- Experience in high-growth infrastructure or platform engineering organizations.
Benefits
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
Nebius is an equal opportunity employer committed to an inclusive and diverse workplace. Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility.
More jobs at Nebius
Director, Forward Deployed Engineering
Nebius · United States
USD 270,800-310,000 per year
Forward Deployment Engineering Manager
Nebius · United States
USD 225,800-281,000 per year
IT Risk & Controls Manager
Nebius · United States
USD 120,000-180,000 per year
IT Infrastructure Engineer – RMA & Hardware Diagnostics
Nebius · Kansas City, United States
USD 112,700-140,800 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Similar jobs
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Engineering Manager, Object Storage - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Staff Software Engineer, AI Reliability
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 325,000-485,000 per year
Service Reliability Engineer
Nvidia · United States
USD 168,000-333,500 per year
Engineering Manager, DGX Cloud Production Engineering
Nvidia · United States
USD 224,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year