Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 6
Compliance @ 3
GPU @ 3
HPC @ 3
Leadership @ 5
Linux @ 5
Machine Learning
Networking @ 6
Project Management @ 6
R
Security @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About Nebius
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The Role
The Data Center Manager owns end-to-end reliability, safety, capacity, and performance for one of our flagship U.S. sites. You’ll lead a high-performing, multi-disciplinary operations team and partner tightly with Design, Build, Network, Security, Capacity Planning, and the DC orgs to deliver world-class availability and cost efficiency.
Responsibilities
- Lead day-to-day data center operations in a 24/7 mission-critical environment
- Manage and develop a team of 15–20 Data Center Technicians
- Oversee installation, break/fix activities, and field change orders
- Ensure timely delivery of tasks aligned with KPIs and operational milestones
- Monitor infrastructure performance; drive troubleshooting, incident response, and root cause analysis
- Own incident management, including resolution and post-incident reviews
- Plan and execute capacity expansion (rack, block, and site growth)
- Maintain physical security, access controls, and compliance with standards
- Partner cross-functionally (Engineering, Build, Site Selection, Operations)
- Manage vendors and contractors to deliver high-quality, cost-effective solutions
- Drive continuous improvement across processes, efficiency, and reliability
- Support hiring and team scaling efforts
Requirements
- 5+ years of experience in data center operations; 2+ years in a leadership role
- Strong knowledge of servers, storage, networking, and data center infrastructure
- Experience with power systems (UPS, backup), cooling, and physical infrastructure
- Proficiency in Linux environments
- Experience with incident management and operational processes in high-availability environments
- Strong project management experience (budgeting, vendor management, resource planning)
- Understanding of security, compliance, and disaster recovery best practices
- Ability to work cross-functionally and drive execution across teams
- Strong leadership, communication, and problem-solving skills
- Ability to lift up to 50 lbs and support on-site operational needs
- Willingness to participate in a 24/7 on-call rotation
- Bachelor’s degree in IT, Computer Science, or related field (or equivalent experience)
It would be an added bonus if you have
- Familiarity with ITIL / ITSM processes
- Experience with GPU clusters, HPC, or cloud infrastructure
- Understanding of data center network traffic patterns (east-west and north-south)
- Experience with data center management and monitoring tools
Key employee benefits
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan: up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: up to $85/month for mobile and internet.
- Disability & life insurance: company-paid short-term, long-term and life insurance coverage.
Compensation
We offer competitive salaries ranging from $115K to $275K OTE, which includes base salary and performance bonus. Equity in the form of RSUs may be available at certain salary grades.
Join Nebius Today!
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams