General Manager, Data Center (New Build)

at Nebius
USD 240,500-300,600 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Change Management @ 4 Compliance GPU @ 4 HPC Leadership @ 8 Networking @ 4 Reporting @ 4 Security

Details

Nebius is building a full-stack AI cloud platform for developers and enterprises, covering data and model training through production deployment. The company operates large-scale GPU orchestration, inference optimization, compute, storage, networking, and applied AI infrastructure.

This role will build and lead the organization responsible for deploying and operating Nebius AI infrastructure at the Independence, Missouri campus. The position is responsible for data center IT operations, critical facilities, networking, logistics, operational readiness, and supporting functions across a gigawatt-scale data center campus.

Responsibilities

Organization Design and Leadership

  • Design and build the campus organization, including functional leadership, building and shift coverage, reporting lines, and decision authority.
  • Hire, lead, and develop directors, data center managers, and operations teams in partnership with Talent Acquisition and HR.
  • Own the phased workforce plan, aligning recruitment, training, contractor support, and leadership coverage with hardware deployment, acceptance testing, and production milestones.
  • Establish performance expectations, technical competency standards, and succession plans.

Operational Deployment and Readiness

  • Own the integrated operational readiness plan covering facility completion, equipment delivery, hardware installation, network and storage integration, staffing, and customer service readiness.
  • Coordinate with construction and facilities teams to verify power, cooling, and space readiness before deployment.
  • Define operational acceptance criteria and coordinate readiness and acceptance testing with technical owners.
  • Ensure trained teams, tested procedures, critical spares, asset records, monitoring, and vendor support are in place before production operations begin.
  • Coordinate GPU, server, network fabric, storage, and cloud readiness with technical owners.
  • Sequence hardware installation, integration, and production changes to protect live environments as additional compute capacity comes online.

Campus Operations and Reliability

  • Own end-to-end campus operations across critical facilities and IT infrastructure, delivering agreed uptime, performance, and service levels.
  • Establish 24/7 coverage, maintenance governance, change control, incident command, root-cause analysis, and business continuity practices.
  • Ensure clear responsibility across power, liquid cooling, controls, hardware, networking, and cloud interfaces.
  • Use monitoring, automation, and performance reporting to improve deployment throughput, repair turnaround, reliability, and operating efficiency.

Financial Capacity and Asset Management

  • Own campus operating budgets, workforce costs, vendor expenditure, and forecasts; contribute to capital planning and investment priorities.
  • Align power, cooling, rack space, equipment, and staffing capacity with deployment plans.
  • Establish asset inventory, receiving, logistics, critical spares, repair, warranty-return, and secure disposition processes.
  • Track budget variance, rework, delayed capacity, and operational losses, and drive corrective actions.

Partners, City Coordination, Safety, and Executive Accountability

  • Serve as the primary operational interface with utilities, infrastructure providers, equipment manufacturers, vendors, and service partners.
  • Coordinate with City of Independence departments, Public Affairs, and Legal on site readiness, permitting, inspections, utility interfaces, emergency preparedness, and ongoing operational matters.
  • Hold providers accountable for safety, quality, delivery milestones, response times, and contractual service performance.
  • Ensure safety, environmental, physical security, and applicable compliance requirements are implemented.
  • Provide executive visibility into launch progress, workforce readiness, operating performance, financial forecasts, and risks.

Operational Scalability and Leadership Development

  • Develop repeatable deployment standards, staffing models, acceptance practices, and operating procedures as compute capacity scales.
  • Build leadership succession and delegation practices.
  • Share lessons and operational improvements with the broader Nebius infrastructure organization.

Requirements

  • 10+ years of experience in data center, cloud infrastructure, or comparable mission-critical operations, including senior leadership responsibility.
  • Experience building and scaling organizations and leading managers across multiple technical and operational functions.
  • Experience leading hardware deployment, infrastructure integration, operational acceptance, or major compute capacity expansion into production service.
  • Strong understanding of critical facilities, including power and cooling, and their impact on IT availability and high-density compute operations.
  • Leadership experience across compute, networking, storage, and infrastructure lifecycle management.
  • Experience owning operating budgets, workforce planning, vendor performance, and executive reporting.
  • Experience leading incident response, change management, risk reduction, and service improvement in high-availability environments.
  • Ability to exercise sound judgment, make decisions with incomplete information, address underperformance, and align stakeholders around clear outcomes.
  • Ability to work on site in Independence, Missouri through deployment, launch, and ongoing operations.

Nice to Have

  • Experience in hyperscale, AI, GPU, or high-performance computing environments.
  • Experience with liquid-cooled environments and the interfaces between cooling systems, GPU platforms, and network fabrics.
  • Experience managing multiple campuses or regional operations.
  • Experience building operating organizations during initial hardware deployment and transition into production service.
  • Experience using infrastructure automation, telemetry, and reliability data to improve capacity delivery and operating performance.

Benefits

  • 100% company-paid medical, dental, and vision coverage for employees and families.
  • 401(k) plan with up to a 4% company match and immediate vesting.
  • Paid parental leave: 20 weeks for primary caregivers and 12 weeks for secondary caregivers.
  • Remote work reimbursement of up to $85 per month for mobile and internet.
  • Company-paid short-term disability, long-term disability, and life insurance.
  • Career growth and learning opportunities.
  • Flexibility and ownership.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment and talented teams.

Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.

More jobs at Nebius

Similar jobs