Bare Metal Infrastructure Engineer

at Nebius
USD 90,000-140,000 per year
MIDDLE
✅ On-site

Tech Stack

AI @ 3 Ansible @ 3 Bash @ 3 Change Management @ 2 Communication @ 6 GPU @ 3 HPC @ 3 InfiniBand @ 2 Linux @ 6 Machine Learning Networking @ 2 Python @ 3 QA @ 2 Terraform @ 3

Details

About Nebius

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The Role

We are seeking a Bare Metal Infrastructure Engineer to lead on site hardware deployment for new data center locations supporting large scale AI and GPU infrastructure.

This is a hands-on infrastructure role focused on bringing new sites from shipment arrival through full operational readiness. You will coordinate and supervise the deployment of servers, storage, networking, cabling, and supporting infrastructure, ensuring equipment is installed, configured, tested, documented, and ready for production use.

This is not a traditional systems administrator role and not a facilities-only role. We are looking for someone who is comfortable working close to the hardware in live data center environments, with strong knowledge of rack-level infrastructure, server installation, cabling, power/cooling readiness, basic Linux-based troubleshooting, and cross-functional deployment coordination.

Responsibilities

Deployment Planning & Coordination

  • Own the hardware deployment plan for new data center sites, including timelines, scope, resources, and readiness milestones.
  • Coordinate with logistics, supply chain, construction, network engineering, infrastructure engineering, and operations teams.
  • Verify site readiness before deployment, including racks, power, cooling, connectivity, cabling pathways, and installation requirements.
  • Review technical documentation, rack elevations, wiring diagrams, BOMs, deployment runbooks, and site layouts.
  • Serve as the on-site point of contact for deployment activities and escalation during installation.

On-Site Hardware Deployment

  • Supervise and participate in receiving, unpacking, racking, stacking, cabling, labeling, and powering hardware.
  • Support deployment of servers, storage systems, network devices, optics, cabling, and related data center infrastructure.
  • Ensure hardware is installed according to design standards, cabling standards, safety requirements, and operational procedures.
  • Coordinate with vendors, contractors, remote engineers, and local data center staff during installation.
  • Support firmware updates, hardware configuration, asset validation, and integration with management or monitoring systems.

Linux, Hardware & Infrastructure Troubleshooting

  • Perform first-line troubleshooting during deployment across hardware, cabling, optics, power, network connectivity, and Linux-based hosts.
  • Validate server health, network connectivity, management access, and readiness for handoff to operations teams.
  • Identify and escalate issues involving hardware faults, cabling defects, configuration gaps, or site readiness blockers.
  • Work with infrastructure, systems, and network engineering teams to resolve deployment issues quickly and safely.

Quality, Readiness & Handoff

  • Conduct functional checks, acceptance testing, and deployment validation before production handoff.
  • Document installation steps, asset data, cabling details, test results, issues, and lessons learned.
  • Support transition of deployed infrastructure to operations teams for steady-state management.
  • Contribute feedback to improve deployment playbooks, hardware standards, and site readiness processes.
  • Maintain high standards for safety, documentation, change control, and operational discipline.

Requirements

  • 5+ years of experience in data center infrastructure, hardware deployment, IT field engineering, infrastructure operations, or related technical environments.
  • Hands-on experience with server, storage, and network hardware installation, including racking, cabling, labeling, powering, and validation.
  • Experience working in data center environments with racks, structured cabling, power distribution, cooling requirements, and physical infrastructure dependencies.
  • Ability to troubleshoot hardware, cabling, optics, connectivity, and deployment issues in a live infrastructure environment.
  • Basic working knowledge of Linux-based systems, including comfort with command-line checks, connectivity validation, logs, or host-level troubleshooting.
  • Ability to read and interpret technical documentation, rack elevations, site layouts, wiring diagrams, and deployment runbooks.
  • Proven ability to coordinate deployments across vendors, contractors, logistics, network teams, construction teams, and operations stakeholders.
  • Strong organizational, communication, and problem-solving skills.
  • Comfortable working in fast-paced deployment environments with changing priorities.
  • Willingness to travel frequently, sometimes on short notice.

Preferred Qualifications

  • Experience supporting GPU, HPC, AI infrastructure, high-density compute, or clustered server environments.
  • Familiarity with high-speed networking, optics, fiber cabling, MPO/MTP, InfiniBand, RoCE, or similar technologies.
  • Experience with hardware diagnostics, firmware updates, out-of-band management, and server validation workflows.
  • Familiarity with asset tracking, change management, QA checklists, and deployment acceptance testing.
  • Experience using Linux, Bash, Python, Ansible, Terraform, or similar tools in infrastructure environments.
  • Experience working in hyperscale, cloud-scale, colocation, or large enterprise data center environments.
  • Understanding of health, safety, and environmental standards in data center or industrial settings.

Key Employee Benefits in the US

  • Health Insurance: 100% company-paid medical, dental, and vision coverage for employees and families.
  • 401(k) Plan: Up to 4% company match with immediate vesting.
  • Parental Leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
  • Disability & Life Insurance: Company-paid short-term, long-term, and life insurance coverage.

Compensation

We offer competitive salaries, ranging from $90 -$140k.

Base Compensation Range: $85,000 — $140,000 USD

More jobs at Nebius

Similar jobs