Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
Ansible @ 3
Bash @ 3
Change Management @ 2
Communication @ 6
GPU @ 3
HPC @ 3
Hiring @ 3
InfiniBand @ 2
Linux @ 3
Networking @ 2
Python @ 3
QA @ 2
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform supporting developers and enterprises from data and model training through production deployment. The company operates globally with R&D hubs across Europe, the UK, North America, and Israel.
The role focuses on leading on-site hardware deployment for new data center locations supporting large-scale AI and GPU infrastructure. This hands-on infrastructure position covers the process from shipment arrival through operational readiness, including server, storage, networking, cabling, configuration, testing, documentation, and production handoff. It is not a traditional systems administrator or facilities-only role.
Responsibilities
Deployment Planning and Coordination
- Own hardware deployment plans for new data center sites, including timelines, scope, resources, and readiness milestones.
- Coordinate with logistics, supply chain, construction, network engineering, infrastructure engineering, and operations teams.
- Verify site readiness, including racks, power, cooling, connectivity, cabling pathways, and installation requirements.
- Review technical documentation, rack elevations, wiring diagrams, bills of materials, deployment runbooks, and site layouts.
- Serve as the on-site point of contact for deployment activities and installation escalations.
On-Site Hardware Deployment
- Supervise and participate in receiving, unpacking, racking, stacking, cabling, labeling, and powering hardware.
- Support deployment of servers, storage systems, network devices, optics, cabling, and related data center infrastructure.
- Ensure hardware is installed according to design standards, cabling standards, safety requirements, and operational procedures.
- Coordinate with vendors, contractors, remote engineers, and local data center staff during installation.
- Support firmware updates, hardware configuration, asset validation, and integration with management or monitoring systems.
Linux, Hardware, and Infrastructure Troubleshooting
- Perform first-line troubleshooting across hardware, cabling, optics, power, network connectivity, and Linux-based hosts.
- Validate server health, network connectivity, management access, and readiness for handoff to operations teams.
- Identify and escalate hardware faults, cabling defects, configuration gaps, and site readiness blockers.
- Work with infrastructure, systems, and network engineering teams to resolve deployment issues.
Quality, Readiness, and Handoff
- Conduct functional checks, acceptance testing, and deployment validation before production handoff.
- Document installation steps, asset data, cabling details, test results, issues, and lessons learned.
- Support transition of deployed infrastructure to operations teams.
- Improve deployment playbooks, hardware standards, and site readiness processes.
- Maintain high standards for safety, documentation, change control, and operational discipline.
Requirements
- 5+ years of experience in data center infrastructure, hardware deployment, IT field engineering, infrastructure operations, or related technical environments.
- Hands-on experience installing server, storage, and network hardware, including racking, cabling, labeling, powering, and validation.
- Experience working with racks, structured cabling, power distribution, cooling requirements, and physical infrastructure dependencies in data centers.
- Ability to troubleshoot hardware, cabling, optics, connectivity, and deployment issues in live infrastructure environments.
- Basic working knowledge of Linux-based systems, including command-line checks, connectivity validation, logs, and host-level troubleshooting.
- Ability to interpret technical documentation, rack elevations, site layouts, wiring diagrams, and deployment runbooks.
- Ability to coordinate deployments across vendors, contractors, logistics, network teams, construction teams, and operations stakeholders.
- Strong organizational, communication, and problem-solving skills.
- Comfort working in fast-paced deployment environments with changing priorities.
- Willingness to travel frequently, sometimes on short notice.
Preferred Qualifications
- Experience supporting GPU, HPC, AI infrastructure, high-density compute, or clustered server environments.
- Familiarity with high-speed networking, optics, fiber cabling, MPO/MTP, InfiniBand, RoCE, or similar technologies.
- Experience with hardware diagnostics, firmware updates, out-of-band management, and server validation workflows.
- Familiarity with asset tracking, change management, QA checklists, and deployment acceptance testing.
- Experience using Linux, Bash, Python, Ansible, Terraform, or similar infrastructure tools.
- Experience in hyperscale, cloud-scale, colocation, or large enterprise data center environments.
- Understanding of health, safety, and environmental standards in data center or industrial settings.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Company-paid short-term disability, long-term disability, and life insurance.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
Compensation
The base compensation range is $85,000–$140,000 USD. The posting also states competitive salaries ranging from $90,000–$140,000. Actual compensation depends on experience, skills, qualifications, hiring level, and geographic location.
Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.