Senior AI Infrastructure Systems Engineer (R&D / GPU / AI)

at Nebius
USD 179,500-224,300 per year
SENIOR
✅ Remote

Tech Stack

AI Communication @ 4 Debugging @ 4 GPU @ 6 Go @ 4 HPC Linux @ 6 NCCL @ 7 NVLink @ 7 Networking Python @ 4

Details

Nebius is building a full-stack AI cloud platform for data and model training through production deployment. The role focuses on diagnosing and resolving complex hardware and platform failures in high-performance AI infrastructure.

The engineer will form technical hypotheses, design targeted tests, analyze system logs and low-level hardware data, and distinguish between GPU, host, firmware, interconnect, thermal, and physical hardware failures. The role requires strong ownership, curiosity, and the ability to turn ambiguous problems into evidence-based conclusions and repeatable troubleshooting procedures.

Responsibilities

  • Diagnose and resolve complex failures across Linux, GPU servers, PCIe, firmware, power, and cooling.
  • Investigate hardware, software, and networking issues through deep problem analysis and root cause investigation.
  • Benchmark systems for performance, stability, power efficiency, thermal behavior, and workload characteristics.
  • Analyze system logs, low-level hardware data, and platform behavior.
  • Develop evidence-based conclusions and repeatable troubleshooting procedures.
  • Optimize performance in cloud or high-performance computing environments.
  • Use scripting and automation to support system investigation and diagnostics.

Requirements

  • At least 5 years of hands-on systems engineering experience across Linux, server hardware, firmware, PCIe, and GPU or high-performance computing platforms.
  • Extensive Linux experience for hardware debugging, system-level troubleshooting, and platform investigation.
  • Knowledge of the Linux kernel and experience with kernel-level debugging or troubleshooting.
  • Strong knowledge of modern server architecture, particularly in high-performance, GPU-based environments.
  • Strong knowledge of NVIDIA GPU platforms and diagnostic tooling, including nvidia-smi, XID and SXID analysis, GSP and driver behavior, NVLink, NVSwitch, Fabric Manager, DCGM diagnostics, and NCCL testing.
  • Deep knowledge of PCIe, including topology, enumeration, root complexes, endpoints, switches, bridges, retimers, link speed and width, AER and DPC errors, completion timeouts, and link failures.
  • Understanding of how protocol-level errors can relate to physical-layer or signal-integrity problems.
  • In-depth understanding of firmware interactions across BIOS/UEFI, BMC, CPLD/FPGA, GPU firmware, NIC firmware, and other platform components.
  • Experience with low-level hardware communication and debugging using I²C, SMBus, PMBus, register maps, bit masks, byte- and word-level data, and device datasheets.
  • Strong analytical and problem-solving skills with a performance-first mindset.
  • Scripting and automation experience using languages such as Python and Go.

Benefits

  • 100% company-paid medical, dental, and vision coverage for employees and families.
  • 401(k) plan with up to a 4% company match and immediate vesting.
  • 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
  • Remote work reimbursement of up to $85 per month for mobile and internet.
  • Company-paid short-term, long-term, and life insurance coverage.
  • Career growth and learning opportunities.
  • Flexibility and ownership.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment and talented teams.

Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.

More jobs at Nebius

Similar jobs