Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 4
Debugging @ 4
GPU @ 6
Go @ 4
HPC
Linux @ 6
NCCL @ 7
NVLink @ 7
Networking
Python @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform for data and model training through production deployment. The role focuses on diagnosing and resolving complex hardware and platform failures in high-performance AI infrastructure.
The engineer will form technical hypotheses, design targeted tests, analyze system logs and low-level hardware data, and distinguish between GPU, host, firmware, interconnect, thermal, and physical hardware failures. The role requires strong ownership, curiosity, and the ability to turn ambiguous problems into evidence-based conclusions and repeatable troubleshooting procedures.
Responsibilities
- Diagnose and resolve complex failures across Linux, GPU servers, PCIe, firmware, power, and cooling.
- Investigate hardware, software, and networking issues through deep problem analysis and root cause investigation.
- Benchmark systems for performance, stability, power efficiency, thermal behavior, and workload characteristics.
- Analyze system logs, low-level hardware data, and platform behavior.
- Develop evidence-based conclusions and repeatable troubleshooting procedures.
- Optimize performance in cloud or high-performance computing environments.
- Use scripting and automation to support system investigation and diagnostics.
Requirements
- At least 5 years of hands-on systems engineering experience across Linux, server hardware, firmware, PCIe, and GPU or high-performance computing platforms.
- Extensive Linux experience for hardware debugging, system-level troubleshooting, and platform investigation.
- Knowledge of the Linux kernel and experience with kernel-level debugging or troubleshooting.
- Strong knowledge of modern server architecture, particularly in high-performance, GPU-based environments.
- Strong knowledge of NVIDIA GPU platforms and diagnostic tooling, including
nvidia-smi, XID and SXID analysis, GSP and driver behavior, NVLink, NVSwitch, Fabric Manager, DCGM diagnostics, and NCCL testing. - Deep knowledge of PCIe, including topology, enumeration, root complexes, endpoints, switches, bridges, retimers, link speed and width, AER and DPC errors, completion timeouts, and link failures.
- Understanding of how protocol-level errors can relate to physical-layer or signal-integrity problems.
- In-depth understanding of firmware interactions across BIOS/UEFI, BMC, CPLD/FPGA, GPU firmware, NIC firmware, and other platform components.
- Experience with low-level hardware communication and debugging using I²C, SMBus, PMBus, register maps, bit masks, byte- and word-level data, and device datasheets.
- Strong analytical and problem-solving skills with a performance-first mindset.
- Scripting and automation experience using languages such as Python and Go.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet.
- Company-paid short-term, long-term, and life insurance coverage.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.