Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
GPU
Go @ 7
HPC
InfiniBand
Linux @ 6
MPI
NCCL
Networking
Performance Optimization @ 6
Python @ 7
Software Development @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About Nebius
Nebius is building a full-stack AI cloud platform for developers and enterprises, supporting workloads from data and model training through production deployment. The platform spans large-scale GPU orchestration, inference optimization, compute, storage, networking, and applied AI.
The company is headquartered in Amsterdam, listed on Nasdaq as NBIS, and has R&D hubs across Europe, the UK, North America, and Israel. Its team of more than 1,500 people includes hundreds of engineers specializing in hardware, software, and AI R&D.
This role will contribute to building Nebius's hyperscaler platform while analyzing and optimizing the performance of large-scale GPU clusters at the intersection of hardware and software. The work spans hardware and system software, networking such as InfiniBand and RoCE, virtualization with KVM/QEMU, and distributed communication layers including MPI and NCCL.
Responsibilities
- Analyze system behavior across multiple layers, identify performance bottlenecks, and drive improvements that influence how clusters are built, operated, tuned, and validated.
- Investigate and troubleshoot GPU cluster performance issues under real training and inference workloads.
- Evaluate and integrate new hardware, system configurations, and tuning approaches throughout the software stack.
- Support complex performance-related escalations from internal teams and customers.
- Work closely with infrastructure, software engineering, and hardware vendor teams, including NVIDIA, Mellanox, and Intel.
- Contribute to hardware and cluster qualification and acceptance, ensuring systems meet performance expectations.
- Participate in coding interviews as part of the hiring process.
Requirements
- 5+ years of professional experience in system-level software development, with a focus on performance optimization and low-level programming.
- 3+ years of hands-on experience with Linux systems, including administration, troubleshooting, and performance tuning.
- In-depth understanding of server architecture, including PCIe devices, NICs, the Linux operating system and kernel, and high-performance computing systems.
- Strong proficiency in one or more performance-oriented programming languages: C, C++, Go, or Python.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet expenses.
- Company-paid short-term disability, long-term disability, and life insurance.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
Nebius is an equal opportunity employer committed to fostering an inclusive and diverse workplace. Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire.