Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
GPU
Go @ 7
HPC @ 4
InfiniBand
Linux @ 6
MPI
NCCL
Networking
Performance Optimization @ 6
Python @ 7
Software Development @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform and is looking for a Lead Software Systems Engineer - GPU Performance to help build its hyperscaler platform. You will work across core components while analyzing and optimizing the performance of large-scale GPU clusters at the intersection of hardware and software.
You will operate across the full stack—from hardware and system software to networking (InfiniBand/RoCE), virtualization (KVM/QEMU), and distributed communication layers (e.g., MPI, NCCL).
Responsibilities
- Focus on understanding system behavior across multiple layers, identifying performance bottlenecks, and driving improvements that shape how clusters are built, operated, tuned, and validated.
- Investigate and troubleshoot performance issues of GPU clusters under real workloads (training and inference).
- Evaluate and integrate new hardware, system configurations, and tuning approaches through the software stack.
- Support complex performance-related escalations from internal teams and customers.
- Work closely with infrastructure, software engineering, and hardware vendor teams (e.g., NVIDIA, Mellox, Intel).
- Contribute to hardware and cluster qualification (acceptance) to ensure systems meet performance expectations.
Requirements
- 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming).
- 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning).
- In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems.
- Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python).
Coding interviews are part of the process.
Benefits
- Health insurance: 100% company-paid medical, dental and vision coverage for employees and families.
- 401(k) plan: up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: up to $85/month for mobile and internet.
- Disability & life insurance: company-paid short-term, long-term and life insurance coverage.
Benefits & perks also include: competitive compensation, career growth and learning opportunities, flexibility and ownership, collaborative and innovative culture, opportunity to work on impactful AI projects, and an international environment and talented teams.