Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
GPU @ 7
Go @ 4
HPC @ 4
InfiniBand @ 4
LLM
Machine Learning
Networking @ 6
Python @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform supporting developers and enterprises from data and model training through production deployment. The platform covers large-scale GPU orchestration, inference optimization, compute, storage, networking, and applied AI.
The GPU Cluster Architect will drive the design of next-generation AI infrastructure and make end-to-end architectural decisions across compute, networking, and storage. The role focuses on platforms that meet the scale, performance, and reliability requirements of modern AI workloads, including the interconnection, cooling, powering, and optimization of tens of thousands of GPUs across multiple data center sites. The position can be performed remotely from the United States.
Responsibilities
- Architect scalable GPU cluster topologies, including compute nodes, interconnects, storage, and control planes.
- Analyze AI/ML workloads, including large language model training and inference, to inform design trade-offs involving latency, bandwidth, and GPU density.
- Align with network architects and validate low-latency, high-throughput interconnects, including InfiniBand HDR/NDR and RoCEv2, at pod and data center scale.
- Work with storage teams to optimize performance for training datasets, checkpointing, and other workloads.
- Analyze monitoring-system signals to detect issues and flows in the design.
- Partner with site reliability, networking, storage, and data center engineering teams to operationalize and scale the architecture.
Requirements
- 5+ years of experience designing clusters.
- Deep understanding of modern GPU architectures, including NVIDIA and AMD.
- Experience with HPC interconnects, including InfiniBand and RoCE.
- Solid background in systems architecture, networking, and hardware reliability.
- Experience with scripting for automation and telemetry pipelines using Python, Go, or similar technologies.
- Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet expenses.
- Company-paid short-term disability, long-term disability, and life insurance.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
Nebius is an equal opportunity employer committed to an inclusive and diverse workplace. Accommodations are available during the application process.