Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CI/CD @ 4
Debugging @ 7
GPU
Hiring @ 4
IaC
Linux @ 7
Networking @ 4
Observability @ 6
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
About Nebius
Nebius is building a full-stack AI cloud platform for developers and enterprises, supporting workloads from data and model training through production deployment. The platform covers compute, storage, networking, and applied AI, including large-scale GPU orchestration and inference optimization.
Headquartered in Amsterdam and listed on Nasdaq as NBIS, Nebius has a global footprint with R&D hubs across Europe, the UK, North America, and Israel. The company has more than 1,500 employees, including hundreds of engineers focused on hardware, software, and AI R&D.
Responsibilities
- Define and own reliability goals for network services and critical paths, including SLIs, SLOs, availability targets, and error budgets where appropriate.
- Drive reliability improvements across the network, including services, site readiness, inter-site connectivity, data center interconnectivity, and operational standards.
- Own incident response for assigned areas, lead investigations and postmortems, and turn failures into durable fixes.
- Build and evolve observability through actionable metrics, logs, traces, alerting, and faster debugging workflows.
- Design safer change workflows using automation, CI/CD, test and staging environments, canarying, rollbacks, and auditability for network changes.
- Collaborate with network engineers and platform teams to embed operability into designs and keep operations practical and efficient.
Requirements
- Strong production Linux fundamentals and a structured approach to debugging complex systems.
- Solid understanding of networking fundamentals and real-world network failures, including control plane and data plane behavior, latency, packet loss, and failure domains.
- Hands-on experience operating high-availability systems and improving them over time.
- Ability to write and maintain software and automation. Go is commonly used at Nebius; Python is also welcome.
- Experience with modern infrastructure tooling, such as infrastructure as code, CI/CD, and container platforms, and comfort automating operational workflows.
Preferred Qualifications
- Experience with high-throughput traffic processing, including load balancers, tunneling or decapsulation, NAT64, or similar datapath-heavy systems.
- Low-level networking performance and debugging experience with technologies such as eBPF/XDP, DPDK, perf/ftrace, or kernel networking internals.
- Experience building network-safe delivery pipelines, including testing labs, staged rollouts, automated verification, and drift detection.
- Background with large-scale network observability and telemetry, such as routing or flow telemetry and regression detection at scale.
Compensation
The base compensation range is $179,500–$224,300 USD per year. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, hiring level, and geographic location.
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
Nebius is an equal opportunity employer committed to fostering an inclusive and diverse workplace. Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Accommodations are available during the application process.