Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Debugging @ 6
Go @ 7
Linux @ 6
Networking @ 4
Python @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform supporting developers and enterprises from data and model training through production deployment. The company operates globally, with R&D hubs across Europe, the UK, North America, and Israel.
The Senior Software Engineer will design, build, and own backend systems that power metrics, monitor large-scale infrastructure, and support a comprehensive infrastructure maintenance platform. The role requires strong production experience, sound system design judgment, and the ability to operate and improve critical services.
Responsibilities
- Design and build services and agents that provide deep visibility into large-scale server fleets and data center engineering systems.
- Evolve metrics, aggregation, and alerting pipelines, with a focus on signal quality and reliability.
- Design and operate maintenance and remediation systems that enable safe, predictable fleet-wide changes and keep infrastructure healthy.
- Investigate production incidents hands-on, including on-host Linux debugging, and drive root-cause fixes.
- Collaborate closely with hardware, networking, and data center operations teams to improve reliability.
Requirements
- 5+ years of professional software engineering experience.
- Strong production experience with Python and Go, or the ability to ramp up quickly.
- Solid Linux fundamentals and comfort debugging live systems.
- Ability to write reliable, maintainable code and investigate complex, ambiguous problems.
- Experience building and operating production systems at scale.
Preferred Qualifications
- Ubuntu experience, including internal tooling and packaging workflows such as building Debian packages.
- CCNA or equivalent networking experience.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- Paid parental leave: 20 weeks for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet.
- Company-paid short-term, long-term, and life insurance coverage.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.