Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Bash @ 3
CI/CD
Distributed Systems @ 3
Linux @ 5
Load Testing
Networking @ 3
Python @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform supporting developers and enterprises from data and model training through production deployment. The Hardware Infrastructure team designs, develops, and supports systems involved in the data-center lifecycle, including functional and load testing systems, engineering and IT equipment monitoring, asset tracking, hardware repair tracking, and server production.
Responsibilities
- Ensure fault tolerance, scalability, and uninterrupted operations for services.
- Use modern technologies to solve infrastructure problems.
- Implement and improve CI/CD processes.
- Monitor engineering equipment in data centers, including power supply and air and water cooling systems.
- Monitor IT equipment, including racks, servers, JBODs, JBOGs, power shelves, and network devices.
- Collaborate with globally distributed engineering and operations teams.
- Travel occasionally to data centers, especially when not located near one.
Requirements
- Proficiency with Linux systems.
- Expertise in Python and Bash scripting for automation.
- Ability to troubleshoot complex hardware, software, and networking issues.
- Strong analytical and problem-solving skills, with a focus on optimizing system performance.
- Working proficiency in English.
Preferred Qualifications
- Desire to contribute to backend development.
- Experience designing, developing, and operating high-load distributed systems.
Working Conditions
- Primarily remote.
- Occasional travel to data centers is required.
- The company also welcomes employees to work from its office in Amsterdam.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- Paid parental leave: 20 weeks for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet expenses.
- Company-paid short-term disability, long-term disability, and life insurance.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Opportunity to work on impactful AI projects in an international environment.
Compensation
- Base salary of $130,000–$180,000 per year, plus quarterly performance bonuses.
Nebius is an equal opportunity employer committed to an inclusive and diverse workplace. Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire.
More jobs at Nebius
Product Marketing Manager - Token Factory
Nebius · United States
USD 117,800-222,200 per year
Senior GTM Analyst
Nebius · United States, Austin, United States
USD 145,000-172,000 per year
Data Center Technician
Nebius · Kansas City, United States
USD 30-45 per hour
Manager, ML Solutions Architecture - Token Factory
Nebius · United States
USD 228,000-285,000 per year
Data Center GM
Nebius · United States
USD 200,000-250,000 per year
Similar jobs
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Senior MLOps Engineer - DSX Enablement
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Senior Engineer System Software, SDN Operations
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Tech Lead Ethernet Networking Verification Engineer
Nvidia · Austin, United States
USD 184,000-287,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior DevOps Engineer, Platform Engineering
Nvidia · Santa Clara, United States
USD 176,000-276,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year