Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Bash @ 6
Debugging @ 4
GPU @ 4
Go @ 6
HPC
Linux @ 6
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform and is looking for an Embedded Software Developer to design and implement firmware and low-level software that powers its next-generation GPU and HPC platforms. This role focuses on embedded control, board management, telemetry, and hardware-firmware integration to ensure reliable operation in high-density, mission-critical environments.
Responsibilities
- Design and implement embedded firmware for server management, telemetry, and control systems.
- Maintain and enhance custom OpenBMC firmware with new features and improvements.
- Enable real-time monitoring of power, thermal sensors, and hardware health.
- Work closely with hardware engineers to validate firmware for existing and future platforms.
- Debug and optimize low-level drivers and protocols.
- Contribute to long-term firmware architecture for GPU cluster reliability.
Requirements
- 5+ years in embedded systems or firmware development.
- Proficiency in embedded Linux.
- Hands-on experience with BMCs, microcontrollers, or SoC firmware.
- Understanding of hardware bring-up and debugging.
- Languages: C, C++, Bash, Go, YAML.
- Firmware: OpenBMC, U-Boot, Linux Kernel.
- Interfaces: I2C, I3C, SPI, eSPI, UART, LPC.
- Protocols: SMBus, PCIe, PMBus, PECI.
- Build Systems: Meson, CMake.
- Descriptors & Formats: FRU, SMBIOS, ACPI, DMI.
Preferred
- Knowledge of the Yocto Project principles.
- Knowledge of systems and D-Bus principles.
- Proficiency in C++.
- Good knowledge of C, sufficient for periodic work with Linux drivers and the U-Boot bootloader.
- Experience in developing Linux drivers implementing sysfs and hwmon interfaces.
- Experience with server BMC firmware: IPMI, IPMB, KCS, SSIF, Redfish, PLDM.
- Knowledge of GPU/CPU telemetry frameworks (e.g., NVML, DCGM).
- Exposure to firmware security (Secure Boot, signed firmware).
- Experience with RAS (Reliability, Availability, Serviceability).
- Background in high-performance computing or data center hardware.
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
More jobs at Nebius
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior Machine Learning Engineer, LLM Inference Optimization
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior Director, Global Systems Integrator (GSI) Partnerships
Nebius · United States
USD 158,800-198,400 per year
Senior GTM Analyst
Nebius · United States
USD 109,500-136,800 per year
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Distinguished Engineer, Storage – AI Cloud
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Principal Software Engineer - Rack Scale Systems Infrastructure
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Developer Tools DevOps Engineer
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior GPU and HPC Infrastructure Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year