Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
CUDA @ 4
Debugging
Deep Learning @ 4
GPU @ 4
HPC @ 6
Marketing @ 6
Observability
Performance Optimization
Software Development @ 8
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Firmware Engineer to join the CSP Engagements team, focusing on system software for data center products such as GB200. This role combines embedded firmware development with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment.
Responsibilities
- Design and develop firmware solutions for manageability and observability of data center servers.
- Participate in hardware bring-up activities, out-of-band firmware development, protocol stacks including Redfish, PLDM, MCTP, and NSM, and hardware-software co-design for Cloud Service Provider deployments.
- Debug and troubleshoot NVIDIA GPU firmware issues, power management, performance, and thermal control problems for data center deployments, providing active support to CSPs.
- Partner directly with CSPs to deliver technical solutions, co-develop and co-debug features and optimizations, and provide support during new product introductions.
- Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments.
- Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation.
Requirements
- Deep expertise in data center server architectures, HPC systems, and hardware-software co-design.
- Deep expertise in embedded firmware, server management controllers, and hardware bring-up, with a proven track record of shipping production BMC solutions.
- Strong knowledge of DMTF protocols, including Redfish, IPMI, PLDM, MCTP, and SPDM; telemetry frameworks; and out-of-band management architectures.
- Expert-level skills in C/C++ in resource-constrained embedded environments, RTOS, device drivers, and low-level protocols including I2C, SPI, UART, PCIe, and MCTP.
- Experience with RAS, including error handling, error injection, fault isolation, and system health monitoring.
- Bachelor's or master's degree in Computer Engineering, Computer Science, or a related field, or equivalent experience.
- 8–12 years of system software development experience.
Preferred Qualifications
- Knowledge of cloud- and cluster-level deployment and management systems.
- Experience with GPU computing and CUDA, including deep learning workloads.
- Knowledge of memory fabric and CXL architectures.
NVIDIA develops products and technologies in artificial intelligence, high-performance computing, and visualization. The company works across Software, Hardware, Firmware, Marketing, and Operations to support the successful introduction of next-generation GPU- and CPU-based products. Knowledge of driver, firmware, diagnostics, and software stack development processes and priorities supports keeping complex projects on track.
Compensation and Benefits
The base salary is determined by location, experience, and the pay of employees in similar positions. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. The role is also eligible for equity and benefits.
Applications will be accepted at least until July 30, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.