Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 7
Debugging @ 4
GPU
Git @ 4
Jira @ 4
Leadership @ 6
Project Management @ 4
Python @ 7
Software Development @ 4
System Architecture @ 4
Technical Leadership @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an expert engineer and technical architect to own the end-to-end manageability architecture for next-generation rack-level AI supercomputing platforms and data center products based on NVIDIA GPUs and Grace solutions. The role involves working with internal component leads, external partners, data center architects, and cloud customers to define requirements, implement firmware and software modules, and deliver reliable, high-quality products to market.
Responsibilities
- Drive server management for large clusters and data centers deploying NVIDIA GPUs and Grace solutions.
- Work with data center architects and cloud customers to define implementation requirements and accelerate product development.
- Coordinate with internal teams to ensure requirements are correctly designed and implemented across firmware and software modules.
- Collaborate with technical leads to design and build data center health management workflows.
- Drive reliability and optimization in firmware architecture from a data center perspective.
- Work closely with cluster bring-up teams to resolve issues quickly.
- Own firmware delivered to data centers, including its quality, reliability, and telemetry performance.
- Work with component leads and customers to align architecture with customer requirements.
Requirements
- 15 or more years of relevant experience working on server firmware, including BMC, and platform software development.
- BS, MS, or PhD in Electrical Engineering, Computer Science, or a related field, or equivalent experience.
- Hands-on experience with data center health management workflows.
- Proven record of delivering server firmware for large data centers.
- Strong knowledge of data center management, server architecture, and server manageability.
- Strong, demonstrable programming skills in C, C++, and Python.
- Experience programming and debugging server platforms.
- Experience with source-control management tools such as Git or Perforce.
- Experience with project management tools such as Jira.
- Excellent written and oral communication skills, strong teamwork, good work ethics, and commitment to delivering high-quality results.
- Self-starter with hands-on coding experience and the ability to find creative solutions to complex problems.
Preferred Qualifications
- Hands-on experience with data center health management.
- Hands-on experience with x86 or ARM system architecture.
- Proven technical leadership driving large, complex problems involving more than 50 engineers.
Benefits
- Equity and employee benefits are available.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
Applications will be accepted at least until August 2, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Systems Software Engineer, Low Latency Streaming Technology - Automotive
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Reinforcement Learning Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Manager, Storage Engineering
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Senior Software and System Architect
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Customer Technical Program Manager
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Similar jobs
Principal Architect, System Software - Orbital Data Center
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Libraries Engineer – AI and HPC
Nvidia · Poland
PLN 221,200-507,000 per year
Senior Software Engineer - Image and Data Processing Libraries
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal System Software Engineer - AV Platform
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Math Libraries Engineer – AI and HPC
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Tech Lead Ethernet Networking Verification Engineer
Nvidia · Austin, United States
USD 184,000-287,500 per year
Senior Manager, Engineering - Data Center Firmware
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Math Libraries Engineer - Sparsity in AI
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year