Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 1
Algorithms @ 7
Communication @ 6
Customer Support
Debugging @ 4
GPU
GenAI
Generative AI
Git @ 4
Grafana @ 7
HPC
Jira @ 4
Machine Learning @ 4
Prometheus @ 7
Python @ 7
System Architecture @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an expert engineer to help design rack-level solutions for next-generation AI supercomputing platforms based on the NVIDIA GH200 superchip, GPUs, and Grace solutions. The role focuses on fleet management, telemetry, fleet health monitoring, and fault remediation at scale for HPC and generative AI workloads.
Responsibilities
- Drive next-generation fleet management solutions for scaling AI infrastructure using NVIDIA GPUs and Grace solutions.
- Work with customers, product management, architects, and engineering teams to define implementation requirements.
- Develop architectures for fleet health monitoring and fault-remediation solutions at scale, using both in-band and out-of-band capabilities.
- Create proofs of concept to validate architecture and product designs.
- Educate customers about product architecture and incorporate feedback into product improvements.
- Write architecture specifications and design documents, and own end-to-end product delivery across teams.
- Review code produced from architecture specifications.
- Work with development teams to improve unit testing and establish comprehensive test plans.
- Drive product lifecycles with QA teams and act as the product owner for productized code.
- Define requirements in Jira and bug-management tools and develop end-to-end execution plans with other managers.
- Contribute to all phases of product development, including product definition, architecture, design, implementation, debugging, testing, and early customer support.
Requirements
- BS, MS, or PhD in Electrical Engineering, Computer Science, or a related field, or equivalent experience.
- At least 5 years of hands-on coding experience.
- Strong knowledge of time-series databases such as InfluxDB and Prometheus.
- Strong knowledge of building and consuming REST APIs; Redfish experience is a plus.
- Strong knowledge of telemetry visualization solutions such as Grafana and InfluxDB.
- Strong knowledge of firmware architecture and optimizing firmware for low-latency APIs.
- Strong knowledge of analyzing algorithms for time and space complexity and projecting system resource requirements.
- Proven record of developing scalable solutions.
- Strong, demonstrable programming skills in C/C++ and Python.
- Experience programming and debugging server platforms.
- Experience with source-code management systems such as Git or Perforce and project-management tools such as Jira.
- Excellent written and oral communication skills, work ethic, teamwork, quality focus, and commitment to completing tasks.
- Self-starter with hands-on coding skills and the ability to develop creative solutions to complex problems.
Preferred Qualifications
- Experience building telemetry collection and analysis engines.
- Experience with Redfish and notification systems such as PagerDuty.
- Active contribution to Open Compute Project or DMTF initiatives in relevant areas.
- Hands-on experience with x86 or ARM system architecture.
- Familiarity with confidential computing.
- Experience with machine learning and multivariable optimization techniques.
Benefits
- Equity and NVIDIA benefits are included.
- NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until July 30, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Software Solutions Engineer
Nvidia · Poland
PLN 230,200-487,500 per year
Senior Offensive Security Engineer, Automotive
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Systems Operations and Administrator
Nvidia · Santa Clara, United States
USD 112,000-218,500 per year
Senior Software Engineer, Agent Simulation and Evaluation
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior AI Product Engineer
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Similar jobs
Principal Platform Software Engineer - RAS
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Systems Software Engineer - Fleet Debuggability
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Software Engineer - Local AI
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Engineering Manager, Agentic GenAI Platform
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year