Senior Platform Telemetry Engineer

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI API @ 1 Algorithms @ 7 Communication @ 6 Customer Support Debugging @ 4 GPU GenAI Generative AI Git @ 4 Grafana @ 7 HPC Jira @ 4 Machine Learning @ 4 Prometheus @ 7 Python @ 7 System Architecture @ 4

Details

NVIDIA is seeking an expert engineer to help design rack-level solutions for next-generation AI supercomputing platforms based on the NVIDIA GH200 superchip, GPUs, and Grace solutions. The role focuses on fleet management, telemetry, fleet health monitoring, and fault remediation at scale for HPC and generative AI workloads.

Responsibilities

  • Drive next-generation fleet management solutions for scaling AI infrastructure using NVIDIA GPUs and Grace solutions.
  • Work with customers, product management, architects, and engineering teams to define implementation requirements.
  • Develop architectures for fleet health monitoring and fault-remediation solutions at scale, using both in-band and out-of-band capabilities.
  • Create proofs of concept to validate architecture and product designs.
  • Educate customers about product architecture and incorporate feedback into product improvements.
  • Write architecture specifications and design documents, and own end-to-end product delivery across teams.
  • Review code produced from architecture specifications.
  • Work with development teams to improve unit testing and establish comprehensive test plans.
  • Drive product lifecycles with QA teams and act as the product owner for productized code.
  • Define requirements in Jira and bug-management tools and develop end-to-end execution plans with other managers.
  • Contribute to all phases of product development, including product definition, architecture, design, implementation, debugging, testing, and early customer support.

Requirements

  • BS, MS, or PhD in Electrical Engineering, Computer Science, or a related field, or equivalent experience.
  • At least 5 years of hands-on coding experience.
  • Strong knowledge of time-series databases such as InfluxDB and Prometheus.
  • Strong knowledge of building and consuming REST APIs; Redfish experience is a plus.
  • Strong knowledge of telemetry visualization solutions such as Grafana and InfluxDB.
  • Strong knowledge of firmware architecture and optimizing firmware for low-latency APIs.
  • Strong knowledge of analyzing algorithms for time and space complexity and projecting system resource requirements.
  • Proven record of developing scalable solutions.
  • Strong, demonstrable programming skills in C/C++ and Python.
  • Experience programming and debugging server platforms.
  • Experience with source-code management systems such as Git or Perforce and project-management tools such as Jira.
  • Excellent written and oral communication skills, work ethic, teamwork, quality focus, and commitment to completing tasks.
  • Self-starter with hands-on coding skills and the ability to develop creative solutions to complex problems.

Preferred Qualifications

  • Experience building telemetry collection and analysis engines.
  • Experience with Redfish and notification systems such as PagerDuty.
  • Active contribution to Open Compute Project or DMTF initiatives in relevant areas.
  • Hands-on experience with x86 or ARM system architecture.
  • Familiarity with confidential computing.
  • Experience with machine learning and multivariable optimization techniques.

Benefits

  • Equity and NVIDIA benefits are included.
  • NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.

Applications will be accepted at least until July 30, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs