Senior Software Engineer, Fleet Intelligence Agent Systems

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API CUDA @ 4 Debugging @ 7 Docker @ 4 GPU @ 4 Go @ 7 HPC @ 4 HTTP Helm @ 4 InfiniBand Kubernetes @ 4 Linux @ 7 Networking @ 7 Observability @ 4 OpenTelemetry @ 4 Prometheus @ 4 Rust @ 7 SRE Security

Details

NVIDIA is looking for a Senior Software Engineer to join the DGX Cloud / Fleet Intelligence team. The role involves building reliable systems software for cloud-scale GPU infrastructure, including agent-side systems that collect GPU health, host telemetry, inventory, attestation evidence, and BMC data across NVIDIA's GPU fleet. The position also includes contributing to open-source software used by developers, operators, and customers working with NVIDIA GPU infrastructure.

Responsibilities

  • Design and develop Fleet Intelligence agent software for Linux hosts, bare-metal systems, and Kubernetes GPU nodes.
  • Build telemetry and health collection for NVIDIA GPUs, DCGM/NVML, drivers, CUDA runtime, InfiniBand, containers, kernel and operating system state, CPU, memory, disk, and networking.
  • Develop inventory, enrollment, node identity, local state, attestation, and backend export workflows.
  • Maintain local API, Prometheus metrics, file export, and OTLP/HTTP export paths.
  • Build and improve Kubernetes DaemonSet deployment, systemd service packaging, .deb/.rpm packaging, and container image workflows.
  • Contribute code, tests, documentation, release artifacts, and community-facing engineering practices to open-source Fleet Intelligence agent and collector software.
  • Contribute to out-of-band collection using Redfish/BMC interfaces for inventory, GPU attestation, BMC metrics, and logs.
  • Improve collector concurrency, rate limiting, retry behavior, credential handling, partial-failure handling, and backend submission semantics.
  • Collaborate with backend, infrastructure, SRE, security, and datacenter operations teams to deliver reliable GPU fleet observability.

Requirements

  • 5 or more years of industry software engineering experience and a Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • Strong Go development experience for Linux services, command-line interfaces, agents, or daemons.
  • Rust experience, or a strong willingness to work in Rust, for out-of-band collector development.
  • Strong Linux systems knowledge, including processes, filesystems, networking, service lifecycle, logs, permissions, and host diagnostics.
  • Experience with telemetry, observability, health monitoring, or fleet-management systems.
  • Experience with Docker, Kubernetes, Helm, and production deployment workflows.
  • Experience contributing to open-source projects or working in public repositories with code review, issue tracking, documentation, release notes, and signed commits.
  • Familiarity with secure credential handling, enrollment flows, tokens/JWTs, and service-to-service authentication.
  • Strong debugging skills across hardware-adjacent software, operating systems, containers, and distributed backend integrations.

Preferred Qualifications

  • Experience with NVIDIA datacenter GPUs, DGX systems, DCGM, NVML, CUDA, GPU drivers, XID/SXID events, or GPU diagnostics.
  • Experience building host agents, node agents, collectors, monitoring daemons, or Kubernetes DaemonSets.
  • Experience with Redfish, BMCs, firmware inventory, secure boot state, PCIe device inventory, or hardware attestation.
  • Experience with OpenTelemetry, Prometheus, OTLP gateways, or metric/log export pipelines.
  • Experience operating software in AI, HPC, cloud, or large-scale datacenter environments.
  • Meaningful open-source contributions in systems software, observability, Kubernetes, Linux, hardware telemetry, Rust, or Go ecosystems.

Compensation and Benefits

  • Base salary for Level 3: USD 152,000–241,500 per year.
  • Base salary for Level 4: USD 184,000–287,500 per year.
  • Equity and benefits are also provided.
  • Applications will be accepted at least until August 28, 2026.
  • This posting is for an existing vacancy.
  • NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs