Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
CUDA @ 4
Debugging @ 7
Docker @ 4
GPU @ 4
Go @ 7
HPC @ 4
HTTP
Helm @ 4
InfiniBand
Kubernetes @ 4
Linux @ 7
Networking @ 7
Observability @ 4
OpenTelemetry @ 4
Prometheus @ 4
Rust @ 7
SRE
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Software Engineer to join the DGX Cloud / Fleet Intelligence team. The role involves building reliable systems software for cloud-scale GPU infrastructure, including agent-side systems that collect GPU health, host telemetry, inventory, attestation evidence, and BMC data across NVIDIA's GPU fleet. The position also includes contributing to open-source software used by developers, operators, and customers working with NVIDIA GPU infrastructure.
Responsibilities
- Design and develop Fleet Intelligence agent software for Linux hosts, bare-metal systems, and Kubernetes GPU nodes.
- Build telemetry and health collection for NVIDIA GPUs, DCGM/NVML, drivers, CUDA runtime, InfiniBand, containers, kernel and operating system state, CPU, memory, disk, and networking.
- Develop inventory, enrollment, node identity, local state, attestation, and backend export workflows.
- Maintain local API, Prometheus metrics, file export, and OTLP/HTTP export paths.
- Build and improve Kubernetes DaemonSet deployment, systemd service packaging,
.deb/.rpmpackaging, and container image workflows. - Contribute code, tests, documentation, release artifacts, and community-facing engineering practices to open-source Fleet Intelligence agent and collector software.
- Contribute to out-of-band collection using Redfish/BMC interfaces for inventory, GPU attestation, BMC metrics, and logs.
- Improve collector concurrency, rate limiting, retry behavior, credential handling, partial-failure handling, and backend submission semantics.
- Collaborate with backend, infrastructure, SRE, security, and datacenter operations teams to deliver reliable GPU fleet observability.
Requirements
- 5 or more years of industry software engineering experience and a Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
- Strong Go development experience for Linux services, command-line interfaces, agents, or daemons.
- Rust experience, or a strong willingness to work in Rust, for out-of-band collector development.
- Strong Linux systems knowledge, including processes, filesystems, networking, service lifecycle, logs, permissions, and host diagnostics.
- Experience with telemetry, observability, health monitoring, or fleet-management systems.
- Experience with Docker, Kubernetes, Helm, and production deployment workflows.
- Experience contributing to open-source projects or working in public repositories with code review, issue tracking, documentation, release notes, and signed commits.
- Familiarity with secure credential handling, enrollment flows, tokens/JWTs, and service-to-service authentication.
- Strong debugging skills across hardware-adjacent software, operating systems, containers, and distributed backend integrations.
Preferred Qualifications
- Experience with NVIDIA datacenter GPUs, DGX systems, DCGM, NVML, CUDA, GPU drivers, XID/SXID events, or GPU diagnostics.
- Experience building host agents, node agents, collectors, monitoring daemons, or Kubernetes DaemonSets.
- Experience with Redfish, BMCs, firmware inventory, secure boot state, PCIe device inventory, or hardware attestation.
- Experience with OpenTelemetry, Prometheus, OTLP gateways, or metric/log export pipelines.
- Experience operating software in AI, HPC, cloud, or large-scale datacenter environments.
- Meaningful open-source contributions in systems software, observability, Kubernetes, Linux, hardware telemetry, Rust, or Go ecosystems.
Compensation and Benefits
- Base salary for Level 3: USD 152,000–241,500 per year.
- Base salary for Level 4: USD 184,000–287,500 per year.
- Equity and benefits are also provided.
- Applications will be accepted at least until August 28, 2026.
- This posting is for an existing vacancy.
- NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Engineering Manager, Data Labeling Platform
Nvidia · Santa Clara, United States
USD 200,000-391,000 per year
Engineering Manager, Local AI Agents
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Deep Learning Software Engineer, DLSim
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Fleet Intelligence Backend
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Business Systems Analyst
Nvidia · Santa Clara, United States
USD 144,000-270,200 per year
Similar jobs
Principal Software Engineer - Rack-Scale Systems Infrastructure
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer - Platform Infrastructure (Rust, C++)
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Distinguished Engineer, Storage – AI Cloud
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior HPC Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year