Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
API @ 4
AWS @ 4
Data Modeling @ 7
Debugging @ 7
Distributed Systems @ 4
Docker @ 4
GPU
Grafana @ 4
Helm @ 4
JWT @ 3
Kafka
Kubernetes @ 4
Linux @ 4
Observability @ 6
OpenTelemetry @ 4
PostgreSQL @ 7
Prometheus @ 4
SRE
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Software Engineer to join the DGX Cloud / Fleet Intelligence team and build backend systems that power GPU health monitoring, telemetry ingestion, operational automation, and fleet visibility for NVIDIA’s high-performance GPU infrastructure.
Responsibilities
- Design and develop Go backend services, REST APIs, and data models for GPU Health and Fleet Intelligence.
- Build customer-facing and agent-facing APIs for compute zones, node groups, nodes, alerts, events, metrics, inventory, attestation, retention policies, and reports.
- Develop high-volume ingestion and persistence paths for in-band agents and out-of-band collectors.
- Work with Aurora PostgreSQL, Kafka/MSK, S3, SQS, Prometheus/AMP, and OpenTelemetry-based observability.
- Build and operate scheduled backend services for liveness tracking, alerting, rollups, cleanup, notification delivery, attestation, and XID analysis.
- Optimize database schemas, partitioned time-series storage, query performance, CTE-heavy queries, and connection pooling for reliable service behavior.
- Collaborate with agent, infrastructure, SRE, UI, and cloud operations teams to turn operational workflows into scalable backend systems.
- Improve service reliability, security, observability, testing, and deployment quality across Docker, Kubernetes, Helm, and cloud environments.
Requirements
- At least 5 years of industry software engineering experience with a Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or equivalent experience.
- Strong Go backend development experience.
- Experience building REST APIs and production services using frameworks such as Gin or similar.
- Strong PostgreSQL experience, including schema design, query tuning, transactions, migrations, and operational data modeling.
- Experience with distributed systems, event pipelines, telemetry ingestion, or operational analytics.
- Experience with cloud infrastructure, Docker, Kubernetes, and Helm.
- Strong debugging skills across APIs, databases, cloud services, and production systems.
- Familiarity with authentication and authorization patterns such as JWT, service account keys, trusted-edge proxies, or customer-scoped APIs.
- Experience with Linux-based development and production environments.
Preferred Qualifications
- Background with telemetry, monitoring, observability, health, or fleet-management platforms.
- Experience with AWS services such as MSK, Aurora, S3, SQS, AMP, or CloudWatch.
- Experience with OpenTelemetry, Prometheus, LightStep, Grafana, or production tracing and metrics systems.
- Background with NVIDIA GPUs, DGX systems, DCGM, XID analysis, attestation, or AI datacenter operations.
- Experience designing APIs and storage systems that support real-time operational workflows at fleet scale.
Compensation and Benefits
The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. The position is also eligible for equity and benefits.
Applications will be accepted at least until August 28, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
More jobs at Nvidia
Engineering Manager, Data Labeling Platform
Nvidia · Santa Clara, United States
USD 200,000-391,000 per year
Engineering Manager, Local AI Agents
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Deep Learning Software Engineer, DLSim
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Business Systems Analyst
Nvidia · Santa Clara, United States
USD 144,000-270,200 per year
Senior Automation and Tools Development Engineer
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Similar jobs
Senior Software Engineer, Fleet Intelligence Agent Systems
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Staff Backend Software Engineer, Agent Platform
SentinelOne · United States
USD 156,000-215,000 per year
Software Engineer - Platform Infrastructure (Rust, C++)
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Senior Engineer System Software, SDN Operations
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year