Senior System Software Engineer - Data Platform Observability

at Nvidia
USD 184,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI API @ 4 Ansible @ 4 Audit Compliance @ 6 Go @ 6 Grafana @ 4 Helm @ 4 Java @ 6 JavaScript @ 6 Kubernetes @ 4 Microservices @ 4 Observability @ 4 OpenTelemetry @ 4 Prometheus @ 4 Python @ 6 React @ 6 Rust @ 6 Software Development @ 7 Spark @ 4 Terraform @ 4

Details

NVIDIA’s Hardware Infrastructure organization is seeking a Senior System Software Engineer to lead the evolution of its next-generation Data and Observability Platform. The platform serves NVIDIA’s AI, hardware, software engineering, and research teams. This role is for a full-stack technical lead who can work deeply with infrastructure and serve as the technical anchor for the observability stack. You will build a centralized platform used by thousands of NVIDIA engineers to visualize chip telemetry, debug distributed pipelines, and ensure platform reliability.

Responsibilities

  • Architect high-performance ingestion and build centralized telemetry pipelines capable of handling massive scale.
  • Solve global latency challenges by implementing modern, push-based edge collection architectures to replace legacy proxy models.
  • Design and implement infrastructure for data governance, policy engines, access control enforcement points, secure credential management, and audit logging.
  • Develop modern web interfaces and APIs that unify distinct observability signals into a consolidated user experience.
  • Implement cost-effective tiered storage architectures, routing high-volume data to cold storage while maintaining multi-year data retention.
  • Architect workflow orchestration systems to automate platform maintenance, data lifecycle management, and complex pipeline operations.
  • Provide operational and strategic data to help engineers and researchers improve performance, productivity, efficiency, quality, workloads, and processes through better observability.

Requirements

  • BS or MS in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 8+ years of full-stack software development experience focused on data platforms or infrastructure tools.
  • Proficiency in high-performance backend systems programming and modern frontend web frameworks, including Python, JavaScript, Java, Rust, Go, React, or similar technologies.
  • Experience with observability platforms such as Apache Spark, Elastic/OpenSearch, Grafana, Prometheus, and similar open-source tools.
  • Hands-on experience operating and extending the Grafana ecosystem or ELK stack at scale.
  • Understanding of time-series database and inverted index internals.
  • Experience deploying complex stateful services on Kubernetes using Helm, Terraform, or Ansible.
  • Familiarity with event streaming and modern data lake formats.

Preferred Qualifications

  • Experience writing custom Grafana data source or backend plugins in Go.
  • Experience migrating legacy monoliths to microservices or Vector-based pipelines.
  • Experience configuring OpenTelemetry collectors, writing custom processors, or developing instrumentation SDKs.
  • Background in data governance, including implementing Policy-as-Code or compliance frameworks in regulated environments.

Compensation and Benefits

  • Base salary: USD 184,000–287,500 per year, determined by location, experience, and compensation for similar positions.
  • Eligible for equity and benefits.
  • NVIDIA is an equal opportunity employer committed to a diverse work environment.

More jobs at Nvidia

Similar jobs