Senior AI and HPC Observability Engineer

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Data Engineering @ 7 Data Modeling @ 7 Data Pipelines @ 4 Debugging @ 7 Distributed Systems @ 6 Flink @ 4 GPU @ 4 Go @ 7 HPC @ 4 Java @ 7 Kafka @ 4 Kubernetes @ 4 Machine Learning @ 4 Observability @ 4 OpenTelemetry @ 6 Prometheus @ 6 Python @ 7 Spark @ 4

Details

NVIDIA's Managed AI Superclusters (MARS) team builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop next-generation AI/ML systems. The team is seeking an AI and HPC Observability Engineer to build and scale observability and telemetry platforms for advanced computing workloads. You will design and develop high-throughput, reliable telemetry pipelines and modern data infrastructure, applying distributed systems fundamentals, production-grade coding, and operational excellence.

Responsibilities

  • Design and scale observability platforms handling high-volume metrics, logs, and traces across distributed environments.
  • Build high-performance backend services for telemetry ingestion, processing, and routing.
  • Develop and extend OpenTelemetry collectors, processors, exporters, and instrumentation libraries.
  • Build and optimize metrics pipelines using large-scale time-series storage systems.
  • Design and operate real-time and batch telemetry pipelines using streaming and distributed data technologies.
  • Improve platform reliability, performance, and cost efficiency through tuning, capacity planning, and system optimization.
  • Develop monitoring, alerting, and service reliability frameworks to ensure platform health and performance.
  • Collaborate with platform engineering, infrastructure, and site reliability teams to deliver production-grade observability solutions.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • At least 5 years of experience building backend or distributed systems in production environments.
  • Strong programming skills in Python, Go, or Java, with experience developing production-quality software.
  • Hands-on experience with modern observability architectures, including metrics, logs, and traces.
  • Solid experience with PromQL and time-series data systems.
  • Experience building or operating distributed data pipelines using technologies such as Kafka, Spark, or Flink.
  • Experience working with Kubernetes and cloud-native infrastructure.
  • Strong understanding of distributed systems, concurrency, and fault-tolerant system design.
  • Strong debugging, performance tuning, and production operations skills.

Preferred Qualifications

  • Experience designing and scaling observability platforms for AI, GPU, or HPC environments.
  • Hands-on expertise with OpenTelemetry, Prometheus, Kafka, and high-volume distributed telemetry pipelines.
  • Strong background in data engineering, time-series data modeling, and real-time performance tuning.
  • Experience integrating observability with AI/ML pipelines, GPU workload monitoring, or intelligent alerting.
  • Experience using statistical or machine learning techniques for anomaly detection, correlation, or predictive insights.

Compensation and Benefits

  • Base salary range for Level 3: $152,000–$241,500 USD per year.
  • Base salary range for Level 4: $184,000–$287,500 USD per year.
  • Eligible for equity and benefits.
  • NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.

More jobs at Nvidia

Similar jobs