Senior Software Engineer, AIOps and Observability

at Nvidia
USD 200,000-322,000 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Agentic AI @ 6 ClickHouse @ 7 Compliance Data Pipelines @ 4 Datadog @ 4 Debugging Distributed Systems Docker @ 4 GenAI Generative AI @ 4 Go @ 6 Grafana @ 7 Java @ 6 Kafka @ 4 Kubernetes @ 4 LLM Machine Learning @ 4 Microservices @ 4 Networking @ 7 Observability @ 4 OpenTelemetry @ 7 Prometheus @ 7 Python @ 6 Security @ 4 VictoriaMetrics @ 7

Details

We are looking for a highly skilled Senior Software Engineer to design and develop AIOps and observability platforms at NVIDIA. These platforms are used by internal teams to monitor, diagnose, and optimize products, millions of assets, and services across cloud, on-premises, data centers, supply chain, and edge environments. You will work with engineers, product managers, and partners to define NVIDIA's observability strategy, roadmap, and standard methodologies. You will also mentor and coach engineers on observability, machine learning, tools, and techniques.

Responsibilities

  • Lead the design, development, and deployment of AIOps and observability platforms, including metrics, logs, traces, events, alerts, dashboards, and visualizations.
  • Drive the technical vision and roadmap for AIOps and observability initiatives, aligning them with business goals and industry best practices.
  • Collaborate with teams and customers to understand observability needs and provide solutions that meet their requirements and expectations.
  • Establish and implement observability standards, guidelines, and processes across NVIDIA.
  • Research, evaluate, and adopt observability technologies and frameworks that can enhance user experience.
  • Provide peer reviews, including feedback on performance, scalability, security, and correctness.
  • Work with data scientists to implement machine learning models for anomaly detection, forecasting, and root cause analysis on logs, metrics, and events.
  • Handle large volumes of data while ensuring data quality, security, and compliance.
  • Develop and operate scalable, reliable, distributed systems capable of handling high traffic and complex workloads.
  • Develop AI agents and AI-native observability tools that help engineers detect, understand, and resolve production issues faster.
  • Build agentic workflows that reason across logs, metrics, traces, events, alerts, topology, and incident history to support anomaly detection, forecasting, root cause analysis, automated debugging, and remediation recommendations.

Requirements

  • Bachelor's degree in computer science and engineering, a related field, or equivalent experience.
  • 12+ years of experience in product development and full-stack engineering.
  • 5+ years of experience developing and operating observability platforms and solutions, preferably in a cloud-native environment.
  • Strong knowledge and experience with observability tools such as Prometheus, VictoriaMetrics, Vector, Loki, Grafana, Alertmanager, ClickHouse, and OpenTelemetry.
  • Hands-on knowledge of AIOps tools such as BigPanda, PagerDuty, and Datadog.
  • Experience with Kubernetes, Nomad, Docker, microservices architectures, and streaming services such as NATS and Kafka for ingesting billions of events.
  • Proficiency in one or more programming languages, such as Go, Python, Java, or C#.
  • Experience developing observability solutions for on-premises and public cloud environments.
  • Experience operating large observability platforms on bare-metal infrastructure.
  • Experience establishing scalable data pipelines and instrumentation for collecting, aggregating, and visualizing telemetry and operational metrics.
  • Passion for observability and delivering high-quality internal platforms.

Preferred Qualifications

  • Deep understanding of implementing observability solutions for large-scale on-premises infrastructure and networking.
  • Hands-on experience managing large-scale observability platforms with LLMs and machine learning models.
  • Experience building custom services to ingest billions of metrics and logs from a wide range of assets.
  • Experience developing unified cloud observability platforms for network, compute, power, storage, operating systems, security, applications, and SaaS platforms.
  • Demonstrated experience using machine learning and generative AI to develop predictive monitoring, incident diagnosis, summarization, and correlation solutions.
  • Proficiency in AI/ML systems, generative AI, or agentic AI frameworks.

Compensation and Benefits

  • Base salary range: USD 200,000–322,000 per year, determined by location, experience, and the pay of employees in similar positions.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until August 4, 2026.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs