Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agentic AI @ 6
ClickHouse @ 7
Compliance
Data Pipelines @ 4
Datadog @ 4
Debugging
Distributed Systems
Docker @ 4
GenAI
Generative AI @ 4
Go @ 6
Grafana @ 7
Java @ 6
Kafka @ 4
Kubernetes @ 4
LLM
Machine Learning @ 4
Microservices @ 4
Networking @ 7
Observability @ 4
OpenTelemetry @ 7
Prometheus @ 7
Python @ 6
Security @ 4
VictoriaMetrics @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for a highly skilled Senior Software Engineer to design and develop AIOps and observability platforms at NVIDIA. These platforms are used by internal teams to monitor, diagnose, and optimize products, millions of assets, and services across cloud, on-premises, data centers, supply chain, and edge environments. You will work with engineers, product managers, and partners to define NVIDIA's observability strategy, roadmap, and standard methodologies. You will also mentor and coach engineers on observability, machine learning, tools, and techniques.
Responsibilities
- Lead the design, development, and deployment of AIOps and observability platforms, including metrics, logs, traces, events, alerts, dashboards, and visualizations.
- Drive the technical vision and roadmap for AIOps and observability initiatives, aligning them with business goals and industry best practices.
- Collaborate with teams and customers to understand observability needs and provide solutions that meet their requirements and expectations.
- Establish and implement observability standards, guidelines, and processes across NVIDIA.
- Research, evaluate, and adopt observability technologies and frameworks that can enhance user experience.
- Provide peer reviews, including feedback on performance, scalability, security, and correctness.
- Work with data scientists to implement machine learning models for anomaly detection, forecasting, and root cause analysis on logs, metrics, and events.
- Handle large volumes of data while ensuring data quality, security, and compliance.
- Develop and operate scalable, reliable, distributed systems capable of handling high traffic and complex workloads.
- Develop AI agents and AI-native observability tools that help engineers detect, understand, and resolve production issues faster.
- Build agentic workflows that reason across logs, metrics, traces, events, alerts, topology, and incident history to support anomaly detection, forecasting, root cause analysis, automated debugging, and remediation recommendations.
Requirements
- Bachelor's degree in computer science and engineering, a related field, or equivalent experience.
- 12+ years of experience in product development and full-stack engineering.
- 5+ years of experience developing and operating observability platforms and solutions, preferably in a cloud-native environment.
- Strong knowledge and experience with observability tools such as Prometheus, VictoriaMetrics, Vector, Loki, Grafana, Alertmanager, ClickHouse, and OpenTelemetry.
- Hands-on knowledge of AIOps tools such as BigPanda, PagerDuty, and Datadog.
- Experience with Kubernetes, Nomad, Docker, microservices architectures, and streaming services such as NATS and Kafka for ingesting billions of events.
- Proficiency in one or more programming languages, such as Go, Python, Java, or C#.
- Experience developing observability solutions for on-premises and public cloud environments.
- Experience operating large observability platforms on bare-metal infrastructure.
- Experience establishing scalable data pipelines and instrumentation for collecting, aggregating, and visualizing telemetry and operational metrics.
- Passion for observability and delivering high-quality internal platforms.
Preferred Qualifications
- Deep understanding of implementing observability solutions for large-scale on-premises infrastructure and networking.
- Hands-on experience managing large-scale observability platforms with LLMs and machine learning models.
- Experience building custom services to ingest billions of metrics and logs from a wide range of assets.
- Experience developing unified cloud observability platforms for network, compute, power, storage, operating systems, security, applications, and SaaS platforms.
- Demonstrated experience using machine learning and generative AI to develop predictive monitoring, incident diagnosis, summarization, and correlation solutions.
- Proficiency in AI/ML systems, generative AI, or agentic AI frameworks.
Compensation and Benefits
- Base salary range: USD 200,000–322,000 per year, determined by location, experience, and the pay of employees in similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until August 4, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior System Software Engineer - CPU SoC Boot Firmware
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
System Software Engineer – Data Center Compute Diagnostics
Nvidia · Durham, United States
USD 152,000-241,500 per year
Senior Research Scientist - Generative World Models for Autonomous Driving and Physical AI
Nvidia · Santa Clara, United States
USD 192,000-356,500 per year
Senior Technical Marketing Engineer - CAE Performance
Nvidia · Santa Clara, United States
USD 136,000-253,000 per year
Senior Software Technical Program Driver - OEM and NCP Escalations
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
Similar jobs
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Full-Stack Software Engineer – Verification Data and Visualization Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Staff Software Engineer, Customer Administration
Coinbase · India
INR 9,424,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Principal Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior AI and HPC Observability Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year