Senior Software Architect, Observability Platform

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

SCRAPED

Used Tools & Technologies

Not specified

Required Skills & Competences ?

ElasticSearch @ 4 Grafana @ 4 Prometheus @ 4 DevOps @ 4 Python @ 4 Spark @ 4 API @ 4 GPU @ 4

Details

NVIDIA’s Infrastructure organization is seeking a Senior Software Architect for our Observability Platform to architect and implement distributed observability systems for data centers enabling EDA workflows. You will collaborate with NVIDIA’s HW and SW engineering teams to develop, deploy, and operate observability solutions for multiple CPU and GPU compute clusters around the world. The role focuses on systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to improve efficiency, performance, and productivity of EDA workloads.

Responsibilities

  • Collaborate with hardware and software engineering teams to deliver observability solutions that meet needs in EDA clusters.
  • Develop, test, and deploy data collectors, pipelines, visualization, and retrieval services.
  • Define data collection and retention policies to balance network bandwidth, system load, and storage costs with analysis requirements.
  • Provide operational and strategic data to empower engineers and researchers to improve performance, productivity, and efficiency.
  • Continuously improve quality, workloads, and processes through better observability.

Requirements

  • 8+ years of proven experience.
  • Experience developing large-scale, distributed observability systems.
  • Ability to collaborate with data scientists, researchers, and engineering teams to identify high-value data for collection and analysis.
  • Experience turning raw data into actionable reports.
  • Experience with observability platforms and tooling such as Apache Spark, Elasticsearch / OpenSearch, Grafana, Prometheus, and similar open-source tools.
  • Python programming experience and experience using API calls.
  • MS (preferred) or BS in Computer Science, Electrical Engineering, or related field, or equivalent experience.
  • Excellent planning and interpersonal skills; flexibility and adaptability in dynamic environments.

Preferred / Ways to Stand Out

  • Background in computer science, EDA software, open-source software, infrastructure technologies, and GPU technology.
  • Prior experience in infrastructure software, production application development, release and support methodology, and DevOps.
  • Experience managing datacenters and large-scale distributed computing.
  • Experience working with EDA developers.
  • Track record of driving process improvements, measuring efficiency, and delivering complex projects end-to-end.

Compensation & Benefits

  • Base salary ranges (determined by location, experience, and comparable employees):
    • Level 4: 184,000 USD - 287,500 USD
    • Level 5: 224,000 USD - 356,500 USD
  • Eligible for equity and benefits (link to NVIDIA benefits in original posting).

Other Details

  • Location: Santa Clara, California, United States (Full-time)
  • Application deadline: Applications will be accepted at least until December 13, 2025.
  • NVIDIA is an equal opportunity employer and committed to fostering a diverse work environment.