Principal Data Platform Architect

at Nvidia
USD 272,000-425,500 per year
SENIOR
✅ On-site

SCRAPED

Used Tools & Technologies

Not specified

Required Skills & Competences ?

Grafana @ 4 Prometheus @ 4 DevOps @ 4 Python @ 4 Spark @ 4 Java @ 4 Machine Learning @ 4 Leadership @ 7 JavaScript @ 4 Data Analysis @ 4 GPU @ 4

Details

NVIDIA’s Hardware Infrastructure organization is seeking a Principal Data Platform Architect to define the vision and roadmap for a distributed data platform and observability systems supporting large-scale AI and HPC clusters. You will architect systems for data collection, aggregation, enrichment, storage, retrieval, and visualization, and lead technical teams to develop, deploy, and operate observability solutions for multiple compute clusters around the world.

Responsibilities

  • Collaborate with AI, hardware, and software engineering and research teams to define a vision and roadmap for AI/HPC cluster observability.
  • Architect and lead teams to develop, test, and deploy data collectors, pipelines, visualization, and retrieval services.
  • Define data collection and retention policies to balance network bandwidth, system load, and storage capacity costs with data analysis requirements.
  • Provide operational and strategic data to empower engineers and researchers to improve performance, productivity, and efficiency.
  • Continuously improve quality, workloads, and processes through better observability.

Requirements

  • Extensive experience designing and building large-scale, distributed observability systems.
  • Ability to collaborate with data scientists, researchers, and engineering teams to identify high-value data for collection and analysis.
  • Experience turning raw data into actionable reports and dashboards.
  • Experience with observability platforms and tooling such as Apache Spark, Elastic/OpenSearch, Grafana, Prometheus, and similar open-source tools.
  • Technical-lead level programming experience in Python, JavaScript, and Java.
  • Thorough understanding of databases (relational and non-relational).
  • Strong planning, interpersonal, and leadership skills; passion for improving team productivity and sharing knowledge.
  • Flexibility and adaptability working in a dynamic environment with changing requirements.
  • MS (preferred) or BS in Computer Science, Electrical Engineering, or related field, or equivalent experience.
  • 15+ years of relevant experience.

Ways to Stand Out

  • Background in computer science, machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology.
  • Prior experience in infrastructure software, production application development, release and support methodology, and DevOps.
  • Experience managing datacenters and large-scale distributed computing.
  • Experience working with AI researchers and/or EDA developers.
  • Demonstrated track record of driving process improvements, measuring efficiency, and delivering complex projects end-to-end.

Compensation & Benefits

  • Base salary range: 272,000 USD - 425,500 USD (determined based on location, experience, and pay of employees in similar positions).
  • Eligible for equity and company benefits.

Additional Details

  • Applications accepted at least until July 29, 2025.
  • NVIDIA is an equal opportunity employer committed to diversity and inclusion.