Senior Observability Architect, AI And HPC

at Nvidia

📍 Santa Clara, United States

USD 224,000-425,500 per year

SENIOR

✅ On-site

SCRAPED

Used Tools & Technologies

Not specified

Required Skills & Competences ^?

Software Development @ 4 Grafana @ 4 Prometheus @ 4 DevOps @ 4 Python @ 4 Spark @ 4 Machine Learning @ 4 Leadership @ 4 Data Analysis @ 4 API @ 4 Technical Leadership @ 4 GPU @ 4

Details

NVIDIA’s Hardware Infrastructure organization is seeking a Senior or Principal Data and Observability Architect. We serve and collaborate directly with NVIDIA’s rapidly growing AI, hardware, and software engineering and research teams across the company. The role involves defining a vision and roadmap for distributed observability systems for large-scale AI and HPC clusters and workloads, and guiding implementation towards this vision.

You will architect systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to significantly improve efficiency, performance, and productivity of AI and HPC workloads. Technical leadership for teams developing, deploying, and operating observability solutions for multiple compute clusters globally is expected.

Responsibilities

Collaborate with AI, hardware, and software engineering and research teams to define a vision and roadmap for AI/HPC cluster observability.
Architect and lead teams to develop, test, and deploy data collectors, pipelines, visualization, and retrieval services.
Define data collection and retention policies balancing network bandwidth, system load, and storage capacity costs with data analysis requirements.
Provide operational and strategic data to empower engineers and researchers to improve performance, productivity, and efficiency.
Continuously improve quality, workloads, and processes through enhanced observability.

Requirements

Experience designing and building large-scale, distributed observability systems.
Ability to collaborate with data scientists, researchers, and engineering teams to identify high-value data for collection and analysis.
Experience transforming raw data into actionable reports.
Experience with observability platforms such as Apache Spark, Elastic/Open Search, Grafana, Prometheus, and similar open-source tools.
Technical lead level Python programming experience and API usage.
Passion for improving the productivity of others.
Excellent planning and interpersonal skills.
Flexibility and adaptability working in a dynamic environment with changing requirements.
MS (preferred) or BS in Computer Science, Electrical Engineering, or related field or equivalent experience.
12+ years of relevant experience.

Ways To Stand Out From The Crowd

Background in computer science, machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology.
Prior experience in infrastructure software, production application software development, software development, release and support methodology, and DevOps.
Experience managing datacenters and large-scale distributed computing.
Experience working with AI researchers and/or EDA developers.
Proven track record of driving process improvements, measuring efficiency, sharing knowledge, and managing complex projects end-to-end.

Compensation & Benefits

The base salary range is 224,000 USD - 425,500 USD. Your salary will be determined based on location, experience, and pay of employees in similar roles. Eligibility for equity and benefits is included.

NVIDIA is committed to fostering diversity and is an equal opportunity employer.