Principal Observability Architect, AI And Hpc
at Nvidia
📍 Santa Clara, United States
$272,000-471,500 per year
SCRAPED
Used Tools & Technologies
Not specified
Required Skills & Competences ?
Grafana @ 4 Prometheus @ 4 Python @ 4 Spark @ 4 Data Analysis @ 4 API @ 4Details
NVIDIA’s Hardware Infrastructure organization is seeking a Senior or Principal Data and Observability Architect. We serve and collaborate directly with NVIDIA’s rapidly growing AI, HW, and SW engineering and research teams across the company. We are looking for a technical leader to define a vision and roadmap for distributed observability systems for large-scale AI and HPC clusters and workloads and guide implementation towards this vision. You will architect systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to spectacularly improve efficiency, performance, and productivity of AI and HPC workloads. You will lead technical teams to develop, deploy, and operate observability solutions for multiple compute clusters around the world.
Responsibilities
- Collaborate with AI, HW, and SW engineering and research teams to define a vision and roadmap for AI/HPC cluster observability.
- Architect and lead teams to develop, test, and deploy data collectors, pipelines, visualization and retrieval services.
- Define data collection and retention policies to balance network bandwidth, system load, and storage capacity costs with data analysis requirements.
- Work in a diverse team to provide operational and strategic data to empower our engineers and researchers to improve performance, productivity, and efficiency.
- Continuously improve quality, workloads, and processes through better observability.
Requirements
- Experience designing and building large scale, distributed observability systems.
- Ability to collaborate with data scientists, researchers, and engineering teams to identify high value data for collection and analysis.
- Experience with turning raw data into actionable reports.
- Experience with observability platforms such as Apache Spark, Elastic/Open Search, Grafana, Prometheus, and other similar open-source tools.
- Technical lead level Python programming experience and use of API calls.
- Passion for improving the productivity of others.
- Excellent planning and interpersonal skills.
- Flexibility/adaptability working in a dynamic environment with changing requirements.
- MS (preferred) or BS in Computer Science, Electrical Engineering, or related field or equivalent experience.
- 12+ years of relevant experience.
Benefits
You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.