Used Tools & Technologies
Not specified
Required Skills & Competences ?
Software Development @ 4 Grafana @ 4 Prometheus @ 4 DevOps @ 4 Python @ 4 Spark @ 4 Machine Learning @ 4 Data Analysis @ 4 API @ 4 GPU @ 4Details
NVIDIA’s AI Infrastructure organization is seeking a Senior AI Observability Engineer to help architect and implement distributed observability systems for AI and HPC clusters. You will work with a team of engineers on systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to improve efficiency, performance, and productivity of AI and HPC workloads. You will develop, deploy, and operate observability solutions for multiple compute clusters around the world.
Responsibilities
- Collaborate with AI, HW, SW engineering and research teams to deliver observability solutions that meet their needs in AI/HPC clusters.
- Develop, test, and deploy data collectors, pipelines, visualization and retrieval services.
- Build a self-serve platform for observability.
- Define data collection and retention policies to balance network bandwidth, system load, and storage capacity costs with data analysis requirements.
- Provide operational and strategic data to empower engineers and researchers to improve performance, productivity, and efficiency.
- Continuously improve quality, workloads, and processes through better observability.
Requirements
- Experience developing large-scale, distributed observability systems.
- Ability to collaborate with data scientists, researchers, and engineering teams to identify high-value data for collection and analysis.
- Experience turning raw data into actionable reports.
- Experience with observability platforms such as Apache Spark, Elastic/OpenSearch, Grafana, Prometheus, and other similar open-source tools.
- Python programming experience and use of API calls.
- Passion for improving the productivity of others.
- Excellent planning and interpersonal skills.
- Flexibility/adaptability working in a dynamic environment with changing requirements.
- MS (preferred) or BS in Computer Science, Electrical Engineering, or related field (or equivalent experience).
- 8+ years of proven experience.
Ways to Stand Out
- Practical experience in machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology.
- Prior experience in infrastructure software, production application software development, release and support methodology, and DevOps.
- Experience in the management of datacenters and large-scale distributed computing.
- Experience working with AI researchers and/or EDA developers.
- Consistent track record of driving process improvements and measuring efficiency and a passion for sharing knowledge and driving complex projects end-to-end.
Compensation & Benefits
- Base salary range (by level):
- Level 4: 184,000 USD - 287,500 USD
- Level 5: 224,000 USD - 356,500 USD
- You will also be eligible for equity and benefits.
Application & Other Information
- Applications for this job will be accepted at least until July 29, 2025.
- NVIDIA is an equal opportunity employer and is committed to fostering a diverse work environment.