Senior AI Tools Engineer, SRE Operations - GeForce NOW

at Nvidia
USD 144,000-230,000 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 AWS @ 4 Data Pipelines @ 7 Go @ 3 Grafana @ 4 Kubernetes @ 4 LLM @ 7 Machine Learning Python @ 3 SRE @ 4 Statistics @ 4

Details

NVIDIA is seeking a passionate AI Tools Engineer to join the Site Reliability Engineering (SRE) Data Team. The role focuses on building and deploying AI-powered tools and products that support the operation and optimization of the critical, global GeForce NOW service. The tools will transform production data streams, including signals, metrics, and logs, into actionable intelligence for automated incident root-cause analysis and service trend prediction.

Responsibilities

  • Build and implement robust AI/ML tools that analyze production data to identify root causes for complex incidents and predict future operational trends.
  • Lead the development of new LLM- and agent-based systems to improve operational efficiency.
  • Establish and maintain data management practices, including workflows for converting and handling large-scale data sources used in model development.
  • Own and improve LLM-based pipelines while incorporating current LLM developments into product development.
  • Serve as an authority on AI frameworks and recommend platforms, toolsets, and architectural approaches that support the product's long-term technical sustainability.
  • Support the advancement of Site Reliability Engineering across production environments.

Requirements

  • Bachelor's degree in Computer Science, Statistics, Engineering, or equivalent experience.
  • 5+ years of experience.
  • Strong proficiency in Python; familiarity with Go or other systems languages is a plus.
  • Practical experience building, optimizing, and deploying AI tools.
  • Strong knowledge of the AI landscape and current developments, including how LLM-based platforms are built and optimized and how to select appropriate platforms.
  • Hands-on experience with Kubernetes and cloud environments, including AWS.
  • Active engagement with developments in AI and the ability to distinguish meaningful advances from noise when making technical decisions.
  • Expertise in automation and large-scale data pipelines.
  • Experience with monitoring and visualization tools such as Grafana.
  • Strong ability to handle, transform, and manage data sources and pipelines.

Preferred Qualifications

  • Current experience with LLM improvement pipelines and a strong understanding of recent developments in LLM training.
  • Understanding of SRE concepts and experience managing production environments.
  • Experience with Kubernetes, AWS, and other cloud technologies.
  • Excellent knowledge of LLMs and AI models, with the ability to recommend sustainable long-term technical approaches and avoid unsuitable platform choices.
  • Proficiency in automation.

Benefits

  • Equity eligibility.
  • NVIDIA benefits package.
  • Competitive salary package.
  • Inclusive work environment and equal employment opportunity.

Applications will be accepted at least until August 31, 2026. This posting is for an existing vacancy.

More jobs at Nvidia

Similar jobs