Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 4
Data Pipelines @ 7
Go @ 3
Grafana @ 4
Kubernetes @ 4
LLM @ 7
Machine Learning
Python @ 3
SRE @ 4
Statistics @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a passionate AI Tools Engineer to join the Site Reliability Engineering (SRE) Data Team. The role focuses on building and deploying AI-powered tools and products that support the operation and optimization of the critical, global GeForce NOW service. The tools will transform production data streams, including signals, metrics, and logs, into actionable intelligence for automated incident root-cause analysis and service trend prediction.
Responsibilities
- Build and implement robust AI/ML tools that analyze production data to identify root causes for complex incidents and predict future operational trends.
- Lead the development of new LLM- and agent-based systems to improve operational efficiency.
- Establish and maintain data management practices, including workflows for converting and handling large-scale data sources used in model development.
- Own and improve LLM-based pipelines while incorporating current LLM developments into product development.
- Serve as an authority on AI frameworks and recommend platforms, toolsets, and architectural approaches that support the product's long-term technical sustainability.
- Support the advancement of Site Reliability Engineering across production environments.
Requirements
- Bachelor's degree in Computer Science, Statistics, Engineering, or equivalent experience.
- 5+ years of experience.
- Strong proficiency in Python; familiarity with Go or other systems languages is a plus.
- Practical experience building, optimizing, and deploying AI tools.
- Strong knowledge of the AI landscape and current developments, including how LLM-based platforms are built and optimized and how to select appropriate platforms.
- Hands-on experience with Kubernetes and cloud environments, including AWS.
- Active engagement with developments in AI and the ability to distinguish meaningful advances from noise when making technical decisions.
- Expertise in automation and large-scale data pipelines.
- Experience with monitoring and visualization tools such as Grafana.
- Strong ability to handle, transform, and manage data sources and pipelines.
Preferred Qualifications
- Current experience with LLM improvement pipelines and a strong understanding of recent developments in LLM training.
- Understanding of SRE concepts and experience managing production environments.
- Experience with Kubernetes, AWS, and other cloud technologies.
- Excellent knowledge of LLMs and AI models, with the ability to recommend sustainable long-term technical approaches and avoid unsuitable platform choices.
- Proficiency in automation.
Benefits
- Equity eligibility.
- NVIDIA benefits package.
- Competitive salary package.
- Inclusive work environment and equal employment opportunity.
Applications will be accepted at least until August 31, 2026. This posting is for an existing vacancy.
More jobs at Nvidia
Senior Systems Software Engineer, Low Latency Streaming Technology - Automotive
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Reinforcement Learning Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Manager, Storage Engineering
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Senior Software and System Architect
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Customer Technical Program Manager
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Similar jobs
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AIOps and Observability
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Infrastructure Security Engineer
SpaceXAI · Washington, United States, Austin, United States, New York City, United States, Palo Alto, United States
USD 100,000-258,000 per year