Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agentic AI @ 4
CI/CD @ 4
Datadog @ 4
Debugging @ 6
DevOps @ 6
Distributed Systems @ 6
Docker @ 4
Grafana @ 4
IaC
Kubernetes @ 4
LLM @ 4
LangChain @ 4
Machine Learning @ 4
Observability @ 4
OpenTelemetry @ 4
Profiling @ 3
Prometheus @ 4
Python @ 7
SRE @ 6
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Develop the future of AI at NVIDIA by joining the BizApps SRE team and advancing the Agentic AI Factory model. This initiative focuses on the fast development, deployment, and operation of AI-powered applications across the enterprise. The role sits at the intersection of agentic AI and business applications, ensuring that AI agents and automation workflows perform reliably, efficiently, and at scale.
Responsibilities
- Build and implement observability solutions for production agentic AI applications, including metrics, tracing, logging, and alerting.
- Develop and maintain deployment pipelines and reliability tools for the Agentic AI Factory model.
- Partner with AI application teams to define SLOs and SLIs and ensure production readiness for new agent deployments.
- Instrument LLM-based workflows for performance, cost, and quality monitoring, including multi-step reasoning chains, token usage, and tool orchestration.
- Drive incident response, root cause analysis, and reliability improvements for BizApps AI services.
- Identify and address observability gaps in agentic AI systems, including hallucination detection and orchestration failure tracing.
- Work with platform, data, and application engineering groups to integrate reliability throughout the development lifecycle.
Requirements
- Bachelor's or master's degree in Computer Science, Software Engineering, or a related field, or equivalent experience.
- At least 8 years of experience and strong software engineering skills in Python.
- Experience with modern CI/CD practices.
- Hands-on experience with container orchestration using Kubernetes and Docker, as well as cloud infrastructure.
- Experience with observability and monitoring platforms such as Datadog, OpenTelemetry, Grafana, or Prometheus.
- An SRE or DevOps approach, including ownership of production systems, automation of toil, and reliability-focused engineering.
- Excellent problem-solving skills and a proven track record of debugging complex distributed systems.
Preferred Qualifications
- Experience with LLM orchestration frameworks such as LangChain, LlamaIndex, or Semantic Kernel, and with agentic AI development patterns.
- Experience operating machine learning or AI systems in production environments.
- Familiarity with AI-specific observability challenges, including inference latency profiling, token economics, and timely response-quality monitoring.
- Contributions to open-source observability or AI tooling projects.
- Experience with infrastructure as code using Terraform or Pulumi, and with GitOps or equivalent workflows.
Compensation and Benefits
- Base salary range: $168,000–$270,250 USD per year, determined by location, experience, and pay for similar positions.
- Eligible for equity and benefits.
- Applications will be accepted at least until September 12, 2026.
- NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.
More jobs at Nvidia
Senior Technical Program Manager – DFX Engineering
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Firmware Verification and Bringup Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer - Platform Software
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior AI Workflow Engineer
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Manager, Software Engineering Productivity and Release Engineering
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Similar jobs
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AIOps and Observability
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior Manager, Software Engineering - Agentic IT Operations
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Cloud Site Reliability Engineer (SRE) - Data Management & Analytics Platform
Bloomberg · Princeton, United States
USD 160,000-240,000 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · Palo Alto, United States, San Francisco, United States
USD 220,000-405,000 per year
Senior Architect, Agentic AI for Marketing
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year