Senior Software Engineer, Agentic AI and Observability

at Nvidia
USD 168,000-270,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Agentic AI @ 4 CI/CD @ 4 Datadog @ 4 Debugging @ 6 DevOps @ 6 Distributed Systems @ 6 Docker @ 4 Grafana @ 4 IaC Kubernetes @ 4 LLM @ 4 LangChain @ 4 Machine Learning @ 4 Observability @ 4 OpenTelemetry @ 4 Profiling @ 3 Prometheus @ 4 Python @ 7 SRE @ 6 Terraform @ 4

Details

Develop the future of AI at NVIDIA by joining the BizApps SRE team and advancing the Agentic AI Factory model. This initiative focuses on the fast development, deployment, and operation of AI-powered applications across the enterprise. The role sits at the intersection of agentic AI and business applications, ensuring that AI agents and automation workflows perform reliably, efficiently, and at scale.

Responsibilities

  • Build and implement observability solutions for production agentic AI applications, including metrics, tracing, logging, and alerting.
  • Develop and maintain deployment pipelines and reliability tools for the Agentic AI Factory model.
  • Partner with AI application teams to define SLOs and SLIs and ensure production readiness for new agent deployments.
  • Instrument LLM-based workflows for performance, cost, and quality monitoring, including multi-step reasoning chains, token usage, and tool orchestration.
  • Drive incident response, root cause analysis, and reliability improvements for BizApps AI services.
  • Identify and address observability gaps in agentic AI systems, including hallucination detection and orchestration failure tracing.
  • Work with platform, data, and application engineering groups to integrate reliability throughout the development lifecycle.

Requirements

  • Bachelor's or master's degree in Computer Science, Software Engineering, or a related field, or equivalent experience.
  • At least 8 years of experience and strong software engineering skills in Python.
  • Experience with modern CI/CD practices.
  • Hands-on experience with container orchestration using Kubernetes and Docker, as well as cloud infrastructure.
  • Experience with observability and monitoring platforms such as Datadog, OpenTelemetry, Grafana, or Prometheus.
  • An SRE or DevOps approach, including ownership of production systems, automation of toil, and reliability-focused engineering.
  • Excellent problem-solving skills and a proven track record of debugging complex distributed systems.

Preferred Qualifications

  • Experience with LLM orchestration frameworks such as LangChain, LlamaIndex, or Semantic Kernel, and with agentic AI development patterns.
  • Experience operating machine learning or AI systems in production environments.
  • Familiarity with AI-specific observability challenges, including inference latency profiling, token economics, and timely response-quality monitoring.
  • Contributions to open-source observability or AI tooling projects.
  • Experience with infrastructure as code using Terraform or Pulumi, and with GitOps or equivalent workflows.

Compensation and Benefits

  • Base salary range: $168,000–$270,250 USD per year, determined by location, experience, and pay for similar positions.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until September 12, 2026.
  • NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.

More jobs at Nvidia

Similar jobs