Generative AI observability

Monitor large language model (LLM) performance, track token usage, analyze costs, and help ensure quality of AI-powered services with comprehensive observability for generative AI applications.

Overview

Instana provides comprehensive observability for generative AI applications. As AI becomes integral to applications, it is important to understand how LLMs perform, what they cost, and how they impact user experience. Generative AI observability provides visibility into prompt engineering effectiveness, model response quality, token usage patterns, and the overall health of your AI services.

Key features

Generative AI observability includes the following key features:

  • LLM performance monitoring: Track response times, throughput, and error rates for AI models.
  • Token usage tracking: Monitor token consumption and costs across models and applications.
  • Prompt analysis: Analyze prompt effectiveness and optimize for better results.
  • Model comparison: Compare performance and costs across different LLM providers.
  • Quality metrics: Track response quality, relevance, and user satisfaction.
  • Cost optimization: Identify opportunities to reduce AI costs without compromising quality.

Common use cases

See the following example use cases:

Cost management
Track token usage and costs across your AI applications. Identify expensive operations, optimize prompts to reduce tokens, and set budgets to control spending.
Performance optimization
Monitor LLM response times and identify slow operations. Optimize prompt engineering, caching strategies, and model selection to improve the user experience.
Quality assurance
Track response quality metrics to help ensure that AI outputs meet standards. Detect degradation in model performance or prompt effectiveness over time.
Model selection
Compare different LLM providers and models based on performance, cost, and quality. Make data-driven decisions about which models to use for different use cases.

Where to find it in the UI

You can access generative AI observability from the following locations in the UI:

  • Gen AI observability dashboard: Click Gen AI Observability in the navigation menu to access the main dashboard.
  • Pricing configuration tab: Configure LLM model pricing and view predefined pricing for common models.
  • Analytics view: Click Analyze gen AI calls to analyze calls by application, service, and endpoint with trace details.
  • Application perspectives: Click Analytics in the navigation menu to view perspectives. Then, filter by service name to organize AI applications.
  • Infrastructure view: Access Infrastructure > Analyze Infrastructure to view GPU metrics, vLLM monitors, and vector database metrics.
  • Trace details: Click Gen AI Observability and select Traces tab. View individual AI operations with prompts, responses, token usage, and costs.

What problems does this solve?

Generative AI observability addresses the following common challenges with AI-powered applications:

  • Cost surprises: Understand and control AI costs before they spiral out of control.
  • Performance issues: Identify slow AI operations that impact the user experience.
  • Quality degradation: Detect when model outputs decline in quality.
  • Prompt inefficiency: Optimize prompts to reduce tokens and improve results.
  • Model selection: Make informed decisions about which models to use.

Related capabilities

The following capabilities work together with generative AI observability:

Learn more

For detailed information about Generative AI observability, instrumentation, and best practices, see Generative AI observability documentation.