For the complete documentation index, see llms.txt. Markdown versions are available by appending .md to documentation URLs.

AI Observability - Monitor LLM Cost, Tokens, Latency

SigNoz Cloud - This page applies to SigNoz Cloud editions.
Self-Host - This page applies to self-hosted SigNoz editions.

AI Observability is a section in SigNoz for LLM and agent workloads. It reads the OpenTelemetry GenAI attributes on your spans and shows cost, token usage, latency, errors, time to first token (TTFT), and tool calls.

Use this page to open AI Observability, read the Overview dashboard, and make sure that your spans carry the attributes that the dashboard needs.

Prerequisites

  • An LLM or agent application that sends traces to SigNoz with OpenTelemetry GenAI attributes. For setup steps for OpenAI, Anthropic, LiteLLM, CrewAI, and other frameworks, see LLM Observability.
  • For cost panels, a pricing rule for each model. SigNoz Cloud includes default prices for common models. In open-source SigNoz, add the price for each of your models. To add or change a price, see Model pricing.
  • Self-hosted SigNoz with a custom collector configuration: add the signozspanmapper and signozllmpricing processors to the traces pipeline. See Add the AI observability processors.

Open AI Observability

  1. In the side navigation, open the More menu.
  2. Select AI Observability.

SigNoz opens the Overview tab. The tab bar at the top has four tabs:

  • Overview: the dashboard that this page describes.
  • Explorer: query LLM and AI trace data. See AI Observability Explorer.
  • Model pricing: set the price for each model. See Model pricing.
  • Attribute Mapping: map attributes from other instrumentation libraries to the GenAI names. See Attribute Mapping.

The Overview dashboard

The Overview tab shows the AI Observability Overview dashboard. SigNoz creates and updates this dashboard for you, so you cannot edit it. To start an alert from a panel, select Create Alerts in the menu of that panel.

AI Observability Overview tab with the model, provider, environment, and service_name filters, six stat cards, and the Cost & tokens panels
The AI Observability Overview dashboard

A row of stat cards is at the top:

CardWhat it shows
Total costThe total LLM cost, in USD.
Total tokensInput tokens plus output tokens.
LLM callsThe number of spans that have gen_ai.request.model.
LLM latency (p95)The p95 duration of LLM calls.
LLM error rateThe percentage of LLM calls with an error status.
TTFT (p95)The p95 time to first token, in seconds.

Below the cards, the panels are in four sections.

Cost & tokens

PanelWhat it shows
Cost over timeLLM cost by model.
Token usage over timeInput, output, cache-read, and cache-write tokens.
Cost breakdownCost split into input, output, cache-read, and cache-write token cost.
Cost per LLM callThe average cost of one LLM call, by model.
Prompt cache hit ratioCache-read tokens as a share of all input tokens.
Cost by serviceLLM cost and call count for each service.
Cost by providerLLM cost and call count for each provider.

Per-trace usage

An AI trace is a trace that has at least one LLM, tool, or agent span.

PanelWhat it shows
AI traces over timeThe number of AI traces in each interval.
Calls per AI traceThe average number of LLM calls and tool calls in a trace.
Cost per AI traceThe average and p95 of the total LLM cost of a trace.
Tokens per AI traceThe average and p95 of the total tokens of a trace.

LLM latency & errors

PanelWhat it shows
LLM latency by modelThe p50, p90, p95, and p99 latency and the call count for each model.
LLM latency over timeThe p95 latency by model.
Slowest LLM call per traceThe p50 and p95 of the longest LLM call in a trace.
Time to first tokenThe p95 TTFT by model.
LLM errors over timeLLM calls with an error status, by model.
Error rate by modelErrored calls, total calls, and the error percentage for each model.
Finish reasonsLLM calls by finish reason. If the share of length or max_tokens goes up, more responses are cut off.

Tool calls

Tool calls section of the AI Observability Overview with Tool call rate, Tool error rate, Tool duration, Top tools, and Top span names panels
The Tool calls section of the Overview dashboard
PanelWhat it shows
Tool call rateTool calls per second, by tool.
Tool error rateThe percentage of tool calls with an error status.
Tool durationThe p50 and p95 duration of tool calls.
Top toolsThe most called tools, with the error count and the median duration.
Top span namesThe most frequent GenAI span names.

Filter the dashboard

The variable bar at the top of the dashboard has four filters. Each filter lets you select more than one value, or All.

FilterAttribute
modelgen_ai.request.model
providergen_ai.provider.name
environmentdeployment.environment (resource attribute)
service_nameservice.name (resource attribute)

The model and provider filters apply to the LLM and per-trace panels only. Tool spans do not carry a model, so the Tool calls panels and Calls per AI trace use only the environment and service_name filters.

Required attributes

The dashboard reads these attributes from your spans. A panel stays empty until your spans carry the attributes that it uses. Most names come from the OpenTelemetry GenAI semantic conventions.

AttributeUsed by
gen_ai.request.modelAll LLM panels and the model filter. SigNoz counts a span as an LLM call when it has this attribute.
gen_ai.provider.nameCost by provider and the provider filter.
gen_ai.usage.input_tokens, gen_ai.usage.output_tokensToken panels and cost.
gen_ai.usage.cache_read.input_tokens, gen_ai.usage.cache_creation.input_tokensCache token panels, cache cost, and Prompt cache hit ratio.
gen_ai.server.ttftTTFT (p95) and Time to first token. Streaming instrumentations, such as OpenLIT, send this attribute in seconds.
gen_ai.response.finish_reasonsFinish reasons.
gen_ai.tool.nameAll Tool calls panels. SigNoz counts a span as a tool call when it has this attribute.
gen_ai.agent.nameMarks agent spans for the Per-trace usage panels.
deployment.environment, service.nameThe environment and service_name filters, and Cost by service.

Error panels use the error status of the span, not an attribute.

You do not send the cost attributes. SigNoz calculates cost from the token counts and the model price when it receives a span, and it writes the result to signoz.gen_ai.usage.tokens.cost and to one cost attribute for each token type.

If your instrumentation uses other attribute names, map them to these names in Attribute Mapping.

Limitations

  • You cannot edit or clone the Overview dashboard. To change a panel, build it in your own dashboard in Dashboards.
  • The earlier in-product /llm-observability/* routes no longer exist. Change saved links to /ai-observability/*.

Get Help

If you need help with the steps in this topic, please reach out to us on SigNoz Community Slack. If you are a SigNoz Cloud user, please use in product chat support located at the bottom right corner of your SigNoz instance or contact us at cloud-support@signoz.io.

Is this page helpful

Last updated—September 29, 2026

Edit on GitHub