Skip to content
Chatbotscape
The OpenTelemetry conventions, the OpenAI Agents SDK tracing guide and Langfuse's documentation were read on 6 October 2026. The OpenTelemetry GenAI conventions are marked as in development, so attribute names can change.
LLM Observability· Operations
LLM observability is the practice of recording what happens inside each request to a language-model application, including the prompt, the model's response, tool and retrieval steps, token use, latency and cost, so that engineers and owners can explain a given answer and track quality over time.
By Chatbotscape Editorial· Methodology· Published 7 October 2026· Updated 7 October 2026

LLM Observability — Being Able to Explain a Bad Answer After It Happens

Quick answer: A large language model gives different answers to the same question, so you cannot reproduce a customer's bad experience by re-running it. Observability solves this by keeping a structured record of every request: what the bot was asked, which steps it took, what each step cost and how long it took, and what came back. Langfuse describes it as structured logs of every request that capture the exact prompt, the response, token usage, latency and any tool or retrieval steps in between. The record is also personal data in many deployments, so what you keep and for how long is a privacy decision as much as an engineering one.

The vocabulary: traces, spans, metrics, scores

TermWhat it isSupport-bot example
TraceThe full record of one request from start to finishOne customer message and everything the bot did to answer it
Span (or observation)One step inside a trace, with start time, end time, input and outputA model call, a knowledge-base lookup, a tool call, a guardrail check
MetricA number aggregated across many requestsRequests per hour, median delay, tokens per day
ScoreA quality judgment attached to a traceA thumbs-down from the customer, or a grade from an automatic judge
SessionTraces grouped into one conversationAll turns of one chat

The OpenTelemetry project describes three signals for generative AI: traces that follow each interaction's lifecycle, metrics that aggregate things like request volume, latency and token counts, and events that capture granular moments such as individual model interactions. Langfuse's documentation lists application tracing, token and cost tracking, performance metrics, scoring, dashboards and alerts as its core pieces. Different products use different words for the same ideas, so check which one your tool means by "run," "trace" or "session."

An illustrative trace

This is invented and simplified, to show the shape of the data. It is not output from any product.

trace: "Can I get a refund on an annual plan?"        total 4.1 s   $0.011
  span 1  guardrail check (input)                      0.2 s
  span 2  knowledge-base search, 4 passages retrieved   0.6 s
  span 3  model call, 2,300 tokens in, 180 tokens out   3.1 s
  span 4  guardrail check (output)                      0.2 s
  score   customer thumbs-down

With this record a reviewer can open the trace for the thumbs-down, see exactly which four passages the model was given, and tell whether the problem was retrieval (wrong passages), the prompt, or the model. Without it, all you have is the final reply.

How it differs from monitoring and from chatbot analytics

Traditional monitoring asks whether the service is up and fast: error rates, server load, uptime. A bot can pass all of that and still tell customers the wrong refund policy, because the failure is in the content of the answer, not in the infrastructure. Observability adds the content and the reasoning path to the picture.

Chatbot analytics measures outcomes across conversations: containment, resolution, handoff and satisfaction, usually by business owners and on a dashboard. Observability works one level down, at individual requests, usually for whoever fixes the bot. The two share some numbers, such as cost and delay, and complement each other: analytics tells you where the problem is, a trace tells you why.

What the OpenTelemetry conventions standardize

OpenTelemetry, a vendor-neutral observability standard, has a set of generative-AI semantic conventions. Their repository says they cover generative AI operations including inference, agents, tool execution, evaluation and Model Context Protocol calls. The spans document lists recommended attributes such as input tokens, output tokens, reasoning output tokens and cache tokens, and notes that the span should cover the duration of the operation. The point of a standard is that a trace produced by one library can be read by many tools. The spans document is marked as in development, so treat attribute names as subject to change.

The privacy catch

A faithful trace contains the customer's message, and often their name, order details or health or payment context. The OpenTelemetry spans document marks capturing input and output messages as opt-in and warns the attribute is likely to contain sensitive information including user and personal data. The OpenAI Agents SDK has a setting that controls whether the inputs and outputs of model generations and function-tool outputs are stored in its traces, and the same documentation says tracing is unavailable for organizations using OpenAI's APIs under a Zero Data Retention policy. Langfuse documents masking that redacts patterns such as card numbers and email addresses before data leaves your application.

Three consequences follow. Decide what to capture before you switch tracing on. Redact before export, not after. And treat the trace store as covered by your data retention policy, subject to deletion requests such as a data subject access request. We are not lawyers; ask yours which rules apply to your customers.

Common mistakes

  • Logging everything forever. More data means more cost and more personal data to protect.
  • Watching averages only. A median hides the 5% of slow or failed requests customers remember.
  • Treating an automatic quality score as truth. LLM-as-a-judge grades are estimates and need a human spot-check against a golden dataset.
  • Tracing without a question. Decide what decision a number will change before you chart it.

FAQ

What is LLM observability?

Recording each request a language-model app handles, with its steps, prompt, response, tokens, delay and cost, so a team can explain individual answers and track quality.

Is LLM observability the same as monitoring?

No. Monitoring checks that the service is up and fast. Observability also records the content and steps of each answer, which is where most chatbot failures happen.

How is it different from chatbot analytics?

Analytics measures outcomes across conversations, such as resolution and satisfaction. Observability inspects individual requests to find out why an answer went wrong.

Do I need a special tool?

Not necessarily. Several vendors and open-source projects offer one, and OpenTelemetry provides a neutral format, but a small bot can start with the traces its platform already exposes.

Is it safe to log customer messages?

Only with a plan: capture the minimum, redact sensitive patterns, set a retention period and restrict who can read the logs.

Sources