
How to Monitor an AI Chatbot in Production
Traces, Alerts and a Weekly Review for Small Teams
Quick answer: Do five things. Decide which few questions your monitoring must answer. Set what gets stored, and redact personal data before it leaves your system. Turn on tracing so each customer message has a record of its steps, tokens, delay and cost. Add alerts for a handful of signals, not dozens. And every week read about 20 real conversations, including the ones customers rated badly, because dashboards show that something changed while only the conversation shows what went wrong. A small bot can start with the traces its platform already exposes; you do not need to buy a dedicated tool on day one.
Step 1: Write down the questions monitoring must answer
Tools offer hundreds of charts. Pick questions first, then record only what answers them. For a small support bot we suggest five:
- Is the bot up and answering quickly?
- Is it costing what we expected per conversation?
- Is it giving wrong or unsafe answers?
- Is it handing off when it should?
- Did something change after our last edit, or after the model provider's last update?
Each question maps to a signal, shown in the table. The thresholds are Chatbotscape's own starting points for a small bot; set yours from a week of normal traffic.
| Question | Signal to record | Example alert (starting point) |
|---|---|---|
| Up and quick? | Error count, median and slowest-5% response time | Errors above 2% of requests for 15 minutes |
| Cost as expected? | Tokens in and out per request, cost per conversation | Daily cost 50% above the seven-day average |
| Wrong or unsafe? | Thumbs-down rate, guardrail triggers, reviewer grades | Thumbs-down rate doubles day over day |
| Handing off? | Escalation rate, fallback rate | Fallback rate above your usual by a set margin |
| Changed? | Same signals, split by prompt version and model name | Any signal moves after a deploy |
Step 2: Decide what to store, and redact before export
A complete record of a customer chat holds personal data. The OpenTelemetry GenAI conventions mark capturing input and output messages as opt-in and say the attribute is likely to contain sensitive information including user and personal data. Treat that as a design instruction.
- Minimum by default. Start by recording the shape of each request (steps, tokens, timing, cost, versions) and capture message text only where you need it for review.
- Redact before it leaves your system. Langfuse's documentation describes masking that redacts patterns such as card numbers, email addresses and phone numbers from trace inputs and outputs before they are exported. Whatever tool you use, check where redaction happens; one applied after storage does not undo the exposure.
- Know your platform's switches. The OpenAI Agents SDK has a setting, on by default, that stores the inputs and outputs of model generations and function-tool outputs in traces, and its documentation says tracing is unavailable for organizations using OpenAI's APIs under a Zero Data Retention policy. If you rely on zero retention, find out what your monitoring path does instead of assuming.
- Set a retention period and access list. Trace data belongs under your data retention policy, and deletion requests must reach it. Our chatbot data privacy guide and PII handling guide cover the wider rules. We are not lawyers; confirm the requirements for your customers with yours.
Step 3: Turn on tracing
Tracing records each request as a trace made of steps. In the OpenAI Agents SDK, for example, each run is wrapped in a trace, model calls get generation spans, function-tool calls get function spans and guardrails get guardrail spans, per its documentation. Other frameworks use similar structures. What you want for a support bot:
- One trace per customer message, with the conversation ID so you can see the whole chat.
- A step for every model call, knowledge-base lookup, tool call and guardrail check, each with its own timing.
- Token counts and cost on every model step. The OpenTelemetry spans document lists recommended attributes for input, output, reasoning and cache tokens, which is useful if you want a tool-neutral format.
- Version labels: the prompt version, the knowledge-base version and the model name on every trace. Without them you cannot answer "did something change?"
- A place for scores: the customer's thumbs rating and, later, reviewer grades.
If your no-code platform exposes only conversation transcripts and a few totals, that is a limit worth knowing before you buy; ask the vendor what a single answer's record contains. If you build your own flow, an open standard such as OpenTelemetry keeps you from being locked to one viewer, with the caveat that its GenAI conventions are marked as in development and may change.
Step 4: Add a few alerts, then prune them
Alerts exist to wake someone for a reason. Start with the table in Step 1 and apply three rules, which are our own practice:
- Alert on change, not on a fixed number, wherever you can. A bot's normal cost and delay vary by hour and season.
- Send each alert to one named person with a one-line note on what to check first, such as "open the slowest three traces from the last hour."
- Delete any alert that fired three times with no action taken. Alert fatigue is how real problems get ignored.
Quality alerts are harder than uptime alerts. A customer thumbs-down is cheap and honest but sparse. An automatic grader, the LLM-as-a-judge approach, can score every trace, but its grades are estimates; use them to decide which conversations a human reads first, not as the verdict. Running a judge on every message also costs tokens, so many teams sample. Sampling 10% of traces for automatic grading and reading 100% of thumbs-down traces is one reasonable split; adjust it to your volume and budget.
Step 5: Run a 20-conversation weekly review (our method)
This is Chatbotscape's own procedure, not a published standard. It takes about an hour.
- Pull 20 traces: all thumbs-down conversations up to 10, the 5 most expensive, the 5 slowest, and the rest at random.
- For each, answer three questions: Was the answer correct? If not, was the problem the retrieved passages, the prompt or the model? Should it have been handed to a person?
- Tag the cause in a shared sheet with one of: missing knowledge, wrong retrieval, prompt gap, model error, tool failure, should-have-escalated.
- Fix the top cause first. If "missing knowledge" leads, update the knowledge base; see build a chatbot knowledge base.
- Turn confirmed failures into tests. Add each fixed question and its correct answer to your golden dataset so the regression tests catch it if it returns.
Tracing and the weekly review work as a loop: traces find the conversation, the review explains it, the test keeps it fixed. For a calendar of other recurring chores, see the chatbot maintenance cadence.
Common mistakes
Watching uptime only. The bot can answer instantly and wrongly. Include content signals.
Storing every message in plain text. It creates a data-protection duty you did not need, and a breach target.
Alerting on everything. Ten ignored alerts are worse than three trusted ones.
Skipping version labels. After a model update or prompt edit, you will not be able to tell which one moved the numbers.
Never reading a conversation. Metrics do not show tone, a confusing reply or a missed handoff.
FAQ
How do I monitor an AI chatbot in production?
Record each request as a trace with its steps, tokens, delay and cost, redact personal data before storage, alert on a few change signals, and read a sample of real conversations weekly.
What should I log for an LLM chatbot?
Request steps, timings, token counts, cost, prompt and model versions, customer ratings and guardrail triggers. Store message text only where you need it, and redact sensitive patterns.
Do I need a dedicated observability tool?
Not at the start. Use the traces your platform exposes, and add a tool when you need to search, compare versions or grade at scale.
How often should I review conversations?
We suggest weekly for a new bot and monthly once the failure causes stop changing, adjusting for volume and risk.
Is monitoring the same as testing?
No. Testing checks known questions before release; monitoring watches real traffic after it. Use both.
Related guides
- Chatbot metrics guide — choose the numbers worth watching.
- Conversational analytics guide — the outcome-level view above traces.
- Chatbot regression testing guide — keeping fixed problems fixed.
- Reduce chatbot costs — acting on what the token data shows.
- LLM observability and chatbot analytics — the glossary entries behind the terms used here.
Sources
- OpenTelemetry, GenAI spans — github.com/open-telemetry/semantic-conventions-genai, read 6 October 2026: development status, opt-in message capture, sensitive-data warning, token attributes.
- OpenTelemetry, Semantic conventions for generative AI (README) — github.com/open-telemetry/semantic-conventions-genai, read 6 October 2026: scope.
- OpenAI, Agents SDK tracing — openai.github.io/openai-agents-python/tracing, read 6 October 2026: default spans, sensitive-data setting, Zero Data Retention limitation.
- Langfuse, LLM observability overview — langfuse.com/docs/observability/overview, read 6 October 2026: tracing, cost tracking, scores, dashboards, alerts.
- Langfuse, Masking — langfuse.com/docs/observability/features/masking, read 6 October 2026: redaction before export.
- Chatbotscape evaluation methodology. /methodology (continuously updated).
About this guide
Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is editorial guidance built from vendor and standards documentation; the five-step method, the weekly review and all thresholds are our own practice, not a published standard. The method does not depend on any platform. To flag an error, write to editorial@chatbotscape.com.
Last updated
7 October 2026 — first published.