
Chatbot Vendor Evaluation Checklist
What to Ask Before You Buy, Trial, or Sign
Quick answer: A sales demo shows you the platform on its best day, on curated data, run by the person paid to close you. Before you sign, ask about pricing mechanics rather than the sticker price (what counts as a "resolution" or a "conversation," and what happens the month you go over), ask what happens to your data and who can see it, ask how the vendor measures its own AI's accuracy and what a wrong answer costs you, and run a pilot on your own messiest real conversations rather than the demo script. Our own review corpus is proof this matters even when we are the ones doing the checking: a hands-on Manychat test found a channel that silently stopped working with zero warning anywhere in the UI, a Botpenguin score didn't match between the review's own prose and its own scoring table, and Gorgias's help center contradicts its own blog on whether an AI-resolved ticket costs you twice. If claims like that can slip past a reviewer paid to check them, they can slip past a sales call too.
Why the demo isn't the product: three things we caught, checking
We publish fifteen hands-on platform reviews, and even under our own methodology — reading vendor documentation, testing live accounts, checking claims against a second source — three real gaps got through the first pass and were only caught on a closer read. They're worth knowing about specifically, because each one is a category of trap a buyer meeting a sales rep for forty-five minutes has no way to catch.
A feature can fail silently, with the UI still saying it works. Our Manychat review documents a Telegram bot that sat idle for about four months, during which the platform-side webhook registration drifted and stopped delivering messages — with, in the review's own words, "no warning, no banner, and no Help Center documentation," and "the channel toggle still reads 'Enabled.'" The operator's own dashboard told them everything was fine. If you are choosing a platform partly on "how will I know if something breaks," the honest answer for at least one platform, on at least one channel, was: you won't, unless you test it yourself on a schedule. Ask any vendor directly what monitoring exists for a channel that goes quiet, and don't accept "you'd notice" as the answer.
A vendor's own numbers can disagree with each other. Our Botpenguin review scores the platform's AI/NLU dimension at 77/100 in its scoring table and frontmatter, and at 76/100 in the same review's own prose two paragraphs later — a minor, honest inconsistency in our own work, caught on review rather than before publishing, and useful precisely because it shows how easily a single-digit number drifts between where it's calculated and where it's written down. A vendor's pricing page, sales deck and support documentation are three separate documents maintained by three different teams; expecting them to agree with each other by default is optimistic.
A vendor's own documentation can contradict itself on what you'll actually be billed. Our Gorgias alternatives research found that Gorgias's help center states plainly that a ticket AI Agent resolves without a human touch is billed both a ticket fee and a separate automation fee under current pricing — while the same vendor's own plan-guide blog post says AI-resolved conversations "don't count as billable tickets." Same company, two live pages, opposite claims, and the help center (the billing authority) is the one that actually applies to your invoice. Nobody on a sales call is going to volunteer that their own blog and their own help center disagree; you find that only by reading both yourself, on the same day, for the plan you're actually about to buy.
None of these were malicious. All three were the kind of thing that only surfaces when someone reads the primary documentation closely instead of trusting the summary — which is exactly what a 45-minute demo doesn't give you time to do.
Ask about pricing mechanics before you ask about price
The number on the pricing page is the least useful number in the whole conversation, because it depends entirely on definitions the page doesn't spell out. Ask, specifically:
- What counts as one billable unit — a "conversation," a "resolution," a "contact," a "ticket," an "AI Agent interaction" — and does that unit change between your trial tier and the plan you'd actually buy? The Gorgias example above is what happens when this isn't asked: "ticket" and "resolved interaction" are two different meters that can both fire on the same conversation.
- What happens the month you go over, not the month you're under. Overage rates, whether the platform auto-recharges without asking, and whether there's a hard cap or just a bigger bill.
- Whether the AI-specific costs are bundled or separate from the base subscription. Some platforms price LLM inference into the plan; others meter it per resolution or per token on top. Ask which, in writing, for the plan tier you'd sign for — not the plan tier in the demo.
- What the free trial or free tier actually requires. Across our fifteen reviews, 8 platforms explicitly document that their free tier or trial requires no credit card at all. A platform confident enough in its product to let you test it without a card first is telling you something different from one that wants your payment method before you've seen anything real.
Ask about data, security and accuracy — then verify, don't just ask
You don't need to become a compliance officer to ask the right four questions; you need to ask them of every finalist and compare the answers side by side, in writing, not from memory of a call.
- Where does our data live, and who can see it — is content used to train the vendor's models for other customers, or kept isolated to your account? Get this in writing; a sales rep's spoken assurance is not a data processing agreement.
- What happens to our data if we cancel — a specific retention window and deletion process, not "we handle that."
- How does the AI decide it doesn't know something, versus guessing? Ask for the platform's own measured hallucination or citation-accuracy rate if they publish one, and ask what a wrong answer costs you operationally (a refunded order, an angry customer, a compliance problem) — not what it costs in tokens.
- What's the actual SLA, in writing, for uptime and support response, and what do you get if it's missed? A verbal "we're very reliable" is not an SLA.
Compliance and security depth beyond this — SOC 2, HIPAA, GDPR data-subject requests, a full four-layer security audit — is a bigger job than a buying checklist should try to compress; that's what our security checklist and GDPR compliance guide are for, and the honest move is to open those before you sign anything regulated.
Run a pilot on your own worst conversations, not the demo script
A demo is scripted by the seller. A pilot should be scripted by you, using your own messiest real transcripts — the ones with typos, ambiguous intent, an angry customer, or a question your team argues about internally. Before you commit budget:
- Bring 15–20 real questions your customers actually ask, including a few you know are hard, and run them against the trial account yourself rather than watching someone else run them.
- Grade each answer on the same three-way scale every time: correct, wrong-but-confident, or an honest "I don't know." A platform that says "I don't know" instead of guessing is safer than one that never admits uncertainty — measure for that directly, don't infer it.
- Test the failure path, not just the happy path. Ask something the bot has no business answering (return policy for a product you don't sell) and see whether it invents an answer or declines. Then test the handoff: does a human actually get notified, and how fast, on the plan tier you'd buy — trial tiers sometimes throttle or omit handoff notifications that the paid tier includes.
- Time yourself building one real flow, not the one in the vendor's own tutorial. Our reviews measure exactly this (a working skeleton in under a minute on some platforms, several minutes on others) precisely because vendor-scripted demos compress the parts that take real users longest.
- Ask what breaks the moment you exceed the trial's limits — contact caps, message caps, seat caps — and whether hitting them mid-pilot changes what you're evaluating.
Red flags worth stopping on
The sales rep can't answer a specific pricing-mechanics question and offers to "follow up." Sometimes true and fine once. A pattern of it, on questions that should be in their own documentation, is the product of a pricing model complicated enough that the people selling it don't fully understand it either.
The trial account behaves noticeably better than every other review or forum thread you can find of the paid tier. Not proof of anything by itself, but worth asking directly: what's different between the trial and what you'd actually run in production.
Nobody will put the SLA, data retention terms, or overage rate in writing before you sign. A legitimate vendor can send you these in an email in the same day. A pattern of "let's discuss that after the contract" is the tell, not the individual instance.
The demo only shows you conversations that go well. Ask to see — or better, to run yourself — what happens when the bot doesn't know something. A vendor unwilling to show a failure mode live is a vendor who hasn't tested their own failure modes either.
FAQ
What questions should I ask a chatbot vendor before signing?
What counts as one billable unit and what happens on overage; whether AI costs are bundled or metered separately; where your data lives and what happens to it if you cancel; the platform's own measured accuracy or hallucination rate if published; and the actual uptime and support SLA in writing. Get answers to all five in writing, not from a call.
How long should a chatbot pilot or trial run?
Long enough to test 15–20 of your own real, hard questions yourself, exercise the failure path (a question the bot should decline) and the handoff path, and hit at least one of the trial's stated limits — usually more informative than a fixed number of days.
What are common red flags when choosing a chatbot vendor?
A sales rep who can't answer specific pricing-mechanics questions from their own documentation; refusal to put SLA, data retention, or overage terms in writing before signing; a demo that never shows what happens when the bot doesn't know something; and vendor documentation (help center vs. blog vs. pricing page) that disagrees with itself on the same fact — verify the fee structure yourself rather than trusting a single page.
Do I need a lawyer to evaluate a chatbot vendor?
Not for the checklist in this guide — asking pricing-mechanics and data-handling questions and running a pilot needs no legal training. Compliance-specific review (a data processing agreement, HIPAA, GDPR data-subject request handling) is a different question worth a specialist's read; see our GDPR compliance guide for what that guide covers and what it explicitly leaves to counsel.
Can I trust a chatbot platform's own case studies and reviews?
Read them, but verify the specific numbers that matter to your decision against the vendor's own primary documentation rather than a marketing page — our own review corpus has caught real disagreements between a vendor's help center and its own blog (see the Gorgias example above), and between a review's own prose and its own scoring table. If it can happen in content built specifically to be checked, it can happen in a page built specifically to sell.
What our own reviews recorded, checked while drafting this
Reproduced 27 September 2026 across the fifteen files matched by sample-reviews/*-review.md: free trial appears in 11 of 15; sales call or book a demo in 6 of 15; and no credit card (in some phrasing) in 8 of 15 — Botpenguin, Botpress, Intercom, SendPulse, Tars, Tidio, Voiceflow and Wati. That last number is the one worth carrying into your own shortlist: nearly half of what we've reviewed lets you run a real pilot before your card is on file at all, which removes "we're already paying, might as well continue" as a reason to skip the pilot step above.
Related guides
- Chatbot security checklist — the four-layer audit to run once you've picked a finalist.
- Chatbot security and PII handling — the data-lifecycle depth this guide's data questions point toward.
- GDPR chatbot compliance guide — where the compliance-specific questions get real depth.
- Chatbot free plans compared 2026 — the side-by-side price sheet this guide assumes you'll consult.
- Chatbot launch checklist — what happens after you've bought and are getting ready to go live.
- Migrating vendors in 14 days — the playbook if this checklist runs late, after you've already signed the wrong contract.
- Chatbot QA testing protocol — the ten-test pass to run once the platform is actually yours, building on the pilot questions above.
- Best AI chatbot — platforms ranked under our published methodology, as a starting shortlist.
Sources
- Chatbotscape, Manychat review — /reviews/manychat-review, read 27 September 2026: the Telegram webhook-desync finding (line 157: "no warning, no banner, and no Help Center documentation," "the channel toggle still reads 'Enabled'"), hands-on section added 27 May 2026 (line 868).
- Chatbotscape, BotPenguin review — /reviews/botpenguin-review, read 27 September 2026: the AI/NLU dimension scored 77/100 in the scoring table and frontmatter (line 20, line 337) versus 76/100 in the review's own prose (line 206).
- Chatbotscape, Gorgias alternatives — /alternatives/gorgias, read 27 September 2026: the help-center-vs-blog billing contradiction on AI-resolved tickets, sourced there to docs.gorgias.com "How you're billed for using Gorgias" and gorgias.com/blog/gorgias-plan, both read 26 September 2026.
- Chatbotscape review corpus, searched 27 September 2026 from the repository root. Denominator:
ls sample-reviews/*-review.md | wc -lreturns 15.grep -liE "free trial" sample-reviews/*-review.mdreturns 11 files;grep -liE "sales call|book a demo" sample-reviews/*-review.mdreturns 6;grep -io -m1 "no credit card[a-z ]*"matches in 8: botpenguin-review.md, botpress-review.md, intercom-review.md, sendpulse-review.md, tars-review.md, tidio-review.md, voiceflow-review.md, wati-review.md. - Chatbotscape evaluation methodology. /methodology (continuously updated).
About this guide
Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is part of our SMB chatbot Academy. It is editorial guidance built from patterns in our own review corpus, not a compliance program or legal advice; where a question needs specialist depth (data processing agreements, HIPAA, GDPR data-subject requests) we say so and point to the guide that covers it. Manychat, Intercom and BotPenguin have Chatbotscape reviews that carry affiliate links; this guide's checklist does not depend on which platform a reader ultimately chooses. To flag an error, write to editorial@chatbotscape.com.
Methodology
The three case-study findings were each re-read directly from their source review files on 27 September 2026 rather than summarized from memory; the corpus-search counts in the section above are reproducible with the commands printed in Sources. The pricing-mechanics, data-question and pilot-design recommendations are editorial working guidance derived from patterns across our published reviews and from the trap categories those three findings represent; they are not the output of a formal buyer survey.
Last updated
28 September 2026 — first published.