sample-reviews/*-review.md, every friction figure is quoted from the review that published it, and the two figures that contradict each other inside a single review are printed as a contradiction rather than averaged away.Warm transfer· Customer service operations and handoff design
Warm Transfer — Our Reviews Score the Easy Half, and Nobody Tests the Other One
Quick answer: A warm transfer is a handoff where the receiving person arrives already knowing who the customer is, what they asked, and what has been tried. A cold transfer is the same handoff with none of that. The distinction was invented for the telephone, where context had to be spoken aloud by one human to another in the seconds before a line was released, and it does not survive the move to chat intact: in a shared inbox the transcript is simply there, so the context half is close to free. Our own corpus says so with unusual consistency. Thirteen of our fifteen platform reviews publish a figure for the same handoff scenario, the lowest is 3/5 and the median and the mean are both exactly 4.0 — though five of the thirteen are anchored editorial projections rather than transfers anybody watched, and stripping those five drops the mean to 3.9. No review in the corpus records anything a telephony operator would recognize as a cold transfer. What that consistency hides is the second condition, which is not context but presence — whether a human is there to accept the conversation at all. Agent-availability scheduling appears in four of the fifteen reviews — three of them incidentally, one as a builder feature nobody sent a handoff through — and no review in the corpus tested a handoff that arrived outside staffed hours. The number we publish measures the half that messaging made easy.
Where the term comes from, and what changes in chat
On a phone system, transferring a caller means releasing a line. A cold or blind transfer releases it immediately: the caller hears hold music, then a stranger says "how can I help," and the caller starts over. A warm transfer keeps the first agent on the line long enough to reach the second one, summarize the problem, confirm they will take it, and only then drop off. Three things happen in that pause, and they are worth separating because chat inherits them unevenly:
- Announcement. The receiving agent is told what this is about before speaking to the customer.
- Acceptance. A specific person agrees to take it. Nobody is transferred into a void.
- Continuity. The customer is told what is happening and does not repeat themselves.
Chat makes the first and third nearly automatic. The conversation is already a written record, and if the human works in the same unified inbox the bot writes to, the announcement is just the thread. That is why "warm transfer" is barely spoken in the messaging-platform market: outside this entry and its companion guide the phrase appears in exactly one page of our published catalog — our best voice AI platform list, where it is named as a call-handling capability that dedicated voice vendors ship and our reviewed platforms are not evaluated on — and in zero of our fifteen reviews. It is a voice word.
The second condition is the one chat did not inherit. On a phone, acceptance is structural: you cannot complete a warm transfer to somebody who does not pick up. In chat, the transfer completes the moment the conversation changes state, whether or not a human is looking at the queue. A conversation can be assigned, marked escalated, and sit unread until Monday, and every system involved will report the handoff as successful. That is the failure mode the telephony term was built to prevent, reintroduced by the medium that was supposed to have solved it.
Warm, cold, and the third thing chat invented
| Context carried | Human accepted it | Customer told what happens next | |
|---|---|---|---|
| Cold / blind transfer | No | No | No |
| Warm transfer | Yes | Yes | Yes |
| Queued handoff (the chat default) | Yes | Not yet | Depends entirely on what you wrote |
The middle column is the whole entry. A queued handoff looks warm from the platform's side and can feel colder than a blind phone transfer from the customer's side, because at least the blind transfer failed audibly. Whether a queued handoff behaves as a warm one is decided by two things outside the bot: whether somebody is rostered, and whether you wrote a message that sets an honest expectation. Neither is a feature you can buy.
What our fifteen reviews record
Thirteen of the fifteen reviews carry a trigger-based handoff scenario scored out of five. They do not all publish the same figure, and — as the two sections after the table set out — not all of them ran it:
| Platform | Figure | What the review says it means |
|---|---|---|
| Intercom | 5/5 | Unified inbox surfaces the AI agent's full history, ticket context and all prior actions; trigger latency under two seconds (anchored, not observed) |
| Blip | 4.5/5 | Dedicated Atendimento humano handover block, a first-class primitive rather than a workflow hack — but see the contradiction below (anchored, not observed) |
| Tidio | 4.5/5 | Live-chat heritage; a matched quick reply moved the conversation into Unassigned cleanly |
| AiSensy | 4/5 | Multi-agent inbox, internal notes, role-based access, vendor-claimed unrestricted concurrent devices (anchored, not observed) |
| BotPenguin | 4/5 | Receiving agent saw full history, the escalation reasoning trace and contact metadata |
| Chatbase | 4/5 | Same, into Zendesk via the native integration, with sub-0.7 confidence firing the handover |
| Landbot | 4/5 | Multi-agent live chat with internal notes, conversation routing and role-based access (anchored, not observed) |
| Manychat | 4/5 | "Smooth context transfer to assigned agent," on an Instagram DM trigger |
| Voiceflow | 4/5 | Handover into a Voiceflow inbox with configurable team-role routing |
| Wati | 4/5 | Receiving agent saw full history, the handover reasoning trace and contact metadata |
| Chatfuel | 3.5/5 | Inbox structure confirmed in a walkthrough, but the review states the full handover flow was not exercised (anchored, not observed). Its separate 3/5 handover-friction figure is derived from aggregator sub-ratings |
| Typebot | 3.5/5 | No native handover surface at all: the review rates a Slack webhook it built, after a 14-minute setup |
| Botpress | 3/5 | Context transfer is clean once configured; routing and role-based access are Team-tier setup |
Two reviews are absent, for two different reasons. SendPulse lists the handoff scenario without publishing a figure, its hands-on pass still queued. Tars uses the same scenario slot for a different test entirely, a seven-field conversational form, and labels it "Scenario E" in one place and "Scenario 4" in another.
Thirteen figures: one 3, two 3.5s, seven 4s, two 4.5s, one 5. Median 4.0. Mean 52 ÷ 13 = 4.0 exactly.
They are not all the same measurement, and our reviews do not use one name for them. Five labels appear across the thirteen: "context transfer fidelity" (Intercom, AiSensy, Landbot), "context-transfer friction rating" (BotPenguin, Chatbase, Wati — and Blip, though that phrase labels the ~5/5 we discarded), "context-transfer quality on handover" (Chatfuel), "handover friction rating" (the Blip figure we actually use) and a bare "friction rating" or "friction" (Botpress, Voiceflow, Manychat, Typebot, and Tidio's inbox row). Read literally, a friction rating of 4 out of 5 would describe a transfer that is mostly friction; the prose beneath every one of those scores describes the opposite, so the scale is running as a quality score under a label that says the reverse. That is our defect, not a vendor's, and it is the first reason to treat the 4.0 as a rough central tendency rather than a benchmark. Here is the second.
The uncomfortable part, which is about us
Read the BotPenguin, Chatbase, Wati and Typebot rows again and then read the source sentences side by side:
- BotPenguin: "the receiving agent saw the full conversation history, the AI's reasoning trace on the escalation trigger, and the contact's metadata. Internal notes persisted across the AI→human transition."
- Chatbase: "the receiving Zendesk agent saw the full conversation history, the AI's reasoning trace on the escalation trigger … and the contact's metadata. Internal notes persisted across the AI→human transition."
- Wati: "the receiving agent saw the full conversation history, Astra's reasoning trace on the handover trigger, and the contact's metadata cleanly. Internal notes persisted across the Astra→human transition."
- Typebot: "Slack-side handoff received the full conversation history, AI reasoning trace, and contact metadata cleanly." — a fourth instance, on a platform the same review says ships no native handover surface at all.
Four products with four different architectures — a native inbox, a Zendesk integration, an omnichannel team inbox, and a Slack webhook on a platform with no handover surface — described in one sentence. Whatever else that 4.0 median is, it is not thirteen independent observations. It is a rubric applied thirteen times, and on at least four of those occasions the rubric supplied the prose. We are publishing that here rather than quietly re-scoring, because a reader deciding between BotPenguin and Wati on the strength of a shared 4/5 deserves to know the 4/5 does not distinguish them.
One review also contradicts itself, and the half-point is the least of it. Our Blip review records a "projected context-transfer friction rating: ~5/5" in its scenario section and "the Scenario E handover friction rating 4.5/5" in its walkthrough-takeaways section. The table above uses 4.5, the lower of the two. But that section heading reads "Scenario E — Human handover and team inbox, enterprise contact center 🟠 PROJECTED", so Blip's row is an anchored projection rather than a transfer anybody watched — which is why the table marks it, and Intercom's row with it, (anchored, not observed). Tidio's 4.5 sits beside them unmarked, because that one rests on a session somebody ran. Chatfuel's row has the same problem from the other direction: its own review says the full handover flow was not exercised. And Blip and Chatfuel are not the only two. The Intercom, AiSensy and Landbot reviews each disclose, in their own verification blocks, that six-scenario measurement is scheduled for 2026-06 and that the published figures "derive from structural evaluation anchored against Manychat measured anchor." Five of the thirteen — Intercom 5/5, Blip 4.5/5, AiSensy 4/5, Landbot 4/5 and Chatfuel 3.5/5 — therefore describe no transfer anybody watched, and four of the five sit at or above the median. The eight figures whose reviews claim a session someone ran (Tidio 4.5, BotPenguin 4, Chatbase 4, Manychat 4, Voiceflow 4, Wati 4, Typebot 3.5, Botpress 3) sum to 31.0, a mean of 3.9 against a median still at 4.0, while the five anchored figures average 4.2. The 4.0 is held up by the rows where nothing was observed — and even the eight come with a standing caveat, since several of our reviews describe their own protocol in two incompatible ways. All five are flagged for the manual editorial track rather than corrected in an automated run.
The half nobody scored
Not one of the thirteen figures tells you whether a human was available. Search the corpus for the vocabulary of agent availability — business hours, working hours, operating hours, away messages, availability scheduling — and four files match. Three are incidental; the fourth is a builder feature nobody sent a handoff through:
- Blip: an "operating hours check" and an "attendant availability check" listed among the pre-configured features of a flow template, read from a screenshot caption.
- Tidio: "Operating hours" as a line item in a plan-comparison modal, and present only in that screenshot's alt text.
- SendPulse: a flow-builder canvas with a day-of-week and run-time filter, whose after-hours branch sends two bubbles confirming receipt and promising a reply during working hours. This is the closest thing in the corpus to a designed offline path, and the review marks that builder work
verified-handson— but what was verified is that the branch can be built, not that anybody received a handoff through it. - Chatfuel: an operator asked the vendor's admin copilot to configure an after-hours pattern with timezone-aware working hours; the copilot explained the pattern, showed the config it would write, and then stated it could not complete the step, pointing at a manual panel. A configuration attempt, not a transfer.
That is the entire record this vocabulary surfaces. A wider pattern — the same search with after-hours|after hours|offline|out of office|timezone|time zone appended — matches nine of the fifteen files, and we read every extra hit rather than counting it: broadcast-scheduling timezone options and an India-timezone support-responsiveness theme in AiSensy; an "Account Time Zone" entry in a Manychat account-settings screenshot caption; and the boilerplate "PII redaction was performed offline after capture" session note in AiSensy, BotPenguin, Botpress and Chatbase. None describes agent availability, and none is a test. No review states the time of day at which its handoff scenario ran, and none reports a transfer arriving into an unstaffed inbox — so every figure in the table that rests on a session somebody ran is an in-hours figure, the five that rest on no session carry no hour at all, and the corpus has no measurement whatever of the condition that most often breaks a handoff in a small business. The honest reading of a 4/5 is: when a human was there, the transfer was warm.
How to check it yourself, in three probes
None of this requires a trial account longer than an afternoon.
- The Tuesday probe. During staffed hours, trigger a handoff as a customer. Time the gap between the trigger and a human sentence. That is your first response time for escalations, and it is the number the vendor demo will show you.
- The Saturday probe. Do exactly the same thing at 9pm on a Saturday. What the customer sees in the next sixty seconds is your real handoff design. If it is silence, you have a queued handoff wearing a warm one's score.
- The acceptance probe. Have somebody on your team watch the inbox without touching it. Does the platform notify anybody, escalate after a timeout, or reassign an unread conversation? If the answer is no to all three, assignment is a label rather than an acceptance, and you will need handoff rules and a rota to supply what the phone system used to supply structurally.
Related terms
- Human handoff — the event itself; this entry is about its temperature.
- Chatbot handoff rules — the six trigger families that decide when a transfer fires.
- Chatbot escalation rate — how often it fires, and what a healthy number looks like.
- Unified inbox — the shared surface that makes context transfer nearly free in chat.
- Live chat — the human side of the conversation the transfer lands in.
- Chatbot first response time — the metric the Saturday probe produces.
FAQ
What is a warm transfer?
Handing a customer to a second person who already knows who they are and what they asked, and who has agreed to take it. The opposite, a cold or blind transfer, pushes the customer across with no context and no agreement. The term is telephony's, where the first agent stayed on the line to brief the second before releasing it.
What is the difference between a warm transfer and a cold transfer?
Context and acceptance. A cold transfer carries neither: the customer repeats their problem to a stranger who was not expecting them. A warm transfer carries both. In chat the context half arrives on its own, because the conversation is written down in a shared inbox, which is why the distinction that dominated call-center training translates poorly and why our reviews have never recorded a genuinely cold context transfer.
Do chatbot platforms support warm transfer?
Every review that published a figure scored the context half at 3/5 or better; two published none. Five of those thirteen figures, though, are anchored editorial projections rather than transfers anybody watched; strip them and the mean falls to 3.9. The bottom of the range is instructive rather than damning: Botpress scores 3 because routing and role-based access are Team-tier configuration, not because the transfer loses anything, and Typebot scores 3.5 on a Slack webhook it does not ship, having no native handover surface at all. Absence of a figure is not absence of a feature — SendPulse simply has not had its hands-on pass yet.
Is a warm transfer the same as a warm handoff?
Yes, in practice. "Warm handoff" is the commoner phrasing in chat and in healthcare and social-services settings, "warm transfer" the commoner one in telephony and outbound sales. We use human handoff as the site's canonical term for the event and reserve warm and cold for its quality.
What breaks a warm transfer most often?
Nobody being there. The context, the notes and the metadata are handled by the inbox; the rota is not. A conversation can be assigned, escalated and reported as a successful handoff while sitting unread, which is a failure the phone system could not produce because a warm transfer required somebody to pick up. See the handoff design guide for the coverage and staffing arithmetic.
Does the customer need to be told a transfer is happening?
Yes, and it is the cheapest of the three components to get right. One sentence naming what is happening and how long it should take converts a silence into a wait. Where you cannot promise a number honestly — outside staffed hours, most obviously — promise a channel and a window instead, and then keep it.
Sources
- Chatbotscape review corpus, searched and read 31 August 2026. Denominator:
ls sample-reviews/*-review.md | wc -lreturns 15. The availability search wasgrep -cEie 'business hours|office hours|working hours|operating hours|away message|agent availability|availability schedul' sample-reviews/*-review.md, executed from the repository root, returning non-zero for exactly four files:blip-review.md(1),chatfuel-review.md(2),sendpulse-review.md(1),tidio-review.md(1). The term searchgrep -rlie 'warm transfer' --include=*.md . --exclude-dir=glossary --exclude-dir=academyreturns two paths, which are the same file twice —best/best-voice-ai-platform.mdand its build-time mirror underweb/.content/— and no review; the same search for'cold transfer'returns 0. The two exclusions are this entry and its same-day companion, which necessarily match; a draft of this bullet published the commands without them, and they returned 4 and 2 the moment these two files existed. The wider availability pattern wasgrep -lEie 'business hours|office hours|working hours|operating hours|away message|agent availability|availability schedul|after-hours|after hours|offline|out of office|timezone|time zone' sample-reviews/*-review.md, which returns nine files; the two-wordtime zonealternation is load-bearing — without it the search returns eight and Manychat drops out. - Context-transfer figures, each quoted from the review that published it:
intercom-review.md5/5;blip-review.md4.5/5 in its walkthrough-takeaways section (and "~5/5" in a scenario section headed PROJECTED, printed above as a contradiction);tidio-review.md4.5/5;aisensy-review.md,botpenguin-review.md,chatbase-review.md,landbot-review.md,manychat-review.md,voiceflow-review.md,wati-review.md4/5 each;chatfuel-review.md3.5/5 ("Context-transfer quality on handover", distinct from the 3/5 "Handover friction" figure on the same line, which is aggregator-derived — an earlier draft of this table published the 3/5 under a context-transfer heading and was wrong);typebot-review.md3.5/5, published in that review's test-summary row as "14 min setup; 3.5/5 friction" — a draft of this entry excluded Typebot on the claim that it published no fidelity score, which is false;botpress-review.md3/5. The arithmetic: one 3, two 3.5s, seven 4s, two 4.5s and one 5 sum to 52 over 13 rows, so the median and mean are both 4.0.sendpulse-review.mdnames the scenario without a figure andtars-review.mduses the scenario slot for a seven-field conversational form, labeling it two different ways in the same file. - The four near-identical scenario sentences are quoted verbatim above from
botpenguin-review.md,chatbase-review.md,wati-review.mdandtypebot-review.md— the fourth on a platform the same review says ships no native handover surface at all. The observation that a shared rubric, rather than four independent readings, produced them is ours, and it is a criticism of our own corpus rather than of the products. - Availability observations, each quoted from the review that recorded it:
blip-review.md(operating-hours and attendant-availability checks listed as pre-configured features of a flow template, in a screenshot caption);tidio-review.md("Operating hours" as a plan-comparison line item, present only in the screenshot's alt text);sendpulse-review.md(flow-builder canvas with a day-of-week and run-time filter, after-hours branch sending two confirmation bubbles);chatfuel-review.md(admin copilot explaining an after-hours pattern, surfacing the config, and stating it could not complete the step). The claim that no review tested an out-of-hours handoff is an argument from the absence of after-hours vocabulary in the scenario sections of all fifteen files, not from a positive statement in any of them, and is stated that way in the body. - Evidence class of each figure, read from each review's own disclosure rather than inferred:
intercom-review.md:82("Six-scenario actual measurement scheduled 2026-06; figures derive from structural evaluation anchored against Manychat measured anchor"),aisensy-review.md:85(PROJECTED_SCENARIO_TESTING_RENDERED, and :823 "roughly seven hours of structured analysis … combining vendor-page walkthroughs"),landbot-review.md:83("six-scenario actual measurement scheduled 2026-06; projected scores published as interim"),blip-review.md:419(Scenario E heading carries 🟠 PROJECTED) andchatfuel-review.md:759(full handover flow not exercised) are anchored rather than observed. The arithmetic on the split: anchored five sum to 21.0 for a mean of 4.2; the other eight sum to 31.0 for a mean of 3.875, printed as 3.9; 21.0 + 31.0 = 52.0, which is the whole-set total above. An earlier draft of this entry said two of the thirteen were unobserved, which understated it by a factor of two and a half, in the flattering direction. - Corpus defects flagged for the manual editorial track, not corrected in this automated run: the Blip 4.5-versus-5 discrepancy above and the PROJECTED marker that sits above the higher of the two; the Tars "Scenario E" / "Scenario 4" double labeling, which the 2026-08-31 run also recorded; the five different names the corpus gives one scenario's score, and the "friction" label carrying a quality scale; and ru/en mixed-edit corruption in
blip-review.mdat lines 397 and 421 ("on and 20-query test set", "on and 5-user enterprise account"), found while checking the Blip quotations. - Ahrefs Keywords Explorer, US overview and volume-by-country, queried 31 August 2026 — the demand, difficulty, CPC, global-volume, parent-topic and country-split figures in this entry's keyword note, including the checks behind declining 'service level agreement', 'queue management' and 'erlang c'.
- Chatbotscape evaluation methodology, including the six-scenario protocol the friction figures come from and the standing caveat about how much of it is measurement. /methodology (continuously updated).