Skip to content
Chatbotscape
Regulation text and our own review corpus re-read 24 August 2026
Data Retention Policy· Privacy and data governance
A data retention policy is the written rule that states, for each category of data you hold, how long you keep it, what justifies that period, and what happens when the period ends — deletion, irreversible anonymization, or archival. For a chatbot the categories that matter are usually the message transcript, the fields extracted from it, the derived analytics, and whatever copy the platform and the model provider keep on their own side.
By Chatbotscape Editorial· Methodology· Published 25 August 2026· Updated 25 August 2026

Data Retention Policy — The Number GDPR Will Not Give You, and the Traffic Level Below Which Deleting Early Blinds Your Bot

Quick answer: A data retention policy is a table, not a paragraph. It lists each category of data you hold, the period you keep it for, the reason that period is the right one, and the end state. Three things on this page are worth more than the definition. The first is that the GDPR names no retention period at all: not thirty days, not six months, not two years. Article 5(1)(e) requires that data be kept "in a form which permits identification of data subjects for no longer than is necessary for the purposes for which the personal data are processed," and most of the specific figures in circulation trace back somewhere else — sector rules, tax and accounting law, or somebody's template. Sector regimes elsewhere do name numbers; the GDPR is not one of them. What it obliges you to do is disclose either the period or the criteria used to set it, which is a documentation duty rather than a numeric one. The second is an audit of our own coverage: across the fifteen platform reviews we have published, three record anything at all about retention, and no two record it in the same form: one an enterprise-tier control, one a control at the top tier plus a disclosed period at the tier below it, one a vendor-set period identical on every tier. The third is arithmetic, and it runs against the instinct most operators bring to this subject: on our own illustrative assumptions, a 30-day transcript window makes transcript-level diagnosis of a specific failure impractical below roughly 14 conversations a day. The smaller the bot, the longer the window it needs before its raw logs can teach it anything.

The number that is not there

Search for a GDPR retention period and you will be given one. Thirty days, ninety days, six months, two years, seven years — each stated with the confidence of a citation and almost never carrying one.

The regulation contains none of them. Article 5(1)(e) sets out the storage limitation principle this way: personal data shall be "kept in a form which permits identification of data subjects for no longer than is necessary for the purposes for which the personal data are processed." The provision then continues with a single carve-out, and it is one this page relies on later: data "may be stored for longer periods insofar as the personal data will be processed solely for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1)," subject to appropriate safeguards. What the article does not contain is a number. There is no annex of periods, no table by data category, no default.

The omission looks intentional, and Recital 39 is the closest the text comes to saying why: it asks that "the period for which the personal data are stored is limited to a strict minimum" and that "time limits should be established by the controller for erasure or for a periodic review." That is an instruction to decide, not a period. A period that is correct for a support transcript is wrong for an invoice and absurd for a job application, and rather than legislate thousands of cases badly the regulation makes the period your decision and then constrains how you make it. Three constraints do the real work:

You have to have decided. Article 5(2) makes the controller responsible for demonstrating compliance with 5(1), which includes storage limitation. "We have never deleted anything" is not a period.

You have to tell people. Article 13(2)(a) requires that, when you collect data from someone, you inform them of "the period for which the personal data will be stored, or if that is not possible, the criteria used to determine that period." That second branch is the practical one and it is widely missed: you are allowed to publish criteria instead of a number, but you are not allowed to publish nothing.

You have to be able to act on it. Article 17 gives a right to erasure on any of six grounds, and the first of them, 17(1)(a), applies where data "are no longer necessary in relation to the purposes […]" — subject to the exemptions in 17(3), which include legal obligations, legal claims and the archiving and statistical purposes mentioned above. A retention schedule you cannot execute is a schedule that generates obligations you will fail.

The practical reading for an operator: the question is not "what period is legal." It is "what period can I write down, justify in a sentence, publish, and actually carry out." That is a shorter list than it looks.

What a chatbot policy has to cover

The mistake that produces unusable policies is writing one period for "chat data." A single conversation splits into at least four categories with genuinely different answers, and they live in different systems.

CategoryWhat it is in a chatbotWhat usually drives the periodEnd state
Raw transcriptThe message text, both sides, in orderQuality review, dispute evidence, bot improvementDelete
Extracted fieldsName, email, order number pulled out by entity extraction and written to a recordThe business purpose the field was collected for — a quote, an order, a ticketDelete or move to the system of record
Derived analyticsIntent label, confidence score, outcome, timestamp, fallback flagTrend measurement, which needs years, not weeksAggregate, and if genuinely non-identifying, keep
Third-party copiesWhat the platform keeps, and what the model provider logsTheir terms, not yoursAsk, then contract for it

That fourth row is the one operators discover late. Your policy governs your systems; the chat platform and, if your bot is generative, the model provider hold their own copies under their own terms. Our PII handling guide works through where those copies sit; the point for a retention policy is narrower. If you promise a customer a ninety-day retention period and your platform holds transcripts for a year, the promise is false and you made it in writing.

The third row is the escape hatch, and it deserves more attention than it gets. Anonymous information sits outside the regulation because it is not personal data within the Article 4(1) definition; Recital 26 spells that out and sets the identifiability test. There is also a narrower route that keeps the data in scope but permits a longer period: the Article 5(1)(e) carve-out for statistical purposes under Article 89(1). It is the closest thing in the operative text to authority for holding a derived record in scope for longer, and it is narrower than it sounds — Recital 162 defines statistical purposes as producing aggregate results that are not used in support of measures or decisions about any particular person, and Article 89(1) attaches safeguards such as minimization and pseudonymization. The example below takes the other route, anonymization, and leaves the regulation's scope entirely. A record saying "24 August 2026, intent refund_status, confidence 0.41, escalated to human, no resolution" carries almost everything you need to improve the bot and identifies nobody, provided you have genuinely severed the link rather than merely dropped the name column. The bar for "genuinely" is high and is a question for someone qualified. But the shape of the answer is the useful part: retention is a per-field decision, not a per-table one.

Our own reviews barely record this

We searched our fifteen published platform reviews on 24 August 2026 for anything they record about retention. The result is thin enough to be worth publishing against ourselves.

Three of fifteen record something, and no two record the same kind of thing.

The Botpress review has a per-tier security and compliance matrix with a "Custom data retention + residency" row marked Enterprise only, read from the vendor's pricing comparison. That is a control, something the customer sets, and there is no number attached to any lower tier.

The Tars review records both kinds. Its compliance narrative lists "Configurable data retention" among the Enterprise items, and its pricing table separately records a flat 12-month retention on Premium, the only self-serve paid tier Tars publishes. So the control is at the top and a disclosed period sits below it.

The AiSensy review records no control but a period, and records it more precisely than either: a Retention period of "36 months past the start of the idle period," attributed to the vendor's privacy policy and marked identically across every tier. The same table also records an open-ended carve-out on deletion, under which the vendor "may retain some information to prevent fraud, troubleshoot problems, assist with investigations" — so the 36 months is the stated rule rather than a guarantee that nothing survives it.

Two things follow.

The first is a finding about our set, stated more narrowly than we first wrote it. Where a retention control appears, it appears only at the top tier. Botpress records it beside data residency; Tars's retention line sits with HIPAA, SSO and role-based access instead, though its review does note residency settings elsewhere, in a caption to an admin screenshot. Two of the three reviews also give an actual figure — 12 months at Tars Premium, 36 months from idle at AiSensy — and in both cases it is a period the customer inherits rather than sets.

The second is about us, and it is the more useful correction. Our review protocol has asked this question three times out of fifteen and never standardized it. Three reviews carry a retention line, in four different places between them: AiSensy in a per-tier compliance table, Botpress in a per-tier security matrix, Tars in its pricing table and again in its compliance narrative. Twelve carry none. That is worse than a blind spot, because the dimension already existed in four places across three reviews and never became a field. Three questions are going into the protocol as a result: what the default transcript retention period is on the entry tier, whether it is configurable and from which tier, and whether a deletion propagates to analytics and to any model provider in the path. Our methodology page carries the current dimension list; these will appear in the next revision of it. The twelve blanks remain evidence that we did not look rather than evidence that those platforms have nothing.

The arithmetic: why small bots need long windows

Here is the part that is genuinely counterintuitive, and it is the reason this entry exists rather than a shorter one.

Every instinct in privacy work points one direction: keep less, keep it briefly. Every instinct in bot improvement points the other: you cannot fix what you cannot read. For a large deployment the tension is mild, because a short window still contains a great deal of traffic. For a small one it is decisive, and the crossover can be located.

Take a concrete question an operator actually asks: why does the bot keep failing on refund requests? Answering it means reading a set of conversations where that failure occurred. Write:

  • C — conversations per day
  • f — the share of them that hit the failure you are studying
  • S — how many examples you want before you change anything
  • R — the retention window, in days, for raw transcripts

At steady state you hold C × R transcripts. The examples you care about accumulate at C × f per day, so the time to gather S of them is D = S ÷ (C × f). The study is possible only if the earliest example is still there when the last one arrives — that is, only if R ≥ D.

Put numbers in. Take a 12 percent failure share and a fifty-conversation sample, with days rounded up to whole days. One note on the 12 percent: our fallback rate entry defines that metric per message, and f here is the per-conversation share of one specific failure mode, so the figure is a stand-in for illustration rather than the same quantity.

Conversations per dayDays to accumulate 50 examplesWorks on a 30-day window?On 90 days?
3139NoNo
584NoYes, with 6 days to spare
1042NoYes
2021YesYes
509YesYes
2003YesYes

Rearranged, the threshold is C ≥ S ÷ (R × f). At S = 50 and f = 0.12 the floors are 13.89 conversations a day on a 30-day window, 4.63 on 90 days and 1.14 on a year — so in whole conversations, 14, 5 and 2. Below the floor, the transcripts that would answer the question are deleted before enough of them exist.

So the operator running three conversations a day on a thirty-day window has, without intending to, guaranteed that the fifty examples this method needs will never coexist in the window. About eleven are present at any moment; the fiftieth arrives on day 139, which is about three and a half months after the first one was deleted. The raw text is not gone; there is just never enough of it at once. The derived record, as the section below the assumptions box explains, is not subject to the constraint at all. And the shorter window feels like the more careful choice.

The resolution is not to keep transcripts longer. It is to notice that the constraint binds on the raw text and not on the derived record. Keep the intent label, the confidence score, the outcome and the timestamp for as long as you need to see a trend; delete the message text on the short schedule. You lose the ability to read the exact words, which is a real loss for the first diagnosis of a new failure mode. You keep the ability to count, which is what most of the work needs. That split is the whole argument for writing the policy per field.

Where policies fail in practice

Four failures account for most of it, and none is exotic.

The backup problem. Transcripts get deleted from the live store and survive in backups, which is the usual reason an erasure request is answered honestly and still leaves data behind. There is no clean answer here beyond documenting the backup cycle, putting a period on it, and being straight with people about it.

Deletion that does not propagate. The transcript goes; the copy in the analytics warehouse, the row in the CRM, the message log at the model provider and the export somebody made to a spreadsheet do not. A retention policy that governs one system out of five is a document about one system out of five.

A published period nobody implemented. The privacy notice says twelve months because that is what the template said. Nothing runs on a schedule. This is the failure that turns a paperwork problem into a statement you cannot support.

Retention set by whoever is loudest. Marketing wants everything forever, legal wants nothing past thirty days, and the period ends up wherever the last argument landed rather than attached to a purpose. The fix is unglamorous: one row per category, one sentence of justification each, and a date for the next review.

FAQ

What is a data retention policy?

It is the written rule stating, for each category of data an organization holds, how long that data is kept, what justifies the period, and what happens at the end — deletion, irreversible anonymization, or archival. For a chatbot the categories usually split into the raw transcript, the fields extracted from it, the derived analytics, and the copies held by the chat platform and any model provider in the path. Each of those four can carry a different period, and writing one period for all of them is the commonest way to end up with a policy nobody can execute.

Does GDPR specify a data retention period?

No, and this is the most commonly misreported thing about the subject. Article 5(1)(e) requires that personal data be kept "no longer than is necessary for the purposes for which the personal data are processed" and names no figure; its only other limb permits longer storage for archiving, research and statistical purposes under Article 89(1), which is again a category rather than a number. The specific periods that circulate under the regulation's name, such as thirty days or six months or seven years, trace back to sector rules, tax and accounting law, national employment law, or somebody's template. What the GDPR does require is that you determine a period, be able to demonstrate you did, and inform people of either the period or, under Article 13(2)(a), "the criteria used to determine that period."

How long should I keep chatbot transcripts?

There is no answer that does not depend on why you are keeping them, so decide the purpose first and let it set the period. In practice three purposes compete: dispute or complaint evidence, which is usually driven by how long a customer might come back; quality review and bot improvement, which the arithmetic on this page shows needs a longer window at low traffic than most operators expect; and analytics trends, which do not need the message text at all. Splitting the raw text from the derived record lets you take a short period on the first and a long one on the second, and that split is usually the whole solution.

Can a short retention period hurt my chatbot?

Yes, and the effect is largest on the smallest deployments. Diagnosing a specific failure by reading conversations means gathering enough examples of it, and at low traffic that takes longer than a short window keeps them. On our own illustrative assumptions (a 12 percent failure share we have not measured on any platform, and a fifty-example sample that is an editorial judgment rather than a statistical derivation), a 30-day window needs roughly 14 conversations a day before that reading is possible at all, and at 90 days the floor drops to about 5. Below the floor the raw evidence is deleted before enough of it exists. The fix is not a longer window on the transcripts but a separate, longer period on the non-identifying derived record, which keeps the ability to count what you can no longer read.

Do chatbot platforms let me set my own retention period?

Sometimes, and in our own reviewed set the control appears only at the top tier. Of the fifteen platform reviews we have published, three record anything about retention, and no two record the same kind of thing. Botpress's per-tier matrix lists custom data retention and residency as Enterprise only, with no period given for any tier. Tars records both: "Configurable data retention" at Enterprise, and a flat 12-month retention on Premium, its only self-serve paid tier. AiSensy records no control but a disclosed period — 36 months past the start of the idle period, attributed to the vendor's privacy policy and identical on every tier. The other twelve record nothing, which reflects a gap in our own review protocol rather than a finding about those platforms. Ask your vendor three questions: what the default period is on your tier, whether it is configurable and from which tier, and whether a deletion propagates to analytics and to any model provider in the path.

What is the difference between deleting and anonymizing data?

Deletion removes the record. Anonymization removes the link between the record and an identifiable person while keeping the rest, and genuinely anonymous information falls outside the regulation because it is no longer personal data within the Article 4(1) definition — Recital 26 spells this out and sets the identifiability test. That is why it is attractive. The catch is in the word "genuinely": dropping a name column while leaving a session identifier, a rare combination of attributes, or free text in which the customer typed their own email is pseudonymization, not anonymization, and pseudonymized data remains personal data. Whether a given transformation clears the bar is a question for a qualified advisor rather than a chatbot review site.

What should a retention policy actually contain?

One row per data category, and for each row: what the data is, where it lives, how long it is kept, the purpose that justifies that period, what happens at the end, and who is responsible for the deletion running. Add a review date. The test of the document is not its length but whether someone who has never seen it can pick a row, find the data, and confirm the deletion happened on schedule.

Sources

  • Regulation (EU) 2016/679 (General Data Protection Regulation), consolidated text, read 24 August 2026; Recital 162 read for the definition of "statistical purposes" relied on in the third section. Specifically: Article 5(1)(e), the storage limitation principle, quoted here in both of its limbs — the "kept in a form which permits identification of data subjects for no longer than is necessary for the purposes for which the personal data are processed" rule, and the carve-out permitting longer storage where data "will be processed solely for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1)"; Article 5(2), the accountability principle placing on the controller the responsibility to demonstrate compliance with 5(1); Article 4(1), the definition of personal data, which is what anonymous information falls outside; Article 13(2)(a), requiring that the data subject be informed of "the period for which the personal data will be stored, or if that is not possible, the criteria used to determine that period"; Article 17(1), the right to erasure, of whose six grounds 17(1)(a) applies where data "are no longer necessary in relation to the purposes for which they were collected or otherwise processed", together with the exemptions in 17(3); Recital 26 on the identifiability test for anonymous information; Recital 39, the source of the "limited to a strict minimum" and "time limits should be established by the controller" language quoted in the second section; and Article 89(1) for the safeguards attached to the statistical-purposes route. Quoted as the regulation's text, not as advice about its application. eur-lex.europa.eu
  • Chatbotscape review corpus, searched 24 August 2026 and published so the count reproduces. Two commands, both run from the repository root. Denominator: ls sample-reviews/*-review.md | wc -l returns 15. Numerator: grep -rliE "retention" sample-reviews/*-review.md returns three paths — aisensy-review.md, botpress-review.md, tars-review.md. The search is deliberately the loose concept word rather than the phrase "data retention", because the narrower string was what our first pass used and it missed the AiSensy row entirely; each of the three matched files was then read end to end and every retention line quoted before classification, a step the first pass skipped and which is why its count was wrong. The classification is ours. Botpress — per-tier security and compliance matrix, row "Custom data retention + residency", marked Enterprise only, read from the vendor's pricing comparison matrix; a control, no period at any tier. Tars — compliance narrative listing "Configurable data retention" at Enterprise, plus a separate pricing-table cell recording "12-month retention" on the Premium tier at $499/mo, which is the only self-serve paid tier Tars publishes; a control and a disclosed period. AiSensy — compliance table row "Retention period", value "36 months past the start of the idle period", attributed to the vendor's privacy policy and marked identically across all tiers; a disclosed period, no control. None of the three was established in a hands-on session. Published as a transparency statement about a gap in our own review protocol: twelve silent reviews mean we did not ask, and we do not present that silence as a finding about those platforms.
  • Chatbotscape retention-window model — the formula D = S ÷ (C × f), the six-row table and the derived thresholds (13.89 conversations a day at a 30-day window, 4.63 at 90 days, 1.14 at a year, all at S = 50 and f = 0.12; table days are rounded up to whole days, and the whole-conversation floors of 14, 5 and 2 are those ratios rounded up in turn) are our own arithmetic, shown in full so they can be checked rather than taken. We have not run this workload and publish no benchmark. The model assumes evenly spaced traffic, a stationary failure rate, that the failing conversations can be located at all, and a fifty-example sample that is an editorial judgment rather than a statistical derivation. The 12 percent failure share is a working figure for the illustration, not a measurement of any platform, and it is not the same quantity as the per-message metric defined at /glossary/chatbot-fallback-rate. Every one of those assumptions is stated where the figures appear.
  • Ahrefs Keywords Explorer, US overview and volume-by-country, queried 24 August 2026 — the search-demand, difficulty, parent-topic and country-split figures in this entry's keyword note, including the checks behind declining 'data minimization', 'privacy by design' and 'right to be forgotten', and behind sending 'data processing agreement' to the same-day companion guide.
  • Chatbotscape evaluation methodology. /methodology (continuously updated).