Skip to content
Chatbotscape
Microsoft's Azure AI Search documentation and Anthropic's Contextual Retrieval research were read on 27 September 2026, and every quotation and figure below carries that date. This is a living engineering practice and the recommended numbers can change as embedding models change.
Chunking· AI infrastructure
Chunking is the step where a source document — a help-center article, a PDF manual, a product page — gets split into smaller passages before each passage is turned into a vector embedding and stored for retrieval. It happens before embedding, not after: get the cut points wrong and no amount of good retrieval or a smarter model fixes what a badly-cut passage lost. Every retrieval-augmented generation (RAG) system that answers from your own documents does this somewhere, whether you ever see it happen or not.
By Chatbotscape Editorial· Methodology· Published 28 September 2026· Updated 28 September 2026

Chunking — Splitting Documents Before a Chatbot Can Retrieve From Them

Quick answer: Chunking is how a RAG chatbot's knowledge base gets built: documents are cut into smaller passages — Microsoft's own Azure AI Search guidance recommends starting around 512 tokens (roughly 2,000 characters) with about 25% overlap between neighboring chunks — before each passage is embedded and indexed. Cut too small and a passage loses the context that made it answerable; cut too large and retrieval drags in unrelated material alongside the useful part. The three real strategies in production are fixed-size (split by a set length, with overlap), structure-aware or semantic (split at paragraph, heading or sentence boundaries so a chunk stays a coherent unit), and a newer technique, Anthropic's Contextual Retrieval, which prepends a short AI-written summary to each chunk before embedding it and measured a 35–67% reduction in retrieval failures depending on configuration.

Why the cut points matter more than the database

Retrieval quality is decided upstream of anything a vector database does. Microsoft's Azure AI Search documentation states the core reason chunking exists at all: "Partitioning large documents into smaller chunks can help you stay under the maximum token input limits of chat completion and embedding models," adding that this "is helpful for retrieval augmented generation (RAG) or agentic retrieval as well" because it "prevents data loss due to truncation." But the practical failure mode people hit is not the token limit — it's meaning. A chunk cut mid-sentence, or one that separates a number from the sentence explaining what the number means, can be embedded and retrieved perfectly and still be useless to the model reading it. The vector database entry puts this plainly: "the database faithfully searches whatever it was given" — it has no way to notice that a chunk was cut badly, because a badly-cut chunk still returns a similarity score.

Fixed-size chunking: the default, with real numbers

The simplest approach splits a document at a fixed length and repeats a portion of each chunk in the next one so meaning doesn't break cleanly at the seam. Azure AI Search's own guidance gives a specific starting point: "We recommend starting with a chunk size of 512 tokens (approximately 2,000 characters) and an initial overlap of 25%, which equals 128 tokens. This overlap ensures smoother transitions between chunks without excessive duplication." The same documentation frames the trade-off directly: "highly structured data might require less overlap, while conversational or narrative text might benefit from more," and separately, on chunk size itself: "If you need intact text or passages, larger chunks and variable chunking that preserves sentence structure can produce better results," while "Large Language Models (LLM) have performance guidelines for chunk size" that constrain how large a chunk can get before it stops helping.

Structure-aware and semantic chunking

Rather than cutting at a fixed character count, structure-aware chunking splits at the document's own boundaries — headings, paragraphs, table rows — so a chunk stays a coherent unit instead of an arbitrary slice. Azure AI Search's Document Layout approach works this way: it "chunk[s] content based on document structure, capturing headings and chunking the content body based on semantic coherence, such as paragraphs and sentences," on the reasoning that "when those chunks are of higher quality and semantically coherent, the overall relevance of the query is improved." The trade-off is that structure-aware chunking needs a document with real structure to exploit — a scanned PDF with no heading tags gets little benefit over a fixed-size cut.

Contextual Retrieval: fixing what chunking loses, after the fact

A chunk taken out of its document can lose the context that made it meaningful even when the cut point was reasonable. Anthropic's own example: a financial-filing chunk reading "The company's revenue grew by 3% over the previous quarter" is a correct, cleanly-cut sentence that answers nothing on its own, because it doesn't say which company or which quarter. Anthropic's Contextual Retrieval technique "solves this problem by prepending chunk-specific explanatory context to each chunk before embedding ('Contextual Embeddings') and creating the BM25 index ('Contextual BM25')," using a model to generate that context (the prompt asks for "a short succinct context to situate this chunk within the overall document for the purposes of improving search retrieval") — and, in Anthropic's own words, "the resulting contextual text, usually 50-100 tokens, is prepended to the chunk before embedding it and before creating the BM25 index." The measured effect, per Anthropic's own published research: "Contextual Embeddings reduced the top-20-chunk retrieval failure rate by 35%," and "Combining Contextual Embeddings and Contextual BM25 reduced [the] failure rate by 49%," rising to a 67% reduction with reranking added. The measured effect, per Anthropic's own published research: "Contextual Embeddings reduced the top-20-chunk retrieval failure rate by 35%," and "Combining Contextual Embeddings and Contextual BM25 reduced [the] failure rate by 49%," rising to a 67% reduction with reranking added. This is a real, named, and measured technique — not a marketing description of "better AI" — and it is worth knowing the name of if a platform's sales team claims to have solved retrieval accuracy without saying how.

What our own reviews show is happening behind the curtain

If you use a chatbot platform's built-in "upload your docs" knowledge-base feature — Chatbase, the knowledge nodes in Botpress or Voiceflow, or the AI-answer step inside a marketing-first builder — chunking already happened for you, and most platforms don't expose the parameters. Our own hands-on Botpress review measured this directly: after testing the platform's Knowledge Base feature against a five-PDF technical corpus, the review states plainly that "specific RAG implementation depth (chunking strategy, retrieval rank, citation format) requires hands-on testing to score precisely" — even a developer-facing platform doesn't document its own chunking parameters anywhere a reviewer could read them in advance. The comparison across two platforms in the same test protocol is the closest thing our corpus has to a chunking-strategy signal: Voiceflow's knowledge-base embedding generation measured about 45 seconds against Botpress's about 30 seconds for an equivalent five-PDF set, a difference our review attributes to "longer chunking strategy" rather than raw processing speed. Neither platform's help documentation states its default chunk size or overlap. If you are building a custom RAG pipeline instead of using a platform's built-in feature, LangChain, LlamaIndex, Pinecone and Azure AI Search all expose chunking as a parameter you set directly — the RAG build guide covers that decision.

What chunking is not

It is not the vector database. A vector database stores and searches chunks after they exist; chunking is the step that produces what gets stored. Changing databases with the same bad chunks produces the same bad retrieval.

It is not embedding. A vector embedding is the numeric representation a chunk is converted into; chunking decides what text goes into that conversion, embedding is a separate, later step performed on the finished chunk.

It is not the dominant sense of the word "chunking." General search volume on the bare term "chunking" is dominated by the cognitive-psychology sense — a memory technique for grouping information (Ahrefs' parent-topic data confirms this: "chunking" resolves to "chunking psychology"). This entry covers only the document-splitting step in AI retrieval systems.

FAQ

What is chunking in RAG?

The step where a source document is split into smaller passages before each one is turned into a vector embedding and indexed for retrieval. It happens before embedding, and it decides what a retrieval system can find later — a badly-cut chunk stays badly cut no matter how good the search that finds it is.

What chunk size should I use?

Microsoft's Azure AI Search documentation recommends starting around 512 tokens (about 2,000 characters) with roughly 25% overlap between chunks, then adjusting for your content: less overlap for highly structured data, more for conversational or narrative text, and larger chunks when a passage needs to stay intact to make sense.

What is semantic chunking?

Splitting text at its own structural boundaries — headings, paragraphs, sentences — rather than at a fixed character count, so each chunk is a coherent unit instead of an arbitrary slice. It works best on documents that have real structure to split along.

What is Contextual Retrieval?

An Anthropic technique that prepends a short, AI-generated summary of where a chunk sits in its source document to the chunk itself before embedding it, so the chunk carries context it would otherwise lose when taken out of the document. Anthropic's own measurements show it reduced retrieval failure rates by 35% alone, 49% combined with a keyword-search layer, and 67% with reranking added.

Can I control chunking on a chatbot platform like Botpress or Chatbase?

Usually not directly. Platform knowledge-base features handle chunking automatically, and our own hands-on testing found that even a developer-facing platform (Botpress) doesn't document its chunking parameters anywhere a buyer could check before signing up. If you need direct control over chunk size and strategy, that is a reason to build a custom RAG pipeline rather than use a platform's built-in feature — see the RAG build guide.

Sources

  • Microsoft Learn, Chunk documents for vector search and RAG workloads — learn.microsoft.com/en-us/azure/search/vector-search-how-to-chunk-documents, read 27 September 2026. Source of the token-limit rationale, the 512-token/2,000-character/25%-overlap recommendation, and the structured-vs-narrative overlap guidance.
  • Microsoft Learn, Semantic chunking with the Document Layout skill — learn.microsoft.com/en-us/azure/search/search-how-to-semantic-chunking, read 27 September 2026. Source of the structure-aware chunking definition and the semantic-coherence quotation.
  • Anthropic, Introducing Contextual Retrieval — anthropic.com/news/contextual-retrieval, read 27 September 2026. Source of the revenue-growth chunk example, the contextualizing-prompt quotation, and the 35%/49%/67% failure-rate reduction figures.
  • Chatbotscape review corpus — Botpress review and Voiceflow review, read 27 September 2026: the Knowledge Base hands-on test sections (five-PDF technical corpus, Claude-routed), the "requires hands-on testing to score precisely" quotation on Botpress's chunking strategy, and the 30-second vs 45-second embedding-generation comparison.
  • Chatbotscape evaluation methodology. /methodology (continuously updated).