A knowledge base (lectures, documentation, articles) is the foundation of any RAG system. The quality of responses directly depends on how you prepare the data. This sounds obvious, but in practice, 80% of problems in RAG are not about choosing the LLM or embedding parameters, but rather incorrect chunking.

What is a chunk, really?

Formally, a chunk is a unit of search, a fragment of text that is turned into a vector (embedding). It’s important to remember: at the semantic search stage, there is no semantics in the usual sense—there is only mathematical comparison of vectors in multidimensional space.

Therefore, a chunk is primarily a unit of comparison. Its parameters should be chosen based on how the embedding is formed and how retrieval is performed (top-k, filters, reranking).

The main question: what size to choose?

Many make an intuitive mistake: they make chunks large, thinking that “the more information, the better.” But the embedding model compresses text into a fixed-dimensional vector.

  • Oversized Chunks: If there are too many topics within a chunk, the vector becomes “averaged,” and the meaning gets blurred. As a result, it poorly matches the specific user query.
  • Small Chunks: A “point” vector is good for searching but loses context. If a thought is split into 3 parts, the LLM will only get a fragment and provide an incomplete answer.

A practical pipeline consists of 4 pillars of configuration

Chunking errors cannot be fully cured by prompt engineering or Query Rewriting. If a relevant piece doesn’t make it into the sample, the LLM simply won’t see it.

1. Adaptive size for the task
There is no “gold standard” of 512 tokens. You need to focus on the semantic density of the text:

  • Legal / Medical: meaning is often blurred across the hierarchy of documents. It’s better to use larger chunks (1500+ characters) with structural analysis.
  • Technical Docs / Instructions: the density of meaning is already high. Chunks of 500–700 characters work better, as the user’s query is usually very specific.
  • Academic lectures: A balance of 1000–1500 characters works well.

2. Overlap

Add about 30% of text from the previous chunk into each subsequent chunk. This is “insurance”: if a thought ends up on the boundary of a cut, it will fully appear in at least one of the neighboring chunks, which is critical for correct searching or reranking.

3. Cutting along logical boundaries (Structural Splitting)

Cut the text using recursive splitters (RecursiveCharacterTextSplitter) that prioritize natural delimiters:

  • Headings and line breaks (\ \ , \ )
  • Punctuation marks (periods, semicolons)

A chunk should be as close as possible to a complete thought.

4. Sometimes use Semantic Chunking

For the most critical tasks, Semantic Chunking is used. Instead of counting characters, we analyze the cosine distance between neighboring sentences. If it sharply increases, it means the topic has changed, and that’s where the cut should be made. This makes the boundaries of the chunk the most meaningful.

What other approaches to text chunking are there?

  • Manual annotation: Experts read documents and highlight semantic blocks into separate pieces of text. The problem is obvious: it’s time-consuming, expensive, and impossible to scale.
  • Annotation through LLM: We provide a large context window to the LLM—asking it to break it into semantic pieces without changing the meaning of the text. It’s expensive and unstable.
  • Semantic search: Even with simple chunking, results can be improved through reranking, query expansion, and multi-step retrieval.

Chunking is a stage that directly affects the geometry of the vector space and the quality of search

Our task is not to “preserve meaning perfectly,” but to increase the likelihood that a relevant piece will make it into retrieval.

An error at this stage cannot be compensated for by either the model or prompts. After all, in the end, if the LLM does not receive relevant chunks, it will not be able to provide a relevant answer to the user.