RAG Chunk Size and Overlap
RAG chunking involves more than cutting a document into equal lengths. Evidence needs enough context to stand on its own, overlap may preserve boundary information, and repeated passages can crowd the final context. Define the length unit before choosing a size.
Contents
See RRF for hybrid retrieval and Cross-Encoder reranking for the later stages. Chunking happens before them. 中文版
Characters and tokens are different units
A chunk_size setting does not inherently mean tokens. The LangChain recursive character splitter example measures length with len, so it counts characters. Changing the length function changes the measurement. Model input limits must be checked with the appropriate tokenizer; 500 characters are not automatically 500 tokens.
Chinese text, English text, code and punctuation can tokenize differently. Equal character counts can produce different token counts. Record both: character offsets help locate the source, while token counts govern input budgets. Titles, prompts and other added text also consume the actual input budget.
Size changes how evidence is packaged
Short chunks can focus on one subject but separate a conclusion from its conditions. Long chunks retain more context but may mix topics or exceed a downstream input budget. Neither direction guarantees better results.
| Observed failure | Inspect first |
|---|---|
| A conclusion appears without its conditions | Boundaries, headings and neighboring context |
| Retrieved passages mix unrelated subjects | Paragraph structure and chunk size |
| Critical evidence is truncated | Actual token budget and truncation position |
| Top results largely repeat each other | Overlap and context deduplication |
These are troubleshooting suggestions, not experimental findings. Preserving headings, paragraphs and source positions before splitting long paragraphs makes evidence easier to trace. Inspect code blocks, tables and conditional instructions separately instead of relying only on average chunk length.
Overlap preserves boundaries at a cost
A fixed window makes the cost visible. Let C be the size and O the overlap, with 0 ≤ O < C. The stride is S = C − O. The diagram uses a 20-unit sequence, C=8 and O=2. Units may be characters or tokens, but one calculation must use the same unit throughout.
The windows are [0,8), [6,14) and [12,20). Their combined length is 24 units, including four units of repeated coverage, not four new facts. For a long sequence, ignoring the tail and metadata, cumulative length relative to the source is approximately C/(C−O). This estimates repetition, not quality improvement or an exact disk-space multiplier.
A recursive splitter tries separators in order and treats overlap as a target. Its outputs need not follow this fixed-stride diagram. Chinese lacks a universal space-delimited word boundary; the official guide recommends adding separators such as the Chinese full stop. Actual semantic boundaries still require sampling.
Verify coverage and cost with code
This independent teaching function returns half-open fixed-window intervals. It is not LangChain’s recursive algorithm and does not detect semantic boundaries. It was executed with coverage checks for lengths 0 through 39, sizes 1 through 11 and every valid overlap, also checking each window’s maximum length.
def spans(length, size, overlap):
if length < 0 or size <= 0 or not 0 <= overlap < size:
raise ValueError('invalid window parameters')
result = []
start = 0
while start < length:
end = min(start + size, length)
result.append((start, end))
if end == length:
break
start += size - overlap
return result
assert spans(0, 8, 2) == []
assert spans(8, 8, 2) == [(0, 8)]
assert spans(20, 8, 2) == [(0, 8), (6, 14), (12, 20)]
for n in range(40):
for size in range(1, 12):
for overlap in range(size):
ranges = spans(n, size, overlap)
covered = {i for a, b in ranges for i in range(a, b)}
assert covered == set(range(n))
assert all(0 < b-a <= size for a,b in ranges)
print('spans:', spans(20, 8, 2))
print('stored units:', sum(b-a for a,b in spans(20, 8, 2)))
print('coverage checks passed')
The output is [(0,8),(6,14),(12,20)] with cumulative length 24. Empty input produces an empty list. Once a window reaches the end, the function stops without adding a tail window containing only repeated material. The fixed-stride window count is zero for N=0; otherwise it is 1 + ceil(max(0,N−C)/(C−O)).
Handle O≥C explicitly: it creates a zero or negative stride, so this function rejects it. A production system also needs document IDs, versions, offsets and access-control metadata, keeping repeated chunks traceable to their source.
Keep comparisons controlled
Changing chunk size, embedding model, retrieval and reranking together makes it impossible to attribute an answer change to chunking. Keep the other components fixed while comparing boundaries or overlap settings on the same query set.
Measure whether complete answer evidence is covered, not merely whether some related chunk was retrieved. Different chunking schemes change chunk IDs and counts, making raw chunk-level hit rates hard to compare. Map evidence labels to source-document intervals and evaluate coverage of those intervals. Track index entries, cumulative tokens, repetition in final context and latency as well.
No length works best for every document. Respect clear document structure first. For dependencies across paragraphs, test modest overlap or expand neighboring passages after retrieval. For redundant context, inspect deduplication and context selection. Change one major variable at a time so the useful intervention is identifiable.
References
LangChain: Recursive text splitter
Publication note: this article was published as a catch-up on October 9, 2026 (Beijing time), retaining the originally scheduled 19:00 article date. Intervals, code and cost calculations are original teaching examples, not model-quality benchmarks.
Support
If this article helped you, you can support this site.
Click a QR code to enlarge. More options: support page。


