Skip to content

Repository files navigation

Adaptive Chunking for RAG

Automatically choose better chunks per document, with explainable scoring and optional OCI support

CI License: MIT Python 3.10+ arXiv

Adaptive Chunking for RAG helps you stop guessing which splitter to use. It tries multiple chunking strategies for each document, scores the candidates with intrinsic quality metrics, and returns the best chunks with diagnostics that explain the choice.

The Python package is published as adaptive-oci-chunking because it includes optional Oracle Cloud Infrastructure adapters, but the core library is cloud-neutral and runs locally. OCI Object Storage and Generative AI support are opt-in extras, not required dependencies.

Use it when:

  • your RAG pipeline contains mixed PDFs, Markdown, policy docs, docs-as-code, or long structured text;
  • recursive/fixed chunking cuts through sections, tables, references, or code blocks;
  • you want chunk-level metadata, offsets, section paths, and selection diagnostics before indexing;
  • you need a drop-in chunking step for LangChain, LlamaIndex, vector databases, or a custom ingestion pipeline.
pip install adaptive-oci-chunking
adaptive-chunk chunk handbook.md --include-candidates --json
from adaptive_chunking import AdaptiveChunker

result = AdaptiveChunker().chunk_file("handbook.md")

print(result.strategy_name, round(result.score, 3))
for chunk in result.chunks:
    print(chunk.metadata.get("section_path"), chunk.text[:120])

Why not just use one splitter?

No single chunking method works best for every document in a RAG pipeline. A Markdown handbook, a legal PDF, a JSON export, and a code file usually need different boundaries. Adaptive chunking treats chunking as a selection problem: generate candidate chunks, score each candidate, and keep the best result for that document.

The selection is inspectable. If a strategy loses, you can see whether it dropped content, duplicated too much text, broke natural boundaries, produced poor cohesion, or missed the target size.

Try the adoption-focused assets

  • Benchmark your own documents against fixed, recursive, Markdown, section-aware, and token-window strategies:

    python benchmarks/compare_chunkers.py examples/benchmark_policy.md --format markdown
  • Wire adaptive chunks into vector database ingestion recipes:

    python examples/vector_store_ingestion.py examples/sample.md
  • Run the local playground for side-by-side chunk inspection:

    pip install "adaptive-oci-chunking[demo]"
    streamlit run demos/streamlit_app.py

Architecture

Adaptive OCI Chunking architecture

This project is inspired by Ekimetrics' adaptive-chunking repository and the paper Adaptive Chunking: Optimizing Chunking-Method Selection for RAG. It keeps the core dependency-light, adds production-oriented metrics, and includes optional adapters for OCI, LangChain, LlamaIndex, APIs, PDFs, and vector-store ingestion recipes.

Features

  • Candidate chunkers:
    • single-document
    • fixed window with overlap
    • token window with token overlap
    • recursive split
    • sentence-aware
    • paragraph-aware
    • split-then-merge
    • section-aware
    • Markdown-aware
    • delimiter-aware
    • page-aware
    • page-index hierarchical
    • semantic lexical drift
    • regex-guided section splitting
    • HTML text extraction
    • JSON structural splitting
    • code symbol splitting
    • hybrid structure-first splitting
  • Metric-guided selection using paper-aligned intrinsic metrics:
    • References Completeness (RC)
    • Intrachunk Cohesion (ICC)
    • Document Contextual Coherence (DCC)
    • Block Integrity (BI)
    • Size Compliance (SC)
  • Additional practical metrics:
    • source coverage
    • overlap control
    • boundary quality
    • semantic drift
    • information density
    • redundancy
  • Weighted strategy selection with explainable per-metric scores.
  • LangChain TextSplitter adapter.
  • LlamaIndex node conversion and parser-style adapter.
  • CLI for local text/Markdown files.
  • Optional OCI Object Storage loader and OCI Generative AI embedding adapter.
  • Small, dependency-light core for local document chunking workflows.
  • Benchmark harness for comparing chunking strategies on your documents.
  • Vector-store ingestion recipe for Chroma, Qdrant, Pinecone-style payloads, and custom databases.
  • Streamlit playground for inspecting selected chunks and candidate rankings.

Evidence and benchmarking

This repository now includes a lightweight benchmark harness in benchmarks/compare_chunkers.py. It compares adaptive selection against named baseline strategies using the same intrinsic metrics that the selector exposes. For a fair RAG-ingestion comparison, the adaptive row selects among the same strategy list passed through --strategies.

Run it on the bundled example:

python benchmarks/compare_chunkers.py examples/benchmark_policy.md

Run it on a folder of your own text/Markdown/PDF documents:

python benchmarks/compare_chunkers.py path/to/docs --strategies fixed-window recursive markdown section-aware token-window --format markdown

For production adoption, pair these intrinsic metrics with task-level retrieval evaluation such as recall@k, answer correctness, RAGAS, DeepEval, or your internal golden-question set. The benchmark script is intentionally small so teams can adapt it to their own documents instead of trusting a canned dataset that may not match their domain.

Contributing

Contributions are welcome for new chunkers, metrics, examples, integrations, benchmarks, documentation, and bug fixes.

See CONTRIBUTING.md for setup instructions, PR expectations, and guidance for adding chunkers or metrics.

Maintained by Yash Shukla, focused on AI, cloud, and RAG systems.

Install

Install the latest release from PyPI:

pip install adaptive-oci-chunking

The package installs the core local chunking toolkit. Optional extras are available for OCI, the API server, and framework integrations:

pip install "adaptive-oci-chunking[oci]"
pip install "adaptive-oci-chunking[api]"
pip install "adaptive-oci-chunking[demo]"
pip install "adaptive-oci-chunking[pdf]"
pip install "adaptive-oci-chunking[langchain,llama-index]"

For local development from a cloned checkout:

pip install -e ".[dev]"

With OCI support from source:

pip install -e ".[oci]"

With the API server from source:

pip install -e ".[api]"

With framework integrations from source:

pip install -e ".[langchain,llama-index]"

Quick Start

Check the installed package version:

python -c "import adaptive_chunking; print(adaptive_chunking.__version__)"

From a cloned checkout, run the bundled sample through the CLI:

adaptive-chunk chunk examples/sample.md --json

After installing from PyPI in any project, create a small Markdown file and chunk it:

printf "# Demo\nAdaptive chunking chooses a splitter per document.\n\n## Details\nChunks keep related context together.\n" > sample.md
adaptive-chunk chunk sample.md --json
adaptive-chunk chunk sample.md --strategy markdown --strategy semantic --json
adaptive-chunk strategies

Python usage:

from adaptive_chunking import AdaptiveChunker, ChunkingConfig

text = "## Introduction\nAdaptive chunking chooses a splitter per document.\n\n## Details\n..."
chunker = AdaptiveChunker(
    config=ChunkingConfig(strategies=["markdown", "token-window", "semantic"])
)
result = chunker.chunk(text, document_id="demo")

print(result.strategy_name)
for chunk in result.chunks:
    print(chunk.text)

print(result.to_json(indent=2))

PDF files

Install the PDF extra, then pass a PDF path directly to the high-level chunker. Pages are separated internally with form-feed boundaries, so the page and page-index candidate strategies retain page_index metadata when selected.

pip install "adaptive-oci-chunking[pdf]"
adaptive-chunk chunk handbook.pdf --json
from adaptive_chunking import AdaptiveChunker

result = AdaptiveChunker().chunk_file("handbook.pdf")
print(result.strategy_name, result.to_json(indent=2))

PDF extraction supports text-based PDFs. Scan- or image-only PDFs must be OCRed before chunking.

Examples

Runnable examples live in examples/:

  • basic_adaptive_chunking.py: end-to-end adaptive selection with metric output.
  • custom_selector.py: custom chunker list and metric weights.
  • langchain_integration.py: LangChain TextSplitter usage.
  • llama_index_integration.py: LlamaIndex TextNode conversion.
  • oci_object_storage.py: loading source text from OCI Object Storage.
  • vector_store_ingestion.py: vector-store-ready records for Chroma, Qdrant, Pinecone-style SDKs.
  • benchmark_policy.md: structured sample document for benchmark comparisons.

Additional adoption assets:

  • benchmarks/compare_chunkers.py: compare adaptive selection with baseline chunkers.
  • demos/streamlit_app.py: local playground for inspecting chunks and candidate rankings.
  • docs/ADOPTION.md: positioning and launch checklist.
  • docs/RELEASE.md: release checklist.

Chunker Options

from adaptive_chunking.chunkers import (
    DelimiterChunker,
    MarkdownChunker,
    PageIndexChunker,
    PageChunker,
    SectionAwareChunker,
    SemanticChunker,
    TokenWindowChunker,
)
from adaptive_chunking.selector import AdaptiveSelector
from adaptive_chunking import AdaptiveChunker

selector = AdaptiveSelector(
    chunkers=[
        MarkdownChunker(max_size=1800),
        TokenWindowChunker(chunk_tokens=240, overlap_tokens=24),
        SectionAwareChunker(max_size=1800),
        DelimiterChunker(delimiter="\n---\n"),
        PageChunker(page_delimiter="\f"),
        PageIndexChunker(page_delimiter="\f"),
        SemanticChunker(max_size=1400, similarity_threshold=0.08),
    ]
)

result = AdaptiveChunker(selector=selector).chunk(text)

PageIndexChunker is separate from PageChunker: it splits by page first, then by heading hierarchy inside each page. Chunks include page_index, heading_path, section_path, section_title, and section_instance_id metadata so repeated headings such as Overview remain tied to the correct page and section occurrence during retrieval.

Every chunk returned through AdaptiveChunker also has a section_path list for retrieval filtering, for example ["Access", "MFA"]. Lines labelled as document headers, footers, page numbers, or standard page-furniture labels are ignored as section boundaries.

Available built-in strategy names can also be discovered at runtime:

from adaptive_chunking import registry

print(registry.names())
chunker = registry.create("markdown", max_size=1200)

The current built-ins are:

code
delimiter
fixed-window
html
hybrid
json
markdown
page
page-index
paragraph
recursive
regex-section
section-aware
semantic
sentence
single
split-then-merge
token-window

Metrics

The selector ranks every candidate by a weighted average of intrinsic scores. The first five metrics follow the paper's evaluation dimensions; the additional metrics make the implementation more practical for production RAG systems where dropped text, excessive overlap, and duplicated chunks are common failure modes.

Weights can be tuned:

from adaptive_chunking.metrics import IntrinsicMetricEvaluator, MetricConfig, MetricWeights
from adaptive_chunking.selector import AdaptiveSelector

weights = MetricWeights(
    block_integrity=1.4,
    coverage=1.5,
    redundancy=0.8,
)
evaluator = IntrinsicMetricEvaluator(MetricConfig(weights=weights))
selector = AdaptiveSelector(evaluator=evaluator)

Adaptive Scoring

For each document, the selector runs every candidate chunker and evaluates the chunks it produces. Each candidate receives a normalized weighted score:

score(candidate) = sum(metric_value_i * metric_weight_i) / sum(metric_weight_i)

Where:

  • metric_value_i is the metric score for a candidate, normalized from 0.0 to 1.0.
  • metric_weight_i controls how important that metric is for selection.
  • Higher scores are better.
  • Candidates are ranked from highest score to lowest score.

For example, a domain that cares about preserving source text and section boundaries might emphasize coverage and block_integrity:

Metric Value Weight Weighted value
coverage 1.00 1.50 1.50
block_integrity 0.90 1.40 1.26
redundancy 0.80 0.80 0.64
score = (1.50 + 1.26 + 0.64) / (1.50 + 1.40 + 0.80)
      = 3.40 / 3.70
      = 0.919

You can inspect every candidate, not just the winner:

from adaptive_chunking import AdaptiveChunker

result = AdaptiveChunker().chunk(text, document_id="demo")

for candidate in result.candidates:
    print(candidate.strategy_name, round(candidate.score, 3), len(candidate.chunks))
    for metric in candidate.metrics:
        print(" ", metric.name, metric.value, "weight=", metric.weight)

This makes the selection process explainable: if a chunker loses, you can see whether it dropped content, produced excessive overlap, cut through structure, or failed a size constraint.

LangChain

from langchain_core.documents import Document
from adaptive_chunking.langchain import LangChainAdaptiveTextSplitter

splitter = LangChainAdaptiveTextSplitter()
documents = splitter.split_documents([
    Document(page_content=text, metadata={"source": "policy.md"})
])

# Or load and chunk a PDF directly (requires the `pdf` extra too).
pdf_documents = splitter.split_pdf("handbook.pdf")

# Every output Document retains source metadata plus document_id, chunk_index,
# start_char, end_char, strategy_name, adaptive_score, and structure metadata.

LlamaIndex

from llama_index.core.schema import Document
from adaptive_chunking.llama_index import LlamaIndexAdaptiveParser

parser = LlamaIndexAdaptiveParser()
nodes = parser.get_nodes_from_documents([
    Document(text=text, metadata={"source": "policy.md"})
])

For a direct PDF-to-node workflow:

from adaptive_chunking.llama_index import pdf_to_llama_nodes

nodes = pdf_to_llama_nodes("handbook.pdf")

LlamaIndexAdaptiveParser is a native LlamaIndex NodeParser, so it can also be passed directly to an IngestionPipeline. Output nodes retain document metadata, adaptive diagnostics, offsets, and standard source/previous/next relationships.

OCI Usage

Copy .env.example and set the values for your tenancy and compartment. The core library does not require OCI credentials unless you instantiate an OCI adapter.

from adaptive_chunking.oci import OCIObjectStorageTextLoader

loader = OCIObjectStorageTextLoader(
    namespace="my-namespace",
    bucket_name="documents",
)
text = loader.load_text("policies/example.md")

API Server

uvicorn adaptive_chunking.api:app --reload

Then post:

curl -X POST http://127.0.0.1:8000/chunk \
  -H "Content-Type: application/json" \
  -d "{\"text\":\"# Title\nBody text\", \"document_id\":\"demo\"}"

Project Layout

src/adaptive_chunking/
  chunkers.py      # candidate splitting strategies
  metrics.py       # intrinsic metric implementations
  selector.py      # weighted adaptive strategy selection
  pipeline.py      # high-level AdaptiveChunker
  langchain.py     # optional metadata-preserving LangChain adapter
  llama_index.py   # optional native LlamaIndex NodeParser and node helpers
  oci.py           # optional OCI adapters
  api.py           # optional FastAPI app
  cli.py           # command line interface
tests/
examples/
benchmarks/
demos/
docs/

Notes

This repo is designed as a clean, extensible foundation rather than a verbatim copy of the reference implementation. The metric implementations are practical approximations intended for engineering use and experimentation. Production RAG deployments should calibrate weights, chunk sizes, and embedding models against their document domains.

References

Citation

If this project helps your work, please cite the original adaptive chunking paper:

@inproceedings{demoura2026adaptive,
    title={Adaptive Chunking: Optimizing Chunking-Method Selection for RAG},
    author={de Moura Junior, Paulo Roberto and Lelong, Jean and Blangero, Annabelle},
    booktitle={Proceedings of the 15th Language Resources and Evaluation Conference (LREC 2026)},
    year={2026},
    url={https://arxiv.org/abs/2603.25333},
}

License

This project is licensed under the MIT License.

About

Adaptive document chunking for RAG with intrinsic evaluation metrics, LangChain/LlamaIndex integrations, and optional OCI support.

Resources

Contributing

Stars

6 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages