DEV Community

Cover image for chunklet-py v3.0.0: a smaller install, a smarter chunker, and no more global state
Speedyk-005
Speedyk-005

Posted on

chunklet-py v3.0.0: a smaller install, a smarter chunker, and no more global state

There's a new chunker that learns its own boundaries, two heavy dependencies moved to optional extras, and a few configuration changes that stop you from repeating yourself.

Here's what was added, changed, or removed.

pip install chunklet-py -U

Release: v3.0.0

This post summarizes the migration guide, what's new and the self-tuning chunker docs. Both have the full detail if you want to go deeper on any section.

The self-tuning chunker (new)

The self-tuning chunker sits on top of DocumentChunker and CodeChunker. It classifies each input as document or code, keeps a running profile per type, and sizes chunks from the profile it has learned so far. Classification goes by extension first, then binary sniffing, then content heuristics for anything extensionless.

from chunklet import SelfTuningChunker

def token_counter(text: str) -> int:
    return len(text.split())

chunker = SelfTuningChunker(lang="en", token_counter=token_counter)

for chunk in chunker.chunk_texts([prose, code]):
    print(chunk.metadata.inferred_type)  # "document" or "code"
Enter fullscreen mode Exit fullscreen mode

It supports the other chunking methods too: chunk_text, chunk_file, chunk_files.

Two profiles, both starting from sensible defaults:

{
    "document": {
        "max_sentences": 7.0,
        "header_density_ratio": 0.05,
        "max_section_breaks": 1,
        "max_tokens": 512.0,
    },
    "code": {
        "max_lines": 15.0,
        "max_functions": 1,
        "max_tokens": 512.0,
    },
}
Enter fullscreen mode Exit fullscreen mode

Each metric updates with a Kaufman Adaptive Moving Average, so stable metrics hold their estimate and swinging ones catch up fast. Code and document profiles are learned independently. You can warm-start from a previous run by passing the saved state back in as initial_state, and a hard_token_limit keeps learned limits from ballooning on huge files.

Note: If you don't pass a token_counter, the chunker ignores max_tokens.

For more detail, see the self-tuning chunker docs, which cover the KAMA mechanics, the mixed-corpus walkthrough, and how this compares to per-document adaptive tools.

The dependency diet

py3langid and indic-nlp-library left the core install, taking their numpy, pandas, and morfessor dependencies with them. Most users touched neither.

They're optional extras now:

pip install 'chunklet-py[lang-detect]'   # for lang="auto", renamed from [auto]
pip install 'chunklet-py[indic]'         # Indic language support
pip install 'chunklet-py[self-tuning]'   # bundles struct-doc + code
Enter fullscreen mode Exit fullscreen mode

lang no longer defaults to "auto" in the affected constructors. In v2 it did, which meant every install pulled in py3langid whether you needed it or not. If you previously passed an explicit language code, nothing changes and you don't need the extra.

Constraints moved to the constructor

Configuration parameters used to be passed on every call. Now they're set once at construction and reused.

# Before
chunker = DocumentChunker()
chunks = chunker.chunk_text(text, lang="en", max_sentences=3, max_section_breaks=1)

# After
chunker = DocumentChunker(lang="en", max_sentences=3, max_section_breaks=1)
chunks = chunker.chunk_text(text)
Enter fullscreen mode Exit fullscreen mode

Two reasons. First, you stop repeating the same five arguments on every call. Second, each chunker owns its own configuration and has plain attributes, so you can still change them between calls:

chunker.max_sentences = 5
Enter fullscreen mode Exit fullscreen mode

max_tokens, max_sentences, max_section_breaks, overlap_percent, and lang all moved. Same for CodeChunker (max_tokens, max_lines, max_functions) and SentenceSplitter (lang).

Registries are per-instance now

The global custom_processor_registry is gone. You create a CustomProcessorRegistry() yourself and pass it to the chunker that should use it.

from chunklet.document_chunker import CustomProcessorRegistry, DocumentChunker

registry = CustomProcessorRegistry()

@registry.register(".json", name="MyJSONProcessor")
def my_json_processor(file_path: str) -> tuple[str, dict]: ...

chunker = DocumentChunker(lang="en", processor_registry=registry)
Enter fullscreen mode Exit fullscreen mode

Share a single registry across chunkers only when you actually want them to share.

Custom sentence splitters are gone

The custom_splitter_registry from v2 is removed in v3.0.0, along with the custom_splitters constructor parameter that preceded it. SentenceSplitter now always uses its built-in language handlers, falling back to a universal rule-based splitter for unsupported languages.

If you were relying on a custom splitter for a specific language, open a feature request or split that language before passing the text to chunklet.

The offset removal

The offset parameter used to skip the first N sentences before chunking. It's gone from DocumentChunker, PlainTextChunker, and the CLI. No drop-in replacement, but there's a route.

yasbd, already a core dependency, exposes sentence boundaries through BoundaryDetector.detect(). Slice the text at one of those boundaries and pass the slice in:

from yasbd import BoundaryDetector

sentences_to_skip = 5
offsets = list(BoundaryDetector(lang="en").detect(text))
start = offsets[sentences_to_skip - 1] if sentences_to_skip else 0
chunks = chunker.chunk_text(text[start:])
Enter fullscreen mode Exit fullscreen mode

detect() returns the end position of each sentence, so the boundary after the fifth sentence is offsets[4]. The variable is named sentences_to_skip on purpose: it counts sentences, not characters.

This splits sentences twice, once in your code and once inside the chunker, so it costs more than offset did. If you were only dropping a fixed preamble, slicing at a boundary you choose is cheaper and does the same thing.

Migrating in five minutes

There's a script that scans your project for v1/v2 patterns:

curl -O https://raw.githubusercontent.com/speedyk-005/chunklet-py/main/audit_migration.py
python audit_migration.py /path/to/your/project
Enter fullscreen mode Exit fullscreen mode

Then the short checklist:

  1. Install the extras you need. [lang-detect] if you use lang="auto", [indic] for Indic languages, [self-tuning] for SelfTuningChunker.
  2. Move sizing arguments to the constructor.
  3. Swap the global registry for a CustomProcessorRegistry() you create.
  4. Handle offset with the yasbd route above, or slice at a boundary you choose.
  5. Rename the old extras: [auto] is [lang-detect], [structured-document] and [document] are [struct-doc], and [visualization] is [viz]. [self-tuning] is new.

If you were already on chunk_text / chunk_texts / split_text and passing specific language codes, most of this doesn't affect you. The full migration guide covers v1 and v2 readers separately.

What v3.0.0 means going forward

The 3.x line is about the self-tuning chunker and keeping the install small. The core library still handles text, documents, and code without the extras, and anything that pulls in a heavy dependency stays opt-in.

The API is more explicit than v2. Constraints are set once, registries are scoped to the chunker that uses them, and lang is required. Every change trades a little convenience at upgrade time for less to think about later.

If you're building RAG pipelines on mixed corpora and tired of tuning max_tokens per source, SelfTuningChunker is the piece to look at. If you just need sentence splitting, SentenceSplitter still provides the same core functionality, with lang now configured on the constructor.

pip install chunklet-py -U
Enter fullscreen mode Exit fullscreen mode

Issues and feedback on GitHub. If the checker missed something in your migration, that's worth filing.

For full documentation, see https://speedyk-005.github.io/chunklet-py/latest/.

Top comments (0)