DEV Community

Cover image for How Many Tokens Is That Elasticsearch Hit? A Reproducible RAG Compression Benchmark
Hermann Samimi
Hermann Samimi

Posted on

How Many Tokens Is That Elasticsearch Hit? A Reproducible RAG Compression Benchmark

This is a follow-up to my earlier post introducing jtoken. That one was the pitch. This one is the actual measurement — the benchmark I run before claiming any savings number, and how you can run it on your own payloads.

The question

When your RAG pipeline pulls documents into a prompt, how much of the context window is you asked for this vs syntax overhead? And when someone (including me) claims "X% fewer tokens," how do you check that claim instead of trusting it?

Two disciplines fixed both problems for me:

  1. Measure with a real tokenizer, not characters. tiktoken (cl100k_base) is seconds to install and its counts are what GPT-4o-class models actually see.
  2. Prove nothing was lost. A compression benchmark for prompts is meaningless if the "compressed" document isn't recoverable.

The methodology

The bundled script — benchmarks/benchmark.py — does this, per payload shape:

  • Generates 50 realistic documents (ES e-commerce hits, Mongo extended-JSON activity docs, deeply-nested SaaS API events)
  • Encodes each document individually (as they'd be injected into a prompt), joins the results
  • Counts tokens for pretty JSON (json.dumps(indent=2)) vs the jtoken representation with tiktoken's cl100k_base
  • Verifies every single payload round-trips exactly: encode_document → decode_document → assert restored == original
python3 benchmarks/benchmark.py
Enter fullscreen mode Exit fullscreen mode

Runs in about a second, cold process, ~0.03 ms per document to encode. Prints a table you can paste anywhere.

The numbers

Payload (50 docs) JSON (pretty) jtoken Saved
Elasticsearch hits 12,038 10,671 11.4%
MongoDB documents (extended-JSON) 9,437 7,630 19.1%
Nested API events 12,773 11,083 13.2%
Total 34,248 29,384 14.2%

What structure actually compresses

Reading these numbers is more instructive than the numbers themselves:

  • My payloads are a floor, not a ceiling. I deliberately wrote prose-heavy documents (product descriptions, log messages with 20-90 character sentences). Prose compresses poorly in both representations — it caps the ceiling. Real-world machine JSON (monitoring events, log payloads, webhook bodies with many identical keys) compresses substantially better. My earlier post's 62%-single-hit number came from a large, highly-structured document; treat claims like that as the best case, not the expectation.
  • Repetition is the lever. Every repeated boolean is nearly free once jtoken collapses it into a shared trues:/falses: line. A JSON doc with 15 true values across 50 docs pays for the word trues once and reuses it as one token.
  • Where jtoken wins vs field selection: if you can drop fields, drop them — that beats any format change. jtoken wins when you can't (you need the data back, an agent will read it programmatically later, or an audit asks what the model actually saw).

A real bug this benchmark caught

The round-trip assertion isn't decoration. Writing this benchmark caught a real bug in jtoken 0.3.5: {"$date": {"$numberLong": "1788220800000"}} (epoch-millis Mongo dates) encoded fine but decoded back as {"$date": "1788220800000"} — the $numberLong wrapper was lost. The fix (a distinct datetime_long typed value in the normalization context) shipped in 0.3.6, with regression tests.

The same round-trip test then surfaced a packaging bug conda-forge's Windows CI caught: a stray scratch module inside the package had import-time side effects (opening a hardcoded file path!) that crashed test collection on Windows only. Both bugs gone in 0.3.6. That's the argument for verification-in-tests over trust: the moment you automate the check, you find the bug you'd have shipped.

Use it from LangChain directly

pip install langchain-jtoken

from langchain_jtoken import JSONTokenDocumentTransformer

compressed = JSONTokenDocumentTransformer().transform_documents(retrieved_docs)
# JSON page_content compresses in place, metadata preserved,
# non-JSON documents pass through untouched
Enter fullscreen mode Exit fullscreen mode

Run it on YOUR payloads

The interesting number isn't in this post — it's the one from your data. Two lines:

git clone https://github.com/HermannSamimi/jtoken && cd jtoken
python3 benchmarks/benchmark.py
Enter fullscreen mode Exit fullscreen mode

Or point the functions at your own documents and print the same table. If you get a number much worse than 10% on machine-generated JSON, that's a bug I want to hear about — open an issue at https://github.com/HermannSamimi/jtoken/issues.

For agents/LLM tooling that want to discover the library programmatically: there's an llms.txt and llms-full.txt at the repo root (the emerging convention for LLM-readable package docs).

TL;DR: lossless JSON compression buys you ~14% on realistic mixed payloads, ~19% on Mongo documents, more on repetitive machine JSON, zero data loss, one line per document. Measure your own — the script takes a second.

Top comments (0)