<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hermann Samimi</title>
    <description>The latest articles on DEV Community by Hermann Samimi (@hermann_samimi).</description>
    <link>https://dev.to/hermann_samimi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3974796%2Fe1cdbc16-9189-40a2-9a29-8736fdd2b928.jpg</url>
      <title>DEV Community: Hermann Samimi</title>
      <link>https://dev.to/hermann_samimi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hermann_samimi"/>
    <language>en</language>
    <item>
      <title>How Many Tokens Is That Elasticsearch Hit? A Reproducible RAG Compression Benchmark</title>
      <dc:creator>Hermann Samimi</dc:creator>
      <pubDate>Tue, 29 Sep 2026 16:03:28 +0000</pubDate>
      <link>https://dev.to/hermann_samimi/how-many-tokens-is-that-elasticsearch-hit-a-reproducible-rag-compression-benchmark-2k6f</link>
      <guid>https://dev.to/hermann_samimi/how-many-tokens-is-that-elasticsearch-hit-a-reproducible-rag-compression-benchmark-2k6f</guid>
      <description>&lt;p&gt;This is a follow-up to my &lt;a href="https://dev.to/hermann_samimi/jtoken-lossless-json-compression-for-llm-prompts-4a6h"&gt;earlier post introducing jtoken&lt;/a&gt;. That one was the pitch. This one is the actual measurement — the benchmark I run before claiming any savings number, and how you can run it on your own payloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question
&lt;/h2&gt;

&lt;p&gt;When your RAG pipeline pulls documents into a prompt, how much of the context window is &lt;em&gt;you asked for this&lt;/em&gt; vs &lt;em&gt;syntax overhead&lt;/em&gt;? And when someone (including me) claims "X% fewer tokens," how do you check that claim instead of trusting it?&lt;/p&gt;

&lt;p&gt;Two disciplines fixed both problems for me:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Measure with a real tokenizer&lt;/strong&gt;, not characters. &lt;code&gt;tiktoken&lt;/code&gt; (&lt;code&gt;cl100k_base&lt;/code&gt;) is seconds to install and its counts are what GPT-4o-class models actually see.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prove nothing was lost.&lt;/strong&gt; A compression benchmark for prompts is meaningless if the "compressed" document isn't recoverable.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The methodology
&lt;/h2&gt;

&lt;p&gt;The bundled script — &lt;a href="https://github.com/HermannSamimi/jtoken/blob/main/benchmarks/benchmark.py" rel="noopener noreferrer"&gt;&lt;code&gt;benchmarks/benchmark.py&lt;/code&gt;&lt;/a&gt; — does this, per payload shape:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Generates 50 realistic documents (ES e-commerce hits, Mongo extended-JSON activity docs, deeply-nested SaaS API events)&lt;/li&gt;
&lt;li&gt;Encodes each document &lt;strong&gt;individually&lt;/strong&gt; (as they'd be injected into a prompt), joins the results&lt;/li&gt;
&lt;li&gt;Counts tokens for pretty JSON (&lt;code&gt;json.dumps(indent=2)&lt;/code&gt;) vs the jtoken representation with &lt;code&gt;tiktoken&lt;/code&gt;'s &lt;code&gt;cl100k_base&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Verifies every single payload &lt;strong&gt;round-trips exactly&lt;/strong&gt;: &lt;code&gt;encode_document&lt;/code&gt; → &lt;code&gt;decode_document&lt;/code&gt; → &lt;code&gt;assert restored == original&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 benchmarks/benchmark.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Runs in about a second, cold process, ~0.03 ms per document to encode. Prints a table you can paste anywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Payload (50 docs)&lt;/th&gt;
&lt;th&gt;JSON (pretty)&lt;/th&gt;
&lt;th&gt;jtoken&lt;/th&gt;
&lt;th&gt;Saved&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Elasticsearch hits&lt;/td&gt;
&lt;td&gt;12,038&lt;/td&gt;
&lt;td&gt;10,671&lt;/td&gt;
&lt;td&gt;11.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MongoDB documents (extended-JSON)&lt;/td&gt;
&lt;td&gt;9,437&lt;/td&gt;
&lt;td&gt;7,630&lt;/td&gt;
&lt;td&gt;19.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nested API events&lt;/td&gt;
&lt;td&gt;12,773&lt;/td&gt;
&lt;td&gt;11,083&lt;/td&gt;
&lt;td&gt;13.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;34,248&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;29,384&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14.2%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What structure actually compresses
&lt;/h2&gt;

&lt;p&gt;Reading these numbers is more instructive than the numbers themselves:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;My payloads are a floor, not a ceiling.&lt;/strong&gt; I deliberately wrote &lt;em&gt;prose-heavy&lt;/em&gt; documents (product descriptions, log messages with 20-90 character sentences). Prose compresses poorly in both representations — it caps the ceiling. Real-world machine JSON (monitoring events, log payloads, webhook bodies with many identical keys) compresses &lt;strong&gt;substantially&lt;/strong&gt; better. My earlier post's 62%-single-hit number came from a large, highly-structured document; treat claims like that as the &lt;em&gt;best case&lt;/em&gt;, not the expectation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repetition is the lever.&lt;/strong&gt; Every repeated boolean is nearly free once jtoken collapses it into a shared &lt;code&gt;trues:&lt;/code&gt;/&lt;code&gt;falses:&lt;/code&gt; line. A JSON doc with 15 &lt;code&gt;true&lt;/code&gt; values across 50 docs pays for the word &lt;code&gt;trues&lt;/code&gt; once and reuses it as one token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where jtoken wins vs field selection:&lt;/strong&gt; if you can &lt;em&gt;drop&lt;/em&gt; fields, drop them — that beats any format change. jtoken wins when you can't (you need the data back, an agent will read it programmatically later, or an audit asks what the model actually saw).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A real bug this benchmark caught
&lt;/h2&gt;

&lt;p&gt;The round-trip assertion isn't decoration. Writing this benchmark caught a real bug in jtoken 0.3.5: &lt;code&gt;{"$date": {"$numberLong": "1788220800000"}}&lt;/code&gt; (epoch-millis Mongo dates) encoded fine but decoded back as &lt;code&gt;{"$date": "1788220800000"}&lt;/code&gt; — the &lt;code&gt;$numberLong&lt;/code&gt; wrapper was lost. The fix (a distinct &lt;code&gt;datetime_long&lt;/code&gt; typed value in the normalization context) shipped in 0.3.6, with regression tests.&lt;/p&gt;

&lt;p&gt;The same round-trip test then surfaced a packaging bug conda-forge's Windows CI caught: a stray scratch module inside the package had import-time side effects (opening a hardcoded file path!) that crashed test &lt;em&gt;collection&lt;/em&gt; on Windows only. Both bugs gone in 0.3.6. That's the argument for verification-in-tests over trust: &lt;strong&gt;the moment you automate the check, you find the bug you'd have shipped.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Use it from LangChain directly
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;pip&lt;/span&gt; &lt;span class="n"&gt;install&lt;/span&gt; &lt;span class="n"&gt;langchain&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;jtoken&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langchain_jtoken&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;JSONTokenDocumentTransformer&lt;/span&gt;

&lt;span class="n"&gt;compressed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;JSONTokenDocumentTransformer&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;transform_documents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;retrieved_docs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# JSON page_content compresses in place, metadata preserved,
# non-JSON documents pass through untouched
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Run it on YOUR payloads
&lt;/h2&gt;

&lt;p&gt;The interesting number isn't in this post — it's the one from your data. Two lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/HermannSamimi/jtoken &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;jtoken
python3 benchmarks/benchmark.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or point the functions at your own documents and print the same table. If you get a number much worse than 10% on machine-generated JSON, that's a bug I want to hear about — open an issue at &lt;a href="https://github.com/HermannSamimi/jtoken/issues" rel="noopener noreferrer"&gt;https://github.com/HermannSamimi/jtoken/issues&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For agents/LLM tooling that want to discover the library programmatically: there's an &lt;code&gt;llms.txt&lt;/code&gt; and &lt;code&gt;llms-full.txt&lt;/code&gt; at the repo root (the emerging convention for LLM-readable package docs).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; lossless JSON compression buys you ~14% on realistic mixed payloads, ~19% on Mongo documents, more on repetitive machine JSON, zero data loss, one line per document. Measure your own — the script takes a second.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>rag</category>
      <category>json</category>
      <category>elasticsearch</category>
    </item>
    <item>
      <title>JTOKEN - Lossless JSON compression for LLM prompts</title>
      <dc:creator>Hermann Samimi</dc:creator>
      <pubDate>Mon, 08 Jun 2026 20:41:22 +0000</pubDate>
      <link>https://dev.to/hermann_samimi/jtoken-lossless-json-compression-for-llm-prompts-4a6h</link>
      <guid>https://dev.to/hermann_samimi/jtoken-lossless-json-compression-for-llm-prompts-4a6h</guid>
      <description>&lt;p&gt;Same data, ~35% fewer tokens. Purpose-built for RAG pipelines, AI agents, and structured prompt engineering.&lt;/p&gt;

&lt;p&gt;Hey Folks 👋&lt;/p&gt;

&lt;p&gt;I'm Hermann Samimi, a data engineer who's been building RAG pipelines professionally. At some point I started looking at what my prompts were actually made of, and noticed something embarrassing: a huge chunk of tokens were going to JSON syntax — braces, quotes, commas, &lt;code&gt;true&lt;/code&gt;, &lt;code&gt;false&lt;/code&gt;, &lt;code&gt;null&lt;/code&gt; — not the actual data I cared about.&lt;/p&gt;

&lt;p&gt;So I built jtoken.&lt;/p&gt;

&lt;p&gt;It encodes JSON into a flat, human-readable format optimized for LLM context windows. The encoding is lossless — you can always round-trip back to the original. It works on plain JSON, but also natively handles MongoDB documents (shell and extended JSON) and Elasticsearch hits, which is where I personally saw the biggest savings.&lt;/p&gt;

&lt;p&gt;Real numbers:&lt;br&gt;
→ Elasticsearch hit: 1,537 → 583 tokens (&lt;code&gt;62%&lt;/code&gt; reduction)&lt;br&gt;
→ MongoDB shell doc: 770 → 508 tokens (&lt;code&gt;34%&lt;/code&gt; reduction)&lt;br&gt;
→ Standard JSON: 617 → 503 tokens (&lt;code&gt;18%&lt;/code&gt; reduction)&lt;/p&gt;

&lt;p&gt;The format is simple enough to understand in 30 seconds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;Input:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Alice"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"active"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"verified"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"ref"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;Output&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Alice&lt;/span&gt;
&lt;span class="na"&gt;trues&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;active&lt;/span&gt;
&lt;span class="na"&gt;falses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;verified&lt;/span&gt;
&lt;span class="na"&gt;nulls&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ref&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Faga067798y15i8ovyf54.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Faga067798y15i8ovyf54.gif" alt=" " width="799" height="416"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There's also a CLI for quick stats and a token counting API that works with tiktoken or a heuristic fallback.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pip install jtoken&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;I'd love feedback — especially on edge cases, format ideas, or use cases I haven't thought of yet. What do you use to pass structured data into LLMs today?&lt;/p&gt;

&lt;p&gt;PyPI:   &lt;a href="https://pypi.org/project/jtoken/" rel="noopener noreferrer"&gt;https://pypi.org/project/jtoken/&lt;/a&gt;&lt;br&gt;
GitHub: &lt;a href="https://github.com/HermannSamimi/jtoken" rel="noopener noreferrer"&gt;https://github.com/HermannSamimi/jtoken&lt;/a&gt;&lt;br&gt;
Linkedin: &lt;a href="https://www.linkedin.com/in/hermann-samimi/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/hermann-samimi/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>python</category>
    </item>
  </channel>
</rss>
