<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: belcore</title>
    <description>The latest articles on DEV Community by belcore (@member_b8352c00).</description>
    <link>https://dev.to/member_b8352c00</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146123%2F025146ca-4251-4716-907e-4c4ae71824b3.png</url>
      <title>DEV Community: belcore</title>
      <link>https://dev.to/member_b8352c00</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/member_b8352c00"/>
    <language>en</language>
    <item>
      <title>Token savings depend on what you count: three numbers from the same benchmark runs</title>
      <dc:creator>belcore</dc:creator>
      <pubDate>Mon, 28 Sep 2026 04:48:48 +0000</pubDate>
      <link>https://dev.to/member_b8352c00/token-savings-depend-on-what-you-count-three-numbers-from-the-same-benchmark-runs-4fje</link>
      <guid>https://dev.to/member_b8352c00/token-savings-depend-on-what-you-count-three-numbers-from-the-same-benchmark-runs-4fje</guid>
      <description>&lt;p&gt;&lt;em&gt;Disclosure: I'm affiliated with Belcore, a memory and context layer for LLM apps. This is a measurement post. The numbers, their limits, and one result that did not hold up are all below.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "tokens saved" is slippery
&lt;/h2&gt;

&lt;p&gt;Chat and agent setups usually keep context by resending history on every call, so input tokens per call grow with the session. A memory layer replaces that with a selected context. How large the saving looks depends on what you compare, so here are the same runs counted three ways.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Dataset: LongMemEval_S, 500 questions, about 109k tokens of prior conversation each (roughly 500 turns).&lt;/li&gt;
&lt;li&gt;Answer model: gpt-5. Fixed seed, frozen harness.&lt;/li&gt;
&lt;li&gt;Tokenizer for context counts: o200k_base.&lt;/li&gt;
&lt;li&gt;Runs: August 2026.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Three numbers
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6h22wo2hs093w3wzytij.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6h22wo2hs093w3wzytij.png" alt=" " width="800" height="336"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What is counted&lt;/th&gt;
&lt;th&gt;Tokens per call&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Full transcript, uncut&lt;/td&gt;
&lt;td&gt;109,079&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assembled memory context&lt;/td&gt;
&lt;td&gt;6,671&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Billed input per call (27-question subset, includes prompt and question)&lt;/td&gt;
&lt;td&gt;9,908&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ul&gt;
&lt;li&gt;Context tokens: &lt;strong&gt;-93.9%&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Billed input on that subset: about &lt;strong&gt;-91%&lt;/strong&gt;. The subset was hand-picked, so treat this one as indicative only.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What this does not show
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxczuz3s0pc39fv1xumo2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxczuz3s0pc39fv1xumo2.png" alt=" " width="800" height="625"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Nothing here says anything about answer quality. Accuracy is a separate metric and I'm not quoting one in this post.&lt;/li&gt;
&lt;li&gt;Cost per successful task, including retries and fixes, is the number that matters for a real bill. I haven't measured it yet.&lt;/li&gt;
&lt;li&gt;LongMemEval_S is a public benchmark, not production traffic.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A result that did not hold up
&lt;/h2&gt;

&lt;p&gt;We tried a multi-pass retrieval step that re-checks and re-retrieves before answering. On 27 questions hand-picked for missing evidence it went from 16 to 20 correct. On a random sample of 103 the attributable effect was +2, inside a ±5.2 noise bar, with one case where extra evidence turned a correct aggregation answer wrong. It also cost about 2.9x the input tokens and 2.6x the latency on the subset. We are not shipping it as an improvement.&lt;/p&gt;

&lt;p&gt;Two more limits: in the wrong answers we audited, 38% involved a gold label that was wrong or defensible either way, and n = 499 cannot resolve small effects.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;There is a free measurement demo on the site: &lt;a href="https://belcore.xyz/?utm_source=devto&amp;amp;utm_campaign=token-counting" rel="noopener noreferrer"&gt;belcore.xyz&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Questions about the method are welcome, especially the ones that make the numbers look worse.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>programming</category>
    </item>
    <item>
      <title>Same benchmark runs, three "tokens saved" numbers: 109,079 (full transcript), 6,671 (assembled memory context, -93.9%), 9,908 (billed input, 27-q subset). LongMemEval_S. Accuracy and cost per successful task not shown. Affiliated with Belcore.</title>
      <dc:creator>belcore</dc:creator>
      <pubDate>Mon, 28 Sep 2026 00:40:08 +0000</pubDate>
      <link>https://dev.to/member_b8352c00/same-benchmark-runs-three-tokens-saved-numbers-109079-full-transcript-6671-assembled-1d6e</link>
      <guid>https://dev.to/member_b8352c00/same-benchmark-runs-three-tokens-saved-numbers-109079-full-transcript-6671-assembled-1d6e</guid>
      <description></description>
    </item>
  </channel>
</rss>
