<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Arjun Shah</title>
    <description>The latest articles on DEV Community by Arjun Shah (@arjunkshah).</description>
    <link>https://dev.to/arjunkshah</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4004457%2Fef60576e-f654-4841-bdc1-a8f82ef6fd73.jpg</url>
      <title>DEV Community: Arjun Shah</title>
      <link>https://dev.to/arjunkshah</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/arjunkshah"/>
    <language>en</language>
    <item>
      <title>I ported my coding-agent benchmark to Kaggle, and the first bugs I found were mine</title>
      <dc:creator>Arjun Shah</dc:creator>
      <pubDate>Sat, 10 Oct 2026 02:18:44 +0000</pubDate>
      <link>https://dev.to/arjunkshah/i-ported-my-coding-agent-benchmark-to-kaggle-and-the-first-bugs-i-found-were-mine-1a5k</link>
      <guid>https://dev.to/arjunkshah/i-ported-my-coding-agent-benchmark-to-kaggle-and-the-first-bugs-i-found-were-mine-1a5k</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;I maintain &lt;a href="https://github.com/arjunkshah12345-hash/cli-bench" rel="noopener noreferrer"&gt;cli-bench&lt;/a&gt;, an open benchmark for coding agents that work in a terminal. Each task is a small repository with a bug to fix, a feature to build, or a function to speed up, plus a verifier script that decides pass or fail. Nothing is graded by another model. Either the tests pass and the extra checks hold, or the run fails.&lt;/p&gt;

&lt;p&gt;For this challenge I ported six cli-bench tasks to Kaggle Benchmarks:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Kaggle task&lt;/th&gt;
&lt;th&gt;What the model has to do&lt;/th&gt;
&lt;th&gt;What the verifier checks&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cb-debug-wrong-answer&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Find the bug in a small stats library&lt;/td&gt;
&lt;td&gt;The shipped pytest suite passes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cb-refactor-deadcode&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Delete the three unused functions and nothing else&lt;/td&gt;
&lt;td&gt;Dead functions gone, six live ones still defined, tests pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cb-feature-rate-limiter&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Write a thread-safe token bucket from an interface spec&lt;/td&gt;
&lt;td&gt;Tests for burst, refill, atomic rollback, and 100 threads at once&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cb-perf-hot-loop&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Make a pair counter at least 10x faster in pure Python&lt;/td&gt;
&lt;td&gt;Stdlib only, same signature, randomized equivalence, 10x on uniform, clustered, and gridded points&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cb-data-log-analysis&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Answer seven exact questions about an application log it never sees&lt;/td&gt;
&lt;td&gt;The model writes &lt;code&gt;solve.py&lt;/code&gt;; the task runs it and compares every answer to the log&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cb-sec-patch-xss&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Close an XSS hole in a comment board&lt;/td&gt;
&lt;td&gt;Exploit tests pass and a fresh payload renders as text&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In cli-bench, an agent gets a shell and a time budget. On Kaggle the model gets one shot: the prompt holds every file in the repo, and the model answers with whole files in &lt;code&gt;FILE: path&lt;/code&gt; blocks. The task writes those files into a temporary directory and runs the original cli-bench verifier gates on the result. A run passes only if every gate passes. If the reply rewrites a test file or an input, that write is thrown away and the run fails, the same way cli-bench treats sabotage.&lt;/p&gt;

&lt;p&gt;So the question this benchmark asks is: &lt;strong&gt;how much of a coding agent's score comes from the model reading code carefully, and how much comes from the agent loop of running tests and trying again?&lt;/strong&gt; I already had one agent run on the full suite (Codex with gpt-5.6-luna, 23 of 36 trials passed). The Kaggle version takes the loop away.&lt;/p&gt;

&lt;h3&gt;
  
  
  Before running a single model, I found two bugs in my own benchmark
&lt;/h3&gt;

&lt;p&gt;Porting meant reading every verifier line by line, and checking each task with a reference solution and a few wrong ones. Two tasks did not hold up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;sec/patch-xss&lt;/code&gt; could not be passed by a correct fix.&lt;/strong&gt; Two shipped tests assert that the words &lt;code&gt;alert&lt;/code&gt; and &lt;code&gt;onerror&lt;/code&gt; appear nowhere in the rendered page. The verifier's own probe then requires the escaped payload text, including &lt;code&gt;alert&lt;/code&gt;, to still be in the page. Escaping the comment with &lt;code&gt;html.escape&lt;/code&gt; is the textbook fix, and it fails the tests, because &lt;code&gt;&amp;amp;lt;script&amp;amp;gt;alert(...)&lt;/code&gt; still contains the word &lt;code&gt;alert&lt;/code&gt;. Deleting the text passes the tests and fails the probe. In the Codex run, two of the three trials used &lt;code&gt;html.escape&lt;/code&gt; and failed the tests, and the third stripped the text and failed the probe. I had written those failures up as the model's fault.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;data/log-analysis&lt;/code&gt; described one rule and graded another.&lt;/strong&gt; The question sheet defined &lt;code&gt;error_rate&lt;/code&gt; as the "fraction of lines with level ERROR". The verifier divides ERROR request lines by request lines, and about one log line in seven is not a request. Both of the Codex trials that failed this task were off on &lt;code&gt;error_rate&lt;/code&gt; and nothing else (0.0434 vs 0.05, 0.038 vs 0.0441). Those are exactly the numbers you get by following the text.&lt;/p&gt;

&lt;p&gt;That is 5 of the 13 failed trials in my first leaderboard run that came from the benchmark, not the model. The Kaggle versions fix both: the XSS tests now check for raw markup instead of words, and the question sheet states the rule the verifier checks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model (Kaggle slug)&lt;/th&gt;
&lt;th&gt;Why it is in the lineup&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-5.6-luna&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The same model my Codex agent run used, so one shot and agent loop can be compared directly.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude-opus-5-5-default&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Anthropic's top tier on Kaggle's model list.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemini-3.1-pro-preview&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Google's Pro tier on Kaggle's model list.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;qwen3-coder-480b-a35b-instruct&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;An open-weights model built specifically for code.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-oss-120b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;An open-weights general model you can run on your own hardware.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemini-3.7-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Not picked: Kaggle runs its default model when a task is pushed, so it came along for free.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each model ran each task once, through Kaggle's model proxy at the SDK's default temperature.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;task&lt;/th&gt;
&lt;th&gt;Opus 5.5&lt;/th&gt;
&lt;th&gt;Gemini 3.1 Pro&lt;/th&gt;
&lt;th&gt;GPT-5.6 Luna&lt;/th&gt;
&lt;th&gt;Qwen3 Coder 480B&lt;/th&gt;
&lt;th&gt;gpt-oss-120b&lt;/th&gt;
&lt;th&gt;Gemini 3.7 Flash&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;cb-debug-wrong-answer&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cb-refactor-deadcode&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cb-feature-rate-limiter&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cb-perf-hot-loop&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;fail&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cb-data-log-analysis&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cb-sec-patch-xss&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;6/6&lt;/td&gt;
&lt;td&gt;6/6&lt;/td&gt;
&lt;td&gt;6/6&lt;/td&gt;
&lt;td&gt;5/6&lt;/td&gt;
&lt;td&gt;6/6&lt;/td&gt;
&lt;td&gt;6/6&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;1. One shot was enough for these six tasks.&lt;/strong&gt; 35 of 36 runs passed. With the files in front of them and no way to run anything, every model fixed the stats bug, pruned exactly the three dead functions, wrote a token bucket that rolls back atomically and survives 100 threads, and escaped the comment board correctly. These tasks are too easy to separate frontier models in one shot. That is a result too: the difficulty I measured in cli-bench was not coming from these tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The same model did better without the agent loop.&lt;/strong&gt; GPT-5.6 Luna passed all six in one shot. As a Codex agent it passed hot-loop 1 time in 3, with a worst-case speedup of 1.2x on the gridded dataset. In one shot it wrote a spatial grid that, timed on my machine against the same verifier, ran 42.5x faster on uniform points, 26.5x on gridded, and 14.6x on clustered. The caveats are real: one run against three, one dataset seed against three, and different machines. But it is the opposite of what I expected. Having a shell did not help the agent find the fast solution, and may have pulled it toward measuring and patching a slow one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Every model reached for the same idea on hot-loop, and the only failure was invisible to the unit tests.&lt;/strong&gt; All six bucketed points into a grid. Clustered points were the worst case for every passing model (14.6x to 41x on my machine), not gridded. Qwen3 Coder's grid passed all nine shipped unit tests and still undercounted: on one random case it found 84 pairs where there are 123. Only the randomized comparison against the naive version caught it. If my verifier had stopped at the shipped tests, that bug would have scored as a pass.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Most of my first-round failures were my harness, again.&lt;/strong&gt; The first batch of Kaggle runs had 9 failures. Seven were mine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Kaggle's task runtime does not have pytest installed. My fallback test runner did not support pytest fixtures (&lt;code&gt;tmp_path&lt;/code&gt;, &lt;code&gt;monkeypatch&lt;/code&gt;), so all six models "failed" the XSS task. Every one of them had escaped the output correctly.&lt;/li&gt;
&lt;li&gt;My reply parser expected &lt;code&gt;FILE: solve.py&lt;/code&gt; and rejected gpt-oss-120b's &lt;code&gt;**FILE: solve.py**&lt;/code&gt;. I ran its script by hand against the same log and every answer was correct.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I fixed both, pushed new versions of those two tasks, and re-ran every model on them. The table shows the re-runs. The other two first-round failures were real. One is the hot-loop bug above. In the other, Qwen3 Coder's log script called &lt;code&gt;statistics.quantiles&lt;/code&gt; with a method name that does not exist and crashed. On the re-run it wrote a different script and passed. That is one run at the default temperature, so treat any single pass or fail as a sample, not a verdict.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. The corrected XSS task is passable by the textbook fix.&lt;/strong&gt; All six models used HTML escaping, and all six passed the corrected tests and the probe. In cli-bench 0.9.1, the same approach failed. That confirms the problem was the tests, not the models.&lt;/p&gt;

&lt;p&gt;What surprised me: across both versions of this benchmark, bugs on my side (two task specs, a test runner, a parser) caused more failures than the models did.&lt;/p&gt;

&lt;p&gt;What I would measure next:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Harder cli-bench tasks: the flaky-test and CI tasks, and the four "houdini" probes that check whether a model games the verifier.&lt;/li&gt;
&lt;li&gt;Three or more runs per model, so a single crash like Qwen's is visible as variance.&lt;/li&gt;
&lt;li&gt;A version that gives the model a tool to run the tests, to see whether one retry helps or, as with hot-loop, hurts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The lesson I am keeping: a benchmark that has never been run against a known-correct answer is not a benchmark yet. Every task in this port now ships with a reference solution, an empty reply, and a test-rewriting reply. A task only counts when the first passes and the other two fail, and that has to hold in the environment where the models actually run, not just on my laptop.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;Kaggle benchmark: &lt;a href="https://www.kaggle.com/benchmarks/aks1321/cli-bench-one-shot" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/aks1321/cli-bench-one-shot&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The six public tasks (each page has the full task code in its published notebook):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/aks1321/cb-debug-wrong-answer" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/aks1321/cb-debug-wrong-answer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/aks1321/cb-refactor-deadcode" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/aks1321/cb-refactor-deadcode&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/aks1321/cb-feature-rate-limiter" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/aks1321/cb-feature-rate-limiter&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/aks1321/cb-perf-hot-loop" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/aks1321/cb-perf-hot-loop&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/aks1321/cb-data-log-analysis" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/aks1321/cb-data-log-analysis&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/aks1321/cb-sec-patch-xss" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/aks1321/cb-sec-patch-xss&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;cli-bench is Apache-2.0: &lt;a href="https://github.com/arjunkshah12345-hash/cli-bench" rel="noopener noreferrer"&gt;https://github.com/arjunkshah12345-hash/cli-bench&lt;/a&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>SuperCompress is now on PyPI! pip install supercompress in 1 line</title>
      <dc:creator>Arjun Shah</dc:creator>
      <pubDate>Fri, 26 Jun 2026 19:55:27 +0000</pubDate>
      <link>https://dev.to/arjunkshah/supercompress-is-now-on-pypi-pip-install-supercompress-in-1-line-20ja</link>
      <guid>https://dev.to/arjunkshah/supercompress-is-now-on-pypi-pip-install-supercompress-in-1-line-20ja</guid>
      <description>&lt;p&gt;I just published &lt;strong&gt;SuperCompress&lt;/strong&gt; to PyPI! 🎉&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pip install supercompress&lt;/code&gt; — that's all it takes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is it?
&lt;/h2&gt;

&lt;p&gt;A tiny ~5K parameter CPU policy that scores every line of context for relevance before sending to the LLM. It keeps only what matters for the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;65% fewer tokens&lt;/strong&gt; → same answers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;100% oracle recall&lt;/strong&gt; → never drops the answer line&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;~60ms CPU latency&lt;/strong&gt; → no GPU needed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open source&lt;/strong&gt; → MIT with non-commercial clause&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Quick Start
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;supercompress

from supercompress import compress
result &lt;span class="o"&gt;=&lt;/span&gt; compress&lt;span class="o"&gt;(&lt;/span&gt;context, question&lt;span class="o"&gt;)&lt;/span&gt;
print&lt;span class="o"&gt;(&lt;/span&gt;f&lt;span class="s2"&gt;"Saved {result['kv_savings_pct']}% tokens"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Live Demo
&lt;/h2&gt;

&lt;p&gt;Try the interactive comparison tool: &lt;a href="https://supercompress.vercel.app/compare" rel="noopener noreferrer"&gt;https://supercompress.vercel.app/compare&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Or read the technical deep-dive: &lt;a href="https://dev.to/arjunkshah/how-i-built-a-prompt-compressor-that-saves-65-on-llm-costs-3m80"&gt;https://dev.to/arjunkshah/how-i-built-a-prompt-compressor-that-saves-65-on-llm-costs-3m80&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/arjunkshah/supercompress" rel="noopener noreferrer"&gt;https://github.com/arjunkshah/supercompress&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;PyPI:&lt;/strong&gt; &lt;a href="https://pypi.org/project/supercompress/" rel="noopener noreferrer"&gt;https://pypi.org/project/supercompress/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>python</category>
    </item>
    <item>
      <title>I Built a Prompt Compressor That Saves 65% on LLM Costs — Here's the Story</title>
      <dc:creator>Arjun Shah</dc:creator>
      <pubDate>Fri, 26 Jun 2026 19:45:49 +0000</pubDate>
      <link>https://dev.to/arjunkshah/i-built-a-prompt-compressor-that-saves-65-on-llm-costs-heres-the-story-2bdp</link>
      <guid>https://dev.to/arjunkshah/i-built-a-prompt-compressor-that-saves-65-on-llm-costs-heres-the-story-2bdp</guid>
      <description>&lt;p&gt;I've been working on a side project called &lt;strong&gt;SuperCompress&lt;/strong&gt; — an intelligent prompt compression system for LLMs. The idea is simple: most tokens you send to an LLM never need to be processed. They're padding, boilerplate, irrelevant context. But they still burn GPU cycles.&lt;/p&gt;

&lt;p&gt;I wanted to fix that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Working with LLM agents, I noticed something: every agent loop was sending massive context through the GPU. 10K tokens. 50K tokens. Sometimes more. Most of it was irrelevant to the specific task.&lt;/p&gt;

&lt;p&gt;Truncation (keeping head + tail) was the standard approach, but it regularly dropped critical information from the middle of the context.&lt;/p&gt;

&lt;p&gt;I thought: what if we could score each line of context for relevance BEFORE sending it to the GPU? A tiny CPU model that decides what matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Build
&lt;/h2&gt;

&lt;p&gt;The technical challenge was:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Train a lightweight policy (~5K params) that runs on CPU in under 60ms&lt;/li&gt;
&lt;li&gt;Score each line of context relative to the user's question&lt;/li&gt;
&lt;li&gt;Evict low-relevance lines while keeping answer-critical ones&lt;/li&gt;
&lt;li&gt;Ensure the compressed output preserves correct answers&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;After a lot of iteration, the results surprised even me:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Policy&lt;/th&gt;
&lt;th&gt;KV Saved&lt;/th&gt;
&lt;th&gt;Oracle Recall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Truncation&lt;/td&gt;
&lt;td&gt;65%&lt;/td&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;H2O&lt;/td&gt;
&lt;td&gt;65%&lt;/td&gt;
&lt;td&gt;98%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SuperCompress&lt;/td&gt;
&lt;td&gt;65%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;100% oracle recall at the same token savings. The policy never dropped a line the answer depended on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Environmental Angle
&lt;/h2&gt;

&lt;p&gt;Here's what hit me hardest: at 50M agent turns per day (a conservative estimate for the industry), we're wasting 100B tokens daily. That's 24K GPU hours, 1,526 tons of CO₂, 6.5M liters of cooling water. Every day.&lt;/p&gt;

&lt;p&gt;Per 1 million compressions, SuperCompress saves:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;800M tokens avoided&lt;/li&gt;
&lt;li&gt;29 kWh energy&lt;/li&gt;
&lt;li&gt;12 kg CO₂&lt;/li&gt;
&lt;li&gt;52 L cooling water&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's tiny per call. It's enormous at scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Current Status
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;✅ Working policy with 100% oracle recall&lt;/li&gt;
&lt;li&gt;✅ Benchmarks and tests (65 passing)&lt;/li&gt;
&lt;li&gt;✅ Hosted API with free tier&lt;/li&gt;
&lt;li&gt;✅ Browser demo (compresses in-browser)&lt;/li&gt;
&lt;li&gt;✅ Python client library&lt;/li&gt;
&lt;li&gt;✅ Integration guides (OpenAI, LangChain, LlamaIndex)&lt;/li&gt;
&lt;li&gt;✅ Open source (MIT)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Currently looking for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;First real users and feedback&lt;/li&gt;
&lt;li&gt;Integration partners&lt;/li&gt;
&lt;li&gt;Contributors to the open-source codebase&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try It
&lt;/h2&gt;

&lt;p&gt;Live demo: &lt;a href="https://supercompress.vercel.app" rel="noopener noreferrer"&gt;https://supercompress.vercel.app&lt;/a&gt;&lt;br&gt;
GitHub: &lt;a href="https://github.com/arjunkshah/supercompress" rel="noopener noreferrer"&gt;https://github.com/arjunkshah/supercompress&lt;/a&gt;&lt;br&gt;
Docs: &lt;a href="https://arjunkshah-supercompress-55.mintlify.app" rel="noopener noreferrer"&gt;https://arjunkshah-supercompress-55.mintlify.app&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The ask:&lt;/strong&gt; If you're building with LLMs, try compressing your next prompt. See if the answers stay the same. I'd love to hear what you think.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Now available on PyPI!&lt;/strong&gt; &lt;code&gt;pip install supercompress&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Links:&lt;/strong&gt; &lt;a href="https://github.com/arjunkshah/supercompress" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; | &lt;a href="https://pypi.org/project/supercompress/" rel="noopener noreferrer"&gt;PyPI&lt;/a&gt; | &lt;a href="https://supercompress.vercel.app" rel="noopener noreferrer"&gt;Live Demo&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>python</category>
    </item>
    <item>
      <title>SuperCompress: Cut LLM Costs by 65% Without Losing Answers</title>
      <dc:creator>Arjun Shah</dc:creator>
      <pubDate>Fri, 26 Jun 2026 19:23:33 +0000</pubDate>
      <link>https://dev.to/arjunkshah/supercompress-cut-llm-costs-by-65-without-losing-answers-2c8n</link>
      <guid>https://dev.to/arjunkshah/supercompress-cut-llm-costs-by-65-without-losing-answers-2c8n</guid>
      <description>&lt;h2&gt;
  
  
  Tweet 1
&lt;/h2&gt;

&lt;p&gt;Every LLM call burns GPU cycles on tokens that never needed to run.&lt;/p&gt;

&lt;p&gt;Padding. Boilerplate. Irrelevant context.&lt;/p&gt;

&lt;p&gt;I built SuperCompress — a tiny CPU policy that cuts 65% of tokens before inference.&lt;/p&gt;

&lt;p&gt;Open source. MIT. Free tier.&lt;/p&gt;

&lt;p&gt;supercompress.vercel.app&lt;/p&gt;

&lt;h2&gt;
  
  
  Tweet 2
&lt;/h2&gt;

&lt;p&gt;The problem is worse than most people realize.&lt;/p&gt;

&lt;p&gt;At ~50M agent turns/day:&lt;/p&gt;

&lt;p&gt;→ 100B tokens wasted daily&lt;/p&gt;

&lt;p&gt;→ 24K GPU hours&lt;/p&gt;

&lt;p&gt;→ 1,526 tons CO₂&lt;/p&gt;

&lt;p&gt;→ 6.5M L cooling water&lt;/p&gt;

&lt;p&gt;We're burning through resources on tokens that don't matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tweet 3
&lt;/h2&gt;

&lt;p&gt;How it works:&lt;/p&gt;

&lt;p&gt;1️⃣ Context + question → CPU policy (5K params)&lt;/p&gt;

&lt;p&gt;2️⃣ Every line scored for relevance to the question&lt;/p&gt;

&lt;p&gt;3️⃣ Low-scoring lines evicted&lt;/p&gt;

&lt;p&gt;4️⃣ Only essential tokens reach the GPU&lt;/p&gt;

&lt;p&gt;CPU first. GPU for what matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tweet 4
&lt;/h2&gt;

&lt;p&gt;The numbers at 35% budget:&lt;/p&gt;

&lt;p&gt;• 65% KV cache saved&lt;/p&gt;

&lt;p&gt;• 100% oracle recall (vs 25% for truncation)&lt;/p&gt;

&lt;p&gt;• ~60ms CPU latency&lt;/p&gt;

&lt;p&gt;Same answers. ⅓ the compute.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tweet 5
&lt;/h2&gt;

&lt;p&gt;Per 1 million compressions:&lt;/p&gt;

&lt;p&gt;→ 800M tokens avoided&lt;/p&gt;

&lt;p&gt;→ 29 kWh saved&lt;/p&gt;

&lt;p&gt;→ 12 kg CO₂ avoided&lt;/p&gt;

&lt;p&gt;→ 52 L cooling water saved&lt;/p&gt;

&lt;p&gt;Scale that across the industry and it's enormous.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tweet 6
&lt;/h2&gt;

&lt;p&gt;SuperCompress is:&lt;/p&gt;

&lt;p&gt;✅ Open source (MIT)&lt;/p&gt;

&lt;p&gt;✅ Free API tier&lt;/p&gt;

&lt;p&gt;✅ Python library&lt;/p&gt;

&lt;p&gt;✅ Browser demo (no install)&lt;/p&gt;

&lt;p&gt;✅ Integration guides for OpenAI/LangChain&lt;/p&gt;

&lt;p&gt;Try it: supercompress.vercel.app&lt;/p&gt;

&lt;p&gt;GitHub: github.com/arjunkshah/supercompress&lt;/p&gt;

&lt;h2&gt;
  
  
  Tweet 7
&lt;/h2&gt;

&lt;p&gt;Built this because I believe we can't scale AI by burning through what we have left.&lt;/p&gt;

&lt;p&gt;Smarter compute means more AI for everyone — without the environmental cost.&lt;/p&gt;

&lt;p&gt;Would love feedback from the community 🙏&lt;/p&gt;

&lt;h1&gt;
  
  
  LLM #AI #OpenSource #MachineLearning
&lt;/h1&gt;




&lt;p&gt;&lt;strong&gt;Links:&lt;/strong&gt; &lt;a href="https://github.com/arjunkshah/supercompress" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; | &lt;a href="https://supercompress.vercel.app" rel="noopener noreferrer"&gt;Live Demo&lt;/a&gt; | &lt;a href="https://supercompress.vercel.app/compare" rel="noopener noreferrer"&gt;Interactive Tool&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>showdev</category>
    </item>
    <item>
      <title>How I Built a Prompt Compressor That Saves 65% on LLM Costs</title>
      <dc:creator>Arjun Shah</dc:creator>
      <pubDate>Fri, 26 Jun 2026 19:15:11 +0000</pubDate>
      <link>https://dev.to/arjunkshah/how-i-built-a-prompt-compressor-that-saves-65-on-llm-costs-3m80</link>
      <guid>https://dev.to/arjunkshah/how-i-built-a-prompt-compressor-that-saves-65-on-llm-costs-3m80</guid>
      <description>&lt;h1&gt;
  
  
  How I Built a Prompt Compressor That Saves 65% on LLM Costs
&lt;/h1&gt;

&lt;p&gt;Every time you call an LLM, tokens that never needed to be processed burn GPU cycles, waste money, and strain the grid. The problem gets worse with every agent loop, every long-context RAG query, every multi-turn conversation.&lt;/p&gt;

&lt;p&gt;I built &lt;strong&gt;SuperCompress&lt;/strong&gt; — a tiny ~5K parameter CPU policy that scores every line of context for relevance before inference, keeping only what the model needs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The results?&lt;/strong&gt; 65% fewer tokens, 100% oracle recall, ~60ms latency. Open source. MIT licensed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem: LLMs Are Wasteful
&lt;/h2&gt;

&lt;p&gt;Modern LLMs process every token you give them. On long contexts (think agent logs, RAG results, codebases), most of those tokens are padding — irrelevant boilerplate that consumes KV cache space without contributing to the answer.&lt;/p&gt;

&lt;p&gt;The standard approaches don't work well:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Tokens Saved&lt;/th&gt;
&lt;th&gt;Answer Quality&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Truncation (keep head/tail)&lt;/td&gt;
&lt;td&gt;~65%&lt;/td&gt;
&lt;td&gt;~25% recall&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FIFO eviction&lt;/td&gt;
&lt;td&gt;~65%&lt;/td&gt;
&lt;td&gt;~25% recall&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;H2O&lt;/td&gt;
&lt;td&gt;~65%&lt;/td&gt;
&lt;td&gt;~98% recall&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SuperCompress&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~65%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100% recall&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At the same KV savings, SuperCompress preserves answer quality dramatically better.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture: CPU-First Eviction
&lt;/h2&gt;

&lt;p&gt;The key insight: &lt;strong&gt;you don't need a GPU to decide what a GPU should process.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌─────────────┐     ┌──────────────┐     ┌──────────┐
│  Context In  │ ──→ │  CPU Policy  │ ──→ │  GPU LLM │
│ (1,247 tok)  │     │  (5K params) │     │ (437 tok) │
└─────────────┘     └──────────────┘     └──────────┘
                          │
                          ↓
                    Score each line
                    Drop low-relevance
                    Keep answer-critical
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The policy is a lightweight neural network (~5,000 parameters) that runs entirely on CPU. It takes each line of context + the user's question, and scores how relevant that line is to answering the question. Lines below a threshold get evicted.&lt;/p&gt;

&lt;h2&gt;
  
  
  Training Approach
&lt;/h2&gt;

&lt;p&gt;The policy was trained on a dataset of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Long-form text passages (books, documentation, code)&lt;/li&gt;
&lt;li&gt;Paired with realistic user questions&lt;/li&gt;
&lt;li&gt;Ground-truth relevance labels from oracle LLM judgments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The training objective balances:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Token savings&lt;/strong&gt; — maximize KV reduction&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recall&lt;/strong&gt; — preserve lines needed for correct answers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency&lt;/strong&gt; — keep inference under 100ms on CPU&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Benchmarks
&lt;/h2&gt;

&lt;p&gt;At a fixed 35% budget (keep 35% of tokens):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csvs"&gt;&lt;code&gt;&lt;span class="k"&gt;Policy&lt;/span&gt;          &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="k"&gt;Oracle&lt;/span&gt; &lt;span class="k"&gt;Recall&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="k"&gt;Entity&lt;/span&gt; &lt;span class="k"&gt;Recall&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="k"&gt;Latency&lt;/span&gt;
&lt;span class="err"&gt;────────────────┼───────────────┼───────────────┼────────&lt;/span&gt;
&lt;span class="k"&gt;FIFO&lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="k"&gt;Truncation&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt;         &lt;span class="mf"&gt;25&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;  &lt;span class="err"&gt;|&lt;/span&gt;         &lt;span class="mf"&gt;73&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;   &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="err"&gt;~&lt;/span&gt;&lt;span class="mf"&gt;57&lt;/span&gt;&lt;span class="k"&gt;ms&lt;/span&gt;
&lt;span class="k"&gt;Summarization&lt;/span&gt;   &lt;span class="err"&gt;|&lt;/span&gt;         &lt;span class="mf"&gt;61&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;  &lt;span class="err"&gt;|&lt;/span&gt;         &lt;span class="mf"&gt;65&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;   &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="err"&gt;~&lt;/span&gt;&lt;span class="mf"&gt;63&lt;/span&gt;&lt;span class="k"&gt;ms&lt;/span&gt;
&lt;span class="k"&gt;H&lt;/span&gt;&lt;span class="mf"&gt;2&lt;/span&gt;&lt;span class="k"&gt;O&lt;/span&gt;             &lt;span class="err"&gt;|&lt;/span&gt;         &lt;span class="mf"&gt;98&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;  &lt;span class="err"&gt;|&lt;/span&gt;         &lt;span class="mf"&gt;73&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;   &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="err"&gt;~&lt;/span&gt;&lt;span class="mf"&gt;56&lt;/span&gt;&lt;span class="k"&gt;ms&lt;/span&gt;
&lt;span class="k"&gt;SuperCompress&lt;/span&gt;   &lt;span class="err"&gt;|&lt;/span&gt;        &lt;span class="mf"&gt;100&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;  &lt;span class="err"&gt;|&lt;/span&gt;         &lt;span class="mf"&gt;73&lt;/span&gt;&lt;span class="err"&gt;%&lt;/span&gt;   &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="err"&gt;~&lt;/span&gt;&lt;span class="mf"&gt;60&lt;/span&gt;&lt;span class="k"&gt;ms&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;100% oracle recall means the policy never dropped a line that the answer depended on. At the same compute savings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Environmental Impact
&lt;/h2&gt;

&lt;p&gt;Per 1 million compressions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;800M tokens avoided&lt;/strong&gt; — that's real GPU time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;29 kWh saved&lt;/strong&gt; — enough to power a home for a day&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;12 kg CO₂ avoided&lt;/strong&gt; — tiny but it adds up&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;52 L water saved&lt;/strong&gt; — datacenter cooling is thirsty&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Getting Started
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Python (in-process)
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;pip&lt;/span&gt; &lt;span class="n"&gt;install&lt;/span&gt; &lt;span class="n"&gt;git&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="n"&gt;https&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;//&lt;/span&gt;&lt;span class="n"&gt;github&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;com&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;arjunkshah&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;supercompress&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;git&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;supercompress&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;compress_context&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;compress_context&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Your long context text here...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What does this code do?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;budget_ratio&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.35&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;compressed_text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;kv_savings_pct&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;% KV saved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Hosted API (no local ML deps)
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://supercompress.vercel.app/api/v1/compress &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"X-API-Key: sc_live_YOUR_KEY"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"context":"...","query":"Summarize this"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Browser demo (no setup needed)
&lt;/h3&gt;

&lt;p&gt;Just visit &lt;a href="https://supercompress.vercel.app" rel="noopener noreferrer"&gt;supercompress.vercel.app&lt;/a&gt; and try the live demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Adaptive compression ratios (not fixed budget)&lt;/li&gt;
&lt;li&gt;Integration with LangChain/LlamaIndex as a built-in compressor&lt;/li&gt;
&lt;li&gt;Quantized policy for even lower latency&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The code is open source under MIT. Contributions welcome!&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/arjunkshah/supercompress" rel="noopener noreferrer"&gt;https://github.com/arjunkshah/supercompress&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Live demo:&lt;/strong&gt; &lt;a href="https://supercompress.vercel.app" rel="noopener noreferrer"&gt;https://supercompress.vercel.app&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Docs:&lt;/strong&gt; &lt;a href="https://arjunkshah-supercompress-55.mintlify.app" rel="noopener noreferrer"&gt;https://arjunkshah-supercompress-55.mintlify.app&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>python</category>
    </item>
  </channel>
</rss>
