Your RAG pipeline retrieves 12 chunks because the retrieval score said "maybe". Your LLM reads all of them. You pay for all of them. And the answer quality was decided by chunks 2 and 7 anyway.
On September 29, OpenAI launched the Decisions API built on Luna, and on September 15, TypeSafe launched Jev. Both are fast decision models for exactly this kind of problem: deciding what matters before the expensive model runs. But they are APIs. Your documents leave your machine, and you pay per decision.
laya-compactor does the same thing locally, for free. It scores your entire retrieval batch in a single forward pass and decides which documents stay in context and which get cut.
70% fewer input tokens
Zero answer-quality loss on SQuAD (exact match: 0.345 vs 0.345)
94.5% of gold documents retained on HotpotQA
$0 per compaction decision
How it works
One principle drives the design: delete, do not rewrite.
The tool never summarizes or paraphrases your documents. It keeps the ones that matter verbatim and removes the rest, with an auditable reason (low score or exhausted budget). The evidence stays intact for the LLM to reason over.
+--------------------------------------------------+
| Retrieval batch (12 docs, ~3,200 tokens) |
| |
| laya decision engine (one forward pass, local) |
| scores each doc: 0=irrelevant .. 3=essential |
| |
| Budget: keep top-scoring docs until limit |
| |
| Output: 4 docs, ~970 tokens, same answer |
+--------------------------------------------------+
Python API and CLI:
from laya_compactor import compact
result = compact(
question="What is the refund policy for annual plans?",
documents=retrieved_chunks,
token_budget=1000,
)
# result.kept: the documents that survived
# result.dropped: [{doc, score, reason}] for the ones cut
The benchmarks (all of them)
200 questions per dataset, BM25 retrieval, generator and blind judge powered by Z.ai's GLM-5.3-flashX:
| Dataset | Full context | Compacted | Tokens saved |
|---|---|---|---|
| SQuAD | EM 0.345, 3,214 avg tokens | EM 0.345, 973 avg tokens | 69.7% |
| HotpotQA | EM 0.230, 3,193 avg tokens | EM 0.200, 947 avg tokens | 70.3% |
On SQuAD the exact match is identical: the compactor removed noise without touching signal. On HotpotQA it dropped 3 points but kept 94.5% of the gold documents, meaning the loss came from multi-hop reasoning, not from cutting the wrong chunks. Truncation baselines (head-only, tail-only) scored worse on both.
Caveats, because they exist: CPU latency is significant (6.3 to 10 seconds p50 for a batch). Multi-hop questions remain hard. The rubric sensitivity study is committed with a hand-labeled 100-row dataset so you can see exactly where the scoring is fragile.
Integrations
Drop-in for the two frameworks you already use:
# LangChain
from langchain.retrievers import ContextualCompressionRetriever
retriever = ContextualCompressionRetriever(
base_compressor=LayaCompactor(token_budget=1000),
base_retriever=your_retriever,
)
# LlamaIndex
from laya_compactor.integrations import LayaNodePostprocessor
query_engine = index.as_query_engine(
node_postprocessors=[LayaNodePostprocessor(token_budget=1000)],
)
Try it
pip install laya-compactor
CLI included: laya-compact --question "..." --docs docs/ --budget 1000
The rest of the series
This is the second of four tools built on the same idea: local, measured decisions in LLM pipelines.
- laya-router: route prompts between cheap and frontier models, 54.9% cost reduction
- laya-triage: support ticket triage, fine-tuned from 51% to 90.5% intent accuracy
- laya-phishield: explainable phishing detection, $0 per 1,000 emails
Each one ships with its benchmarks in the README, including the numbers that hurt.
Top comments (0)