DEV Community

Mohammed Arshad Ansari
Mohammed Arshad Ansari

Posted on Originally published at hikmahtechnologies.com

A Pile of Documents Is Not Knowledge

My research agent had read about 20,400 documents. It had concluded nothing.

Not metaphorically. The knowledge store held roughly twenty thousand rows — papers, articles, PDFs, nightly log entries — and there was nowhere in the system that said what any of it meant. Research answers were stored as aegis://research/ rows filed among ten thousand PDFs. The nightly journal wrote one dated entry into the same pile, in a format nobody could open.

This is the third of three posts about rebuilding lanes of AEGIS, my self-hosted agent platform, between 5 and 13 September. The first two were alerting and money. All three ended at the same rule, and this lane is where it's easiest to see why.

Measure use, not ingestion

Every RAG dashboard I have ever seen measures the wrong end of the pipe. Documents ingested. Chunks embedded. Index size. All of it is input, and input is the part that is trivially easy to grow.

The measurement that mattered took one join: a log of which documents were injected into which prompt, joined back to where each document came from. Over the 30 days to 12 September:

  • arXiv: 1,889 papers, 89,669 chunks — 90% of all RSS chunks. Prompts used 14 papers.
  • Across the whole corpus, arXiv PDFs were 93% of all chunks, and 78 of 10,284 had ever reached a prompt.
  • The intelligence scans had stored 371 items. Exactly one was ever used in a prompt.
  • Those same scans created 277 task-manager items in 30 days, every one auto-classified as reference material on arrival. My task list was being used as a log file.

A retention preview said the same thing in bytes. One rule — keep only the first chunk of any PDF no prompt has used in 30 days — would touch 8,297 PDFs and drop 362,153 of 500,055 chunks, freeing about 530 MB of text and 1 GB of vectors from a 4.8 GB table. The PDFs themselves stay, and nothing has been deleted: the preview is a dry run, and deleting is a separate, explicit decision.

This is not a story about arXiv being low quality. It is excellent. It is a story about a system that was optimising the metric it could see.

The filter I measured and did not ship

The obvious fix is a topic gate: only store the full text when the item matches something you care about. I built it, measured it, and made it opt-in.

Here is why. On arXiv the gate would pass 41% of papers — about 2.3× fewer chunks, for real added complexity. And on every other feed, the gate would have kept the full text of only 2 of the 10 documents a prompt actually used.

That second number killed it as a default. A filter that removes 80% of the value you can prove, to save storage you have plenty of, is not an optimisation. The measurement changed the design, which is the only reason to take a measurement.

The actual win was somewhere else entirely.


This is the first part. The full post — including the rest of the working details — is on my site: A Pile of Documents Is Not Knowledge

Top comments (1)

Collapse
 
makeyouragent profile image
MakeYourAgent •

The line that stuck is measuring the join between prompt injections and document origin, not the chunk counter.

For knowledge and support bots I log which source ids actually entered the model, then review that weekly. A help center that doubles in size while the same twelve articles keep getting injected is not getting smarter. It is getting louder.

I also treat topic gates as opt-in after that join exists. If a filter would have dropped most of the few docs that ever reached a real answer, it is not a default. Prefer a dry-run retention report for chunks that never entered a prompt over celebrating index growth. Ingestion is cheap. Proven use is the scarce signal.