DEV Community

Sharon Basovich
Sharon Basovich

Posted on Fully Autonomous

Claim Ledger gives a project review a paper trail

What I Built

Claim Ledger is a browser-based review desk for a teammate who does an independent second pass on my projects. A handoff can include a README, setup notes, test records and known limitations. Claim Ledger helps connect a specific claim to the passages worth checking, then keeps the review decision with that evidence.

Fully Autonomous: AI agents researched, implemented, tested and drafted this project under my direction. The independent QA described here was automated. My teammate has not tried the tool yet, so I have no recipient feedback or measured time savings to report.

The workflow is straightforward: load documents, enter or extract claims, find candidate receipts, inspect their source context, record a verdict and export the handoff. Each receipt includes its filename, line range, retrieval mode, score and content fingerprints. The reviewer supplies the verdict and note.

A README assertion, a supplied test record and an actual reviewer observation have different meanings. The interface labels those distinctions. A similar passage gives someone a place to investigate; it doesn't establish that a feature works.

Demo

Try Claim Ledger. No account or API key is needed. Start with Load fictional example packet, then enable on-device AI ranking or use the labelled keyword mode. The included TrailMend, PantryPal and Inkwell documents are synthetic examples.

Try “Can I use the trail planner on a plane with no signal?” TrailMend's README discusses downloaded routes, while its limitations document excludes offline route planning and GPX import. Open the source context before deciding what the broad phrase “works offline” should mean.

A separate automated check exercised a deliberately small evidence-change scenario: a claim said an import limit was 200 items; its source was replaced with a version saying 210. The repaired app invalidated the earlier confirmation through replacement, reranking and reload. The old note remained, labelled as having been recorded against earlier evidence. Explicitly choosing Confirmed again records a new review.

The review can be exported as Markdown or JSON. These are review records, not certificates of truth.

A supplementary AI review exercised nine frozen, AI-authored public-document cases on the deployed app. All 11 supplied excerpts and their exported receipts retained exact text and manually entered provenance. The stale 48-second demo description still ranked first in the version-check case; reviewer judgment and a Follow-up note were needed. With only one to four chunks per case and expected guidance visible to the AI reviewer, this is workflow evidence, not accuracy or human validation.

Use unique filenames for different versions or projects: same-named sources replace one another. An initial harness run made this mistake; both its invalid outcomes and the corrected filename run were retained.

Code

View the application source on GitHub. The application is Apache-2.0 licensed. Source, tests, benchmark cases, raw results and correction history are included. Model and runtime binaries are distributed separately from the small source archive.

How I Built It

The retrieval core combines BM25 keyword search with the open mixedbread-ai/mxbai-embed-xsmall-v1 embedding model. The pinned revision is e6ac24e5d6efb8782b59de1647b3ececb4ece94e, using q8 ONNX weights through Transformers.js 4.3.0 and CPU/WASM in a Web Worker. The model is Apache-2.0 licensed.

The model ranks candidate excerpts. It doesn't write answers or set verdicts. Hybrid mode combines lexical and semantic rankings with reciprocal-rank fusion. Keyword mode remains available without enabling the model.

Independent automated QA caught problems that mattered to the workflow: older receipts could pick up replacement text, confirmation could survive changed evidence, and multiline Markdown excerpts could escape their intended formatting. Those cases are now covered by targeted regression checks. A separate reviewer rebuilt the repaired source archive and matched its compiled JavaScript byte-for-byte to the deployed application. Typecheck, build, 48 unit tests and 16 end-to-end tests passed in that review. The browser regression runs used the byte-identical local build.

The benchmark has a less tidy story, and the source archive keeps it. It uses 48 AI-created questions over fictional documents: 40 answerable and eight no-answer cases. On the preserved exploratory run, BM25 retrieved the labelled passage first for 20/40 answerable questions; semantic and hybrid retrieval each did so for 23/40. All three methods correctly abstained on 0/8 no-answer questions. Keyword search outperformed semantic search on the negation category.

These are exploratory results. Chunking and the threshold were changed after inspecting results, on the same set. The original first raw output was overwritten and lost. A later rerun and an older differently chunked result remain separately labelled, alongside the correction manifest. The repaired release preserves those benchmark files unchanged. This is not a held-out evaluation or evidence of improved review quality.

Why Does Open Innovation Matter?

An openly licensed model makes this narrow use of AI inspectable and replaceable. Developers can compare semantic search against the keyword baseline, inspect every failed query, and decide whether the model helps their documents.

Inference processes the documents in the browser. The host still serves the site, model and runtime assets and receives ordinary network metadata. In the tested canary scenario, the reviewer observed same-origin GET requests and no transmission of the canary text. That observation is scoped to the tested flow, not a universal privacy guarantee.

Limits and Next Steps

Only Chromium has been tested; Firefox, Safari and real phones remain untested. The model-loading progress display can sit at zero while WASM initializes. An immediate reload can lose the most recent approval, which fails safely by losing the approval rather than inventing one. Deleting and re-adding byte-identical evidence can restore the earlier confirmation while also showing a stale indicator. One split receipt covers only part of its cited line, and JSON exports omit a claim's original source-line field.

The first trial with the intended teammate is still ahead. For now, the result is a working, inspectable claim-to-evidence review workflow with automated checks and visible limitations.

Top comments (0)