DEV Community

Shaurya Mehta
Shaurya Mehta

Posted on

Receipts: I Built a Local AI Verifier for My Friend's Project Reports


Why I Built It for a Friend

My friend was working on project documentation where the README, configuration, results, and datasets could easily drift apart.

Instead of manually checking every number and configuration value, Receipts provides a quick, auditable report showing which claims are supported and which need attention.


Why Local Open-Source AI?

Project folders can contain coursework, datasets, configuration files, and other information that shouldn't automatically be uploaded to a cloud AI provider.

Receipts uses Gemma locally through Ollama, so:

  • Project data stays on the machine.
  • No cloud AI API is required.
  • There is no per-request API cost.
  • The model can be changed/configured locally.
  • The extraction process can be reproduced.

Most importantly:

The AI is not the final authority.


Four Possible Verdicts

Verdict Meaning
VERIFIED Evidence supports the claim.
CONFLICT Evidence contradicts the claim.
AMBIGUOUS Multiple plausible pieces of evidence exist.
UNVERIFIABLE Suitable evidence could not be found.

Each result keeps provenance so the user can understand where the claim came from and what evidence was used.


Technical Stack

  • Python
  • Ollama
  • Gemma
  • Python standard library
  • Local HTTP inference
  • HTML report generation
  • unittest

Receipts does not execute project code, notebooks, or scripts while inspecting a project.


Demo

Receipts generates a self-contained HTML verification report containing:

  • verification results
  • verdicts
  • claim information
  • evidence
  • provenance
  • summary information

The controlled demo contains examples of all four verdict types:

VERIFIED · CONFLICT · AMBIGUOUS · UNVERIFIABLE


Source Code

GitHub: https://github.com/sm4006/Receipts


Verification Report

The CLI generates a self-contained report.html file containing the verification results and provenance.


How I Built It

I used GitHub Copilot as a coding agent throughout development.

I worked from a written specification, implemented the system incrementally, reviewed each stage, added regression tests, and performed adversarial testing before moving to the next stage.

The main lesson from building this was that adding an LLM isn't enough.

You have to define exactly what the model is allowed to do.

For Receipts, that boundary is simple:

  • AI proposes the claims.
  • Deterministic code makes the decision.

That makes the result easier to test, audit, and trust.


The Testing

Final test suite covered:

  • All verdict types
  • normalization
  • CSV/JSON/notebook parsing
  • malformed input
  • invalid AI output
  • fabricated evidence
  • Ollama failures
  • path traversal
  • unsupported files
  • security boundaries

Result: 85 tests → OK


The Part I Actually Like

The boundary between probabilistic AI and deterministic Python.

  • LLMs → good at structuring messy docs.
  • Python → reliable for factual verification.

So:

AI → understand

Python → verify


Why "Receipts"?

Because when a project says:

"Our model achieved 94.2% accuracy."

I want the project to have the receipts — the evidence behind it.


What I Learned

Hardest part: deciding what the LLM was allowed to do.

  • More authority → harder to trust.
  • More deterministic Python → easier to test.

The Handover

Receipts was built for students/researchers who want a second check before submission.

Point it at the folder.

It checks.

If it can’t prove something, it says so.


Hacktoberfest Partner Categories

Receipts directly qualifies for the Hacktoberfest $100 partner categories:

  • Best Use of Gemma → Gemma 3 4B via Ollama for claim extraction.
  • Best Use of GitHub Copilot → Copilot coding agent used throughout incremental development, testing, debugging, and refinement.

Both categories emphasize open-source AI and Copilot-assisted development, which were central to how Receipts was built.


Future Vision

Extensions:

  • experiment reports
  • research documentation
  • configuration drift detection
  • reproducibility checks
  • CI verification
  • submission pre-flight checks
  • larger project evidence graphs

Principle stays the same:

AI extracts. Deterministic code verifies.

Top comments (0)