<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Shaurya Mehta</title>
    <description>The latest articles on DEV Community by Shaurya Mehta (@shaurya_mehta).</description>
    <link>https://dev.to/shaurya_mehta</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4152880%2F6db3e205-d4e9-4c84-a233-56b55beeb965.jpg</url>
      <title>DEV Community: Shaurya Mehta</title>
      <link>https://dev.to/shaurya_mehta</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shaurya_mehta"/>
    <language>en</language>
    <item>
      <title>Receipts: I Built a Local AI Verifier for My Friend's Project Reports</title>
      <dc:creator>Shaurya Mehta</dc:creator>
      <pubDate>Mon, 05 Oct 2026 00:56:25 +0000</pubDate>
      <link>https://dev.to/shaurya_mehta/receipts-i-built-a-local-ai-verifier-for-my-friends-project-reports-1614</link>
      <guid>https://dev.to/shaurya_mehta/receipts-i-built-a-local-ai-verifier-for-my-friends-project-reports-1614</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft6plwo6lv0pgs5b3gfqx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft6plwo6lv0pgs5b3gfqx.png" alt=" " width="799" height="393"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Why I Built It for a Friend
&lt;/h2&gt;

&lt;p&gt;My friend was working on project documentation where the README, configuration, results, and datasets could easily drift apart.&lt;br&gt;&lt;br&gt;
Instead of manually checking every number and configuration value, &lt;strong&gt;Receipts&lt;/strong&gt; provides a quick, auditable report showing which claims are supported and which need attention.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Local Open-Source AI?
&lt;/h2&gt;

&lt;p&gt;Project folders can contain coursework, datasets, configuration files, and other information that shouldn't automatically be uploaded to a cloud AI provider.  &lt;/p&gt;

&lt;p&gt;Receipts uses &lt;strong&gt;Gemma locally through Ollama&lt;/strong&gt;, so:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Project data stays on the machine.&lt;/li&gt;
&lt;li&gt;No cloud AI API is required.&lt;/li&gt;
&lt;li&gt;There is no per-request API cost.&lt;/li&gt;
&lt;li&gt;The model can be changed/configured locally.&lt;/li&gt;
&lt;li&gt;The extraction process can be reproduced.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Most importantly:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The AI is not the final authority.&lt;/p&gt;




&lt;h2&gt;
  
  
  Four Possible Verdicts
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;VERIFIED&lt;/td&gt;
&lt;td&gt;Evidence supports the claim.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CONFLICT&lt;/td&gt;
&lt;td&gt;Evidence contradicts the claim.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AMBIGUOUS&lt;/td&gt;
&lt;td&gt;Multiple plausible pieces of evidence exist.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UNVERIFIABLE&lt;/td&gt;
&lt;td&gt;Suitable evidence could not be found.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each result keeps &lt;strong&gt;provenance&lt;/strong&gt; so the user can understand where the claim came from and what evidence was used.&lt;/p&gt;




&lt;h2&gt;
  
  
  Technical Stack
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Python
&lt;/li&gt;
&lt;li&gt;Ollama
&lt;/li&gt;
&lt;li&gt;Gemma
&lt;/li&gt;
&lt;li&gt;Python standard library
&lt;/li&gt;
&lt;li&gt;Local HTTP inference
&lt;/li&gt;
&lt;li&gt;HTML report generation
&lt;/li&gt;
&lt;li&gt;unittest
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Receipts does &lt;strong&gt;not&lt;/strong&gt; execute project code, notebooks, or scripts while inspecting a project.&lt;/p&gt;




&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;Receipts generates a self-contained &lt;strong&gt;HTML verification report&lt;/strong&gt; containing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;verification results
&lt;/li&gt;
&lt;li&gt;verdicts
&lt;/li&gt;
&lt;li&gt;claim information
&lt;/li&gt;
&lt;li&gt;evidence
&lt;/li&gt;
&lt;li&gt;provenance
&lt;/li&gt;
&lt;li&gt;summary information
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The controlled demo contains examples of all four verdict types:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;VERIFIED · CONFLICT · AMBIGUOUS · UNVERIFIABLE&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Source Code
&lt;/h2&gt;

&lt;p&gt;GitHub: &lt;a href="https://github.com/sm4006/Receipts" rel="noopener noreferrer"&gt;https://github.com/sm4006/Receipts&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Verification Report
&lt;/h2&gt;

&lt;p&gt;The CLI generates a self-contained &lt;code&gt;report.html&lt;/code&gt; file containing the verification results and provenance.&lt;/p&gt;




&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;I used &lt;strong&gt;GitHub Copilot&lt;/strong&gt; as a coding agent throughout development.&lt;br&gt;&lt;br&gt;
I worked from a written specification, implemented the system incrementally, reviewed each stage, added regression tests, and performed adversarial testing before moving to the next stage.&lt;/p&gt;

&lt;p&gt;The main lesson from building this was that &lt;strong&gt;adding an LLM isn't enough&lt;/strong&gt;.&lt;br&gt;&lt;br&gt;
You have to define exactly what the model is allowed to do.&lt;/p&gt;

&lt;p&gt;For Receipts, that boundary is simple:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI proposes the claims.
&lt;/li&gt;
&lt;li&gt;Deterministic code makes the decision.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes the result easier to test, audit, and trust.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Testing
&lt;/h2&gt;

&lt;p&gt;Final test suite covered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;All verdict types
&lt;/li&gt;
&lt;li&gt;normalization
&lt;/li&gt;
&lt;li&gt;CSV/JSON/notebook parsing
&lt;/li&gt;
&lt;li&gt;malformed input
&lt;/li&gt;
&lt;li&gt;invalid AI output
&lt;/li&gt;
&lt;li&gt;fabricated evidence
&lt;/li&gt;
&lt;li&gt;Ollama failures
&lt;/li&gt;
&lt;li&gt;path traversal
&lt;/li&gt;
&lt;li&gt;unsupported files
&lt;/li&gt;
&lt;li&gt;security boundaries
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Result:&lt;/strong&gt; 85 tests → OK&lt;/p&gt;




&lt;h2&gt;
  
  
  The Part I Actually Like
&lt;/h2&gt;

&lt;p&gt;The boundary between &lt;strong&gt;probabilistic AI&lt;/strong&gt; and &lt;strong&gt;deterministic Python&lt;/strong&gt;.  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LLMs → good at structuring messy docs.
&lt;/li&gt;
&lt;li&gt;Python → reliable for factual verification.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;AI → understand&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Python → verify&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Why "Receipts"?
&lt;/h2&gt;

&lt;p&gt;Because when a project says:&lt;br&gt;&lt;br&gt;
&lt;em&gt;"Our model achieved 94.2% accuracy."&lt;/em&gt;  &lt;/p&gt;

&lt;p&gt;I want the project to have the &lt;strong&gt;receipts&lt;/strong&gt; — the evidence behind it.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Learned
&lt;/h2&gt;

&lt;p&gt;Hardest part: deciding &lt;strong&gt;what the LLM was allowed to do&lt;/strong&gt;.  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;More authority → harder to trust.
&lt;/li&gt;
&lt;li&gt;More deterministic Python → easier to test.
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The Handover
&lt;/h2&gt;

&lt;p&gt;Receipts was built for students/researchers who want a &lt;strong&gt;second check&lt;/strong&gt; before submission.&lt;br&gt;&lt;br&gt;
Point it at the folder.&lt;br&gt;&lt;br&gt;
It checks.&lt;br&gt;&lt;br&gt;
If it can’t prove something, it says so.&lt;/p&gt;




&lt;h2&gt;
  
  
  Hacktoberfest Partner Categories
&lt;/h2&gt;

&lt;p&gt;Receipts directly qualifies for the &lt;strong&gt;Hacktoberfest $100 partner categories&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma&lt;/strong&gt; → Gemma 3 4B via Ollama for claim extraction.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of GitHub Copilot&lt;/strong&gt; → Copilot coding agent used throughout incremental development, testing, debugging, and refinement.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both categories emphasize &lt;strong&gt;open-source AI&lt;/strong&gt; and &lt;strong&gt;Copilot-assisted development&lt;/strong&gt;, which were central to how Receipts was built.&lt;/p&gt;




&lt;h2&gt;
  
  
  Future Vision
&lt;/h2&gt;

&lt;p&gt;Extensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;experiment reports
&lt;/li&gt;
&lt;li&gt;research documentation
&lt;/li&gt;
&lt;li&gt;configuration drift detection
&lt;/li&gt;
&lt;li&gt;reproducibility checks
&lt;/li&gt;
&lt;li&gt;CI verification
&lt;/li&gt;
&lt;li&gt;submission pre-flight checks
&lt;/li&gt;
&lt;li&gt;larger project evidence graphs
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Principle stays the same:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
AI extracts. Deterministic code verifies.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
