This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built SAMJHO, a privacy-first, local AI document-understanding MVP designed to simplify confusing notices, policies, and forms into clear, actionable insights for everyday users.
Core Components Built
Local Document Pipeline (parser.py + OCR): Extracts text page-by-page from standard text PDFs and handles scanned image-based documents using Tesseract OCR.
Local AI Engine (ai.py + Gemma 3 1B): Powered entirely offline via Ollama to generate structured JSON data—including summaries, required actions, deadlines, and costs—without relying on external proprietary APIs.
Evidence Validator (evidence.py): Powers the PROVE IT feature, cross-referencing AI-generated claims directly against original document quotes to guarantee reliability and prevent hallucinations.
Streamlit Interface: A simple, user-friendly UI featuring transparent processing steps, bilingual support (English and Hindi), and verifiable source inspection.
By combining local open-weight models with strict evidence validation, SAMJHO bridges the gap between complex paperwork and real-world understanding while keeping all data completely offline.
Demo
<!-- Share a deployed link or a video demo. -->https://drive.google.com/file/d/1euuxA1SqSctYp01NZTVP9wyLy_nPUnxK/view?usp=drive_link
Code
<!-- Show us the code! You can embed a GitHub repo directly into your post. --> https://github.com/ViV1-siNgh/Samjho.git
How I Built It
<!-- Which open-source AI did you use (open-weight models, agent harnesses, frameworks, local inference), and how is your project built around it? --> I built SAMJHO using a strict local-first, modular architecture to ensure complete privacy and reliability.
First, I created a robust document parser (parser.py) using PyMuPDF to extract text page-by-page, integrating Tesseract OCR for scanned PDFs. Next, I connected the local open-weight model Gemma 3 1B via Ollama (ai.py) to process the text offline, instructing it to return clean, structured JSON without hallucinating missing details. To guarantee accuracy, I built an evidence validator (evidence.py) that powers a "PROVE IT" feature, cross-referencing AI claims against exact document quotes. Finally, I unified the pipeline into a simple Streamlit interface offering bilingual support.
Why Does Open Innovation Matter?
<!-- Why does open innovation matter for what you built? What did it make possible that a closed API wouldn't? -->Open innovation matters because complex, real-world problems—like making bureaucratic paperwork accessible to everyday people—cannot be solved in isolation. By building SAMJHO as an open-source project, open innovation drives several key benefits:
Transparency and Trust: Open-source architectures allow anyone to inspect the codebase, verify that document processing happens entirely locally, and ensure that private data is never sent to external servers.
Community Collaboration: It enables developers, designers, and domain experts to build upon existing foundations, share improvements, and adapt solutions to new languages, regions, or document formats.
Rapid Iteration: Sharing code and methodologies openly fosters quick feedback loops, turning MVPs into reliable, robust tools much faster than closed-door development.
Accessibility: Open innovation democratizes AI technology, ensuring that practical tools like evidence-backed document explainers remain free, modular, and accessible to everyone who needs them.
My Agent Session
Prize Categories
<!-- Which partner categories are you entering? List every one that applies, or remove this section. -->Best Use of Tinker
Use Thinking Machines' Tinker to fine-tune a model for a specific task, and show a clear improvement in performance, latency, or cost over a baseline.
ViV1-siNgh - https://github.com/ViV1-siNgh
Himanshu699-cyber - https://github.com/Himanshu699-cyber
Top comments (1)
use it and find what you have