Primary Tags: #devchallenge, #hacktoberfest, #ai, #opensource
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built Learning Ledger for my close friend and lab partner who spends 20+ hours a week reading academic preprints in computer science and bioinformatics.
Last semester, they experienced a researcherβs worst nightmare: weeks before submitting a thesis proposal, they found out that a primary empirical baseline cited across their work had been quietly retracted due to irreproducible data. They had hundreds of reading notes stored across Notion and markdown files, but no system in their workflow was tracking whether the scientific beliefs they had recorded months ago were still true.
When researchers and students read literature, they formulate beliefs:
- "LoRA reduces trainable parameters by 10,000x with comparable accuracy"
- "Molecule X inhibits Receptor Y in vitro"
Yet science evolves rapidly: papers get retracted, errata and corrections are issued, or the student slightly misinterprets or exaggerates the authorβs exact text.
Learning Ledger is an intelligent research memory system that remembers what you learn from scientific papers and actively alerts you when it becomes wrong. It indexes academic PDFs, lets students ask grounded questions with verbatim citations, verifies that recorded beliefs strictly entail source quotes using typed decision models (Jev 1.13), and continuously monitors literature against retractions, errata, and confidence decay so you know when to unlearn or revise.
Demo
The application runs locally with a modern Scholar Edition dashboard (http://localhost:8000) built with Swiss Modernism 2.0 design principles:
Core Features:
- ποΈ Student Scholar Interface: Editorial layout featuring Newsreader serif typography, gold leaf accents, and a subtle millimeter-grid notebook texture.
- π΄ Study Cards & Active Recall Mode: A revision toggle that hides verbatim quotes and prompts students: "Can you cite the proof?", allowing active recall testing before revealing the supporting quote.
- βοΈ Jev 1.13 Decision Gate: Evaluates quote-to-claim entailment via typed probability ($0.0 - 1.0$), ensuring students don't exaggerate findings.
- π‘οΈ Continuous 4-Layer Audit: Scans claims against a local 300MB Retraction Watch SQLite database, Crossref
update-toerrata, and field-velocity temporal decay. - π One-Click APA Citations: Copies formatted academic citations directly to the clipboard.
(Add your deployed URL or video demo link here if available)
Code
π Learning Ledger (AI)
An intelligent research memory system that remembers what you learn from scientific papers and tells you when it becomes wrong.
Live Dashboard β’ Architecture β’ Flowchart β’ Tech Stack β’ Installation β’ Git Push Guide
π The Problem
When students, graduate researchers, and engineers read scientific literature, they synthesize key findings into mental models or notes:
- "Model X reduces fine-tuning parameters by 10,000x with zero accuracy loss"
- "Molecule A inhibits Target B in vitro"
- "Algorithm Q scales quadratically rather than exponentially"
However, scientific knowledge is not permanent:
- The Retraction & Replication Crisis: Thousands of peer-reviewed papers are retracted, withdrawn, or flagged every year due to data manipulation, ethical lapses, or flawed methodologies.
- Post-Publication Errata & Amendments: Authors routinely publish corrections and errata via Crossref that alter experimental parameters, margins of error, or code baselines.
- Semantic Drift & Human Overstatement: When students summarizeβ¦
Learning Ledger is 100% open-source under the MIT License.
How I Built It
[ Research PDF ] βββΊ pypdf parser βββΊ Academic Watermark Stripper
β
Header-Based Semantic Chunking βββ
β
βΌ
ChromaDB Vector Store ββββ Gemini embedding-001 (Dense Vectors)
β
βββΊ Student Question βββΊ IBM Granite Classifier (Local Ollama / HF)
β β
β βΌ
βββΊ Grounded Retrieval βββΊ Gemini 3.8 Flash (Verbatim Citations)
β
βΌ
User records: "Claim X backed by Quote Y"
β
βΌ
Jev 1.13 Decision Gate (TypeSafe SDK via OpenRouter)
ββ Noul(claim_support) > 0.50 βββΊ Stored in SQLite Ledger (Active)
ββ Noul(claim_support) <= 0.50 βββΊ Overstatement Alert (Suspect)
β
βΌ
Continuous Multi-Source Audit (Retraction Watch + Crossref API + Cascade Decay)
-
Agent Tool Harness (
openharness-ai): All user actions (drop_paper,ask_question,learn_claim,revalidate_claims,check_doi) are defined as typed Pydantic tool schemas insrc/tools.py, enabling autonomous tool-calling and agent orchestration. -
Local Model Routing (
IBM Granite): Student questions are classified locally by IBM Granite (granite3-dense:2bvia Ollama orcnmoro/granite-question-classifieron Hugging Face). This instantly routes small talk away from the vector database, keeping the pipeline snappy and offline-friendly. -
Typed Decision Gate (
Jev 1.13viatypesafe-sdk): Instead of asking an LLM to generate unstructured prose that requires brittle regex parsing, Jev returns typed mathematical evaluations:-
Noul($0.0 - 1.0$): Determines whether the verbatim paper quote strictly supports the student's claimed finding without overstatement. -
Choice: Determines retrieval routing (local_retrieve,llm_synthesize,stop).
-
-
4-Layer Continuous Integrity Engine:
- Retraction Watch SQLite Index: Full local offline database generated from the Crossref Retraction Watch dataset (~300MB).
-
Crossref REST API: Queries the live
update-toendpoint for freshly published errata. - OpenAlex & Unpaywall: Resolves open-access PDFs, citation counts, and publication version history.
- Field-Velocity Confidence Decay: Models disciplinary aging rate ($\text{Decay} = \frac{1}{1 + \text{velocity} \times \text{age}}$), decaying fast-moving ML papers ($0.50/\text{year}$) faster than pure mathematics ($0.05/\text{year}$).
Why Does Open Innovation Matter?
In scientific education and research, academic reproducibility and trust are paramount.
Closed, proprietary AI systems treat knowledge as black-box text generation. When a commercial model cites a retracted paper or hallucinates a benchmark number, it offers zero visibility into why. Furthermore, locked APIs force students to upload sensitive or unpublished thesis research to third-party clouds.
Open innovation made Learning Ledger possible:
- Open Science Data: Crossref and Retraction Watch open datasets allow students to run millisecond-fast integrity audits completely offline without recurring API costs.
- Local Inference for Student Privacy: Running local query classification on-device (via IBM Granite and Ollama) ensures draft research questions never leave the student's machine.
-
Verifiable Probabilities Over Black Boxes: Combining open-source agent harnesses (
openharness-ai) with typed decision gates replaces "vibe-based" answers with verifiable numbers.
Prize Categories
- Main Category: Build for a Friend (Education & Student Researcher Integrity)
-
Open-Source AI Track: Best Use of Open-Source AI, Local Inference & Agent Harnesses (
openharness-ai,IBM Granite,ChromaDB)
Top comments (0)