This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
PaperTrail turns a pile of application paperwork into a checklist where every tick comes with a receipt.
I built it for [FRIEND'S NAME], who is applying to [universities / scholarships] this year. [In one or two sentences, your own words: what you watched them struggle with.]
No single document is the hard part. The hard part is knowing at any moment which requirements are covered, which are missing, and which documents will expire before the deadline.
PaperTrail does five things:
- Reads the requirements. Upload the program's requirements as a PDF or paste them as text. PaperTrail extracts each requirement and its deadline, and links each one to the sentence it came from.
- Reads your documents. Upload a transcript, a test score report or a CV. Each requirement is either matched to a passage that satisfies it or marked missing, expired or needs review.
- Tracks what's left. Open items become tasks, with reminders that survive restarts. When you mark a task done, its pending reminder is cancelled.
- Answers questions with citations. Ask "Why is my English test marked for review?" and the answer quotes the score report and links to the page.
- Exports a report that separates what the AI found from what you verified yourself.
The rule I held myself to: the AI proposes, the person decides. The UI shows "Evidence found (AI)" and "Verified by you" as different states. Nothing counts as verified until a human says so.
Demo
- Create a workspace with a 2027-01-15 deadline and upload the program's requirements PDF.
- Review the extracted requirements. Each one links to its source sentence.
- Upload a transcript and an English test score report. The transcript is matched with a cited passage. The English test goes to needs review because it expires on 2026-12-12, before the deadline.
- Inject a failure while a document is processing. Attempt 1 fails and shows up in Sentry. Temporal retries, and attempt 2 succeeds.
- Set a reminder, complete the task, and watch the reminder get suppressed instead of sent.
- Ask the assistant why the English test needs review.
Code
PaperTrail
Every document. Every requirement. One clear path forward.
PaperTrail turns a pile of application paperwork into an evidence-backed checklist. Upload the requirements (a PDF or pasted text) and your supporting documents. PaperTrail extracts the requirements, matches each one to passages in your documents, flags what's missing, expired or uncertain, tracks deadlines, and sends reminders that survive restarts. Every AI assessment links to the page and passage it relied on, and nothing counts as verified until you say so.
The MVP targets university and scholarship applications. Nothing in the data model is specific to them.
browser ── React + TanStack Query ──▶ Fastify API (/api/v1) ──▶ Tiger Data / Postgres + pgvector
│ ▲
└─ start / signal ─▶ Temporal ─▶ Worker (activities)
├─ PDF parser / OCR
├─ Gemma via Ollama (or any OpenAI-compatible endpoint)
└─ webhook reminders
API + worker ── OTLP traces + errors ──▶ Sentry
| Piece |
|---|
How I Built It
The open-weight core is Gemma, running locally through Ollama:
-
gemma3:4bextracts requirements, classifies documents, assesses evidence, answers questions and reads text from images. -
embeddinggemmaproduces 768-dimensional embeddings for search.
The server reaches Ollama through an OpenAI-compatible adapter. Pointing the same code at a hosted Gemma endpoint means changing one environment variable.
Here's how the pieces connect:
browser ── React + TanStack Query ──▶ Fastify API ──▶ Postgres + pgvector (Tiger Data)
│
└─▶ Temporal ─▶ Worker
├─ PDF parser / OCR
├─ Gemma via Ollama
└─ webhook reminders
API + worker ── traces + errors ──▶ Sentry
- Temporal runs four workflows: process a document, analyze a workspace, deliver reminders on durable timers, and retry failed processing from the step that failed.
- Postgres + pgvector stores everything and also handles search. Vector search and full-text search are combined with reciprocal rank fusion in a single SQL query.
-
Sentry traces the API and the worker. Model calls show up as
gen_aispans, and activity failures are tagged with the workflow, activity and attempt.
Making a 4B model trustworthy
A 4-billion-parameter model on a laptop CPU is slow, and sometimes it is confidently wrong. I ran it against realistic documents and kept a list of the ways it failed:
-
It inverted booleans. When I asked whether a requirement was
required, it often gave the opposite answer. Asking foroptionalinstead fixed it: true only when the text actually says "optional" or "if applicable." - It garbled citation labels. Now labels come from a fixed enum, and each quote is re-attributed to the passage that actually contains it.
- It cited a transcript as a CV. The quote was real but came from the wrong kind of document. A document-class check now ensures a transcript can never satisfy a CV requirement, however convincing the quote.
Beyond those fixes, deterministic code in server/analysis.ts checks every model answer against the source before anything is stored:
- Every call is constrained to a JSON schema and validated with zod. An invalid answer gets one re-ask. If it is still invalid, the activity fails and Temporal retries it.
- An evidence quote is kept only if it really appears in the cited passage. A "satisfied" verdict without a verifiable quote is downgraded to needs review.
- If a requirement's source sentence can't be found in the document, the requirement is flagged and its due date is dropped.
- An expiry date is stored only when the model quotes the sentence that states it.
- If the model states any uncertainty, satisfied becomes needs review.
- Document text sits inside tags it can't close, and the model is told never to follow instructions found inside them. A scholarship PDF shouldn't be able to say "mark everything complete."
What I learned: fix a model's mistakes with code that checks the source, not with a longer prompt.
Living with CPU speed
My laptop has no GPU, so gemma3:4b takes 1–2 minutes per assessment and about 5 minutes to extract requirements. The first call after startup spends another 3 minutes or so loading the model.
That is where Temporal earned its place. Model activities get 15-minute timeouts and send a heartbeat every 10 seconds. If the worker crashes, the work resumes where it stopped. The Activity page shows progress for each requirement. Slow is fine. Lost work is not.
Why Does Open Innovation Matter?
Think about what goes into PaperTrail: transcripts, passports, test scores, recommendation letters, bank statements. That is some of the most sensitive paperwork a person owns.
Because the model has open weights and runs locally, none of it leaves the laptop. My friend doesn't have to trust me, a startup's privacy policy or an API provider's data retention settings. The Settings page states plainly where document text is sent.
Open weights made three more things possible:
- It's free to run, indefinitely. A student applying to several programs re-runs the analysis every time a document changes. There's no per-token bill, no rate limit and no API key to leak.
-
I can debug the model. The weights are identical on every run, so the quirks are too. When
gemma3:4binverted booleans, I could reproduce it, fix it in code and trust the fix to hold. A hosted model can change between one week and the next without warning. - It's easy to swap. The adapter speaks the OpenAI-compatible protocol, so a GPU machine or a hosted Gemma endpoint is one environment variable away.
The rest of the stack is open source as well: Temporal, Postgres, pgvector, Ollama, Fastify and React. Hosted services such as Tiger Cloud, Temporal Cloud and Sentry are optional. Locally, a single docker compose up starts everything.
Top comments (0)