DEV Community

Rahul Kumar
Rahul Kumar

Posted on AI-assisted

Padh Ke Batao: an AI that reads my grandfather's documents to him, on a laptop with no GPU

Hacktoberfest Weekend Challenge: Build for a Friend Submission ЁЯдЭ

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My Nana (my mother's father) gets a steady stream of paper he can't fully read: bank notices, hospital reports, pension circulars, government letters. They're dense, formal, often in English, and the one thing that matters (do I need to do something, and by when?) is usually buried in paragraph three.

Recently it was a notice about SIR, the Election Commission's Special Intensive Revision of the voter rolls: the kind of notice where missing a step can keep your name off the voter list. He couldn't read it. So he asked the way he always does: "padh ke batao", read it and tell me.

That's the name of what I built. Padh Ke Batao (рдкрдврд╝ рдХреЗ рдмрддрд╛рдУ) does exactly what he asks. He chooses a photo of the document and, about a minute later, gets:

  • What it is and who sent it
  • What it says, including what happens if he ignores it
  • What to do, step by step, with every document he needs to bring written out
  • The last date, with a countdown ("15-10-2026 ┬╖ 10 рджрд┐рди рдмрд╛рдХреА", 10 days left)
  • How urgent it is: рдЬрд▓реНрджреА рдХрд░реЗрдВ (act soon), рдзреНрдпрд╛рди рджреЗрдВ (needs attention) or рд╕рд┐рд░реНрдл рдЬрд╛рдирдХрд╛рд░реА (information only)
  • A big ЁЯФК рд╕реБрдиреЗрдВ (Listen) button that reads it all out in a natural Hindi voice

Everything that reads the document runs on an ordinary laptop: an open-weight vision model in Ollama, no GPU, no cloud AI. The photo of his bank letter never leaves the machine.

Nana is far from the only one. In millions of Indian homes, an official letter arrives and sits on a shelf until someone has time to read it to the person it was sent to. I built this for him, but it works for anyone in that spot.

Padh Ke Batao explaining a bank KYC notice in Hindi: urgency badge, deadline with days left, what it says, what to do, key details

The letter in these screenshots is made up, so no real account appears anywhere. The explanation is the model's real output.

Demo

Start Reading Result (English, dark mode)
Start screen in Hindi Progress while the model reads the document English result in dark mode

Code

GitHub logo Rahullkumr / padh-ke-batao

Hacktoberfest 2026 week 1 challenge - Build for a Friend

Hacktoberfest 2026 Weekend Challenge: Open-Source AI Challenge

рдкрдврд╝ рдХреЗ рдмрддрд╛рдУ ┬╖ Padh Ke Batao

Padh ke batao (рдкрдврд╝ рдХреЗ рдмрддрд╛рдУ) is Hindi for "read it and tell me". It's what my Nana says whenever he hands someone a letter he can't read.

Take a photo of a document (a bank notice, hospital report, pension circular or government letter). Get it explained in simple Hindi or English: what it says, what to do, and by when.

The reading happens on your own computer, with an open-weight vision model running in Ollama. The photo never leaves the machine.

Built for my grandfather for DEV's Hacktoberfest 2026 Weekend Challenge: Build for a Friend.

Hindi explanation of a bank KYC notice: urgency, deadline with days left, what it says, what to do, key details

Demo

Watch the Padh Ke Batao demo on YouTube

тЦ╢ Watch the demo on YouTube (sound on for the Hindi voice)

Why

My grandfather gets a steady stream of letters he can't fully read: bank notices, hospital reports, pension circulars, government letters. They're dense, formal, often in English, and the part thatтАж

How I Built It

Architecture: photo to browser page, POST /explain to a Fastify server on 127.0.0.1, image and prompt to Ollama running qwen3.5:4b, all inside

The open-source AI at the core is qwen3.5:4b, an open-weight vision model (3.4 GB), running locally in Ollama. Around it:

  • A small Fastify server, listening only on 127.0.0.1. It sends the photo to Ollama with a prompt and a JSON schema passed to Ollama's format field, so every answer has the same six fields: document_type, summary, key_details, action_needed, deadline, urgency.
  • Streaming. The answer streams back as newline-delimited JSON, so the page shows "readingтАж" then "writingтАж" with a progress bar instead of a frozen screen for a minute.
  • A plain HTML/CSS/JS page. No framework, no build step, big text, high contrast, Hindi/English and light/dark toggles.
  • Built for a CPU-only laptop: an i7-13700H with 16 GB RAM and integrated graphics, the kind of machine a family already owns.

Making a 4B model trustworthy on a CPU

Getting an answer was easy. Getting an answer I'd trust in front of Nana took four rounds of testing on six real documents: a bank KYC reminder, the Election Commission's press note about the same SIR drive, a Hindi government message, a medical check-up report, a pension notice and a YouTube copyright email.

1. Speed: the image was the bottleneck, not the model. My first run took 147 seconds. Ollama's timing stats showed why: the photo became 2,391 image tokens that took 95 s just to read, while writing the answer took only ~35 s. Shrinking the photo to 768 px brought a full explanation down to 50 s. Then I found out what that cost.

2. The scariest bug wasn't a crash. At 768 px, a harmless Hindi announcement from a state government department came back as a warning that "you may be punished if you don't submit your details", complete with an invented city and a phone number, 1234567890. Nothing in the letter said any of that. The model couldn't read dense Devanagari at that size, so it confidently made something up. At full resolution it read the same letter perfectly: "only information, nothing to do."

A fixed width was the wrong knob. The page now shrinks photos to a budget of ~900,000 pixels: big enough for dense Hindi text, small enough to read in about 30 s. Final runs on the six documents took 56тАУ91 seconds. I'll take a minute of waiting over a fast answer that frightens him.

3. Language drift. In Hindi mode, an English press note got an English answer, and others came back in Roman-letter Hinglish ("Aapko branch jaakarтАж"). The fix was a rule that every field must be in the chosen language and Hindi must be Devanagari, repeated in Hindi at the end of the prompt.

4. Temperature, in both directions. At 0.2 the model got stuck repeating itself until the JSON broke. At 0.7 it translated YouTube's "claimed" as рдЕрдкрд╣рд░рдг (kidnapping). 0.4, with top_p 0.8, top_k 20 and a 900-token cap, was the balance.

5. Small rules that matter. I'd put today's date in the prompt so it could flag past deadlines; it started reporting today's date as a date in the letter. I removed it, and the page now counts the days itself. Other rules came straight from failures:

  • A printed or test date is not a deadline.
  • Write out every step and document; never say "the items below".
  • At most six key details, preferring a phone number over an IFSC code.
  • For medical reports, explain the values but never diagnose or suggest changing medicines.

After the final round, the urgency was right on all six documents, with no invented threats, names or numbers. It isn't perfect: it still picks an odd word sometimes, and it once translated the bank name "Sunrise" as рд╕реВрд░рдЬ (sun). So the page always ends with "if in doubt, check with family or the office that sent it." The final prompt is short enough to read in a minute.

The voice: ElevenLabs

A lot of the reason Nana hands letters over is that reading is the hard part, so the explanation needs to be heard, not just shown. The ЁЯФК рд╕реБрдиреЗрдВ button uses ElevenLabs (eleven_multilingual_v2, slowed to 0.9├Ч speed) for a natural Hindi or English voice. I was careful about what crosses the internet:

  • Sent: only the urgency, deadline, summary and steps. Never the photo, and never the key-details list, which is where names and account numbers live.
  • Server-side key: the API key stays on the local server; the browser never sees it.
  • Replay is free: pressing Listen again reuses the audio instead of using more credits.
  • Offline fallback: with no internet or no key, it falls back to a voice installed on the computer.
  • Honest UI: a line under the button says exactly what is sent to ElevenLabs.

Why Does Open Innovation Matter?

Because these are Nana's bank letters, pension papers and medical reports.

  • His documents stay home. The model runs on the laptop, and the server only listens on 127.0.0.1. No account, no upload, no third party keeping a copy of his bank details under a retention policy he never agreed to.
  • It runs on what we already own. No GPU, no cloud bill, no rented server. An open 4B model on an ordinary laptop CPU is enough, once you tune it.
  • Nothing costs anything per document. He can check every envelope, not just the scary-looking ones, and nobody has to think about a bill.
  • I could look inside and fix it. Ollama told me exactly where the time went (reading the image vs writing the answer), which is how I found the resolution trade-off. I could tune temperature, sampling, output length and the schema until the answers were trustworthy, then pin that exact model so it doesn't change under him.
  • It works with the internet unplugged. Reading and explaining need no connection; only the ElevenLabs voice does, and it falls back to a local one.

A closed vision API would have meant uploading photos of a family's financial and medical paperwork to someone else's servers, and paying for every one.

Handing It Over

I showed it to Nana, and he loved it. The part he liked most was the Hindi voice. He speaks Hindi most of the time, so hearing the explanation out loud, in his own language, is what made it click for him. He didn't have to read anything, not even the screen.

That changed how I see the app. I'd treated the Listen button as an add-on; for Nana it's the main way in. It's also why I'm comfortable with the one thing that leaves the laptop: a few sentences of summary sent to ElevenLabs, in exchange for a voice he actually enjoys listening to.

My Agent Session

I built this with Claude Code as a pair programmer for planning, debugging and the test rounds. That was during development only: at runtime the app uses no closed-model AI. All the reading and explaining is done by the open Qwen model in Ollama.

Building Padh Ke Batao: a local Hindi document explainer for my Nana claude-opus-5-5
You

Continue my Hacktoberfest "Build for a Friend" entry (letter explainer for my grandfather). Read my project memory first; everything is decided, so don't redo it. Deadline is Oct 5, 12:29 IST.

Start with the pending repo housekeeping, then build step 1: install Ollama, pull gemma4:e4b, and time it on one sample letter. Ask before installing anything.

When we finish, save this session to DEV via DevRelay so I can embed it in my post.

Agent

Housekeeping done. Package files stay uncommitted until real app code lands (per plan), now clean of Backboard.

Step 1 check before install:

Fine. While you install Ollama (ollama.com/download, Windows installer), I'll make a synthetic letter and the timing script.

Ready on my side. Waiting on your Ollama install.

Done:
- Housekeeping: most of it was already finished last session (5b3d3d0 has .gitignore with .agents/ and .env.example; backboard-test.js was already deleted). Today I removed the leftover backboard-sdk dependency and the test:backboard script from package.json, deleted the empty scripts/ folder, and added samples/ to .gitignore. Committed as 01fe098. Nothing pushed. The package files stay uncommitted until real app code lands.
- Sample letter: made a fake "final reminder" bank KYC notice at samples/synthetic-kyc-notice.jpg. The bank and person are invented. It has a deadline, a list of documents to bring, a warning about debit restrictions, and an offer of a doorstep visit for seniors, so it tests the action, deadline and urgency fields.
- Timing script: scripts/time-letter.js sends the photo to Ollama with the agreed JSON schema and prints the explanation. It reports model load time, image-reading time and generation speed separately. It also doubles as a prompt playground for step 4. Syntax checks out; I haven't run it yet because Ollama isn't installed.

Your turn:
1. Install Ollama from ollama.com/download (Windows installer, native, not Docker).
2. Tell me when it's done. I'll then run ollama pull gemma4:e4b (about 10 GB) and the timing script twice. The first run includes loading the model, so the second run is the real per-letter speed:

node scripts/time-letter.js samples/synthetic-kyc-notice.jpg hi

If the second run takes more than about 1 minute, we try qwen3.5:4b or shrink the image before rethinking anything.

I'll upload the session to DevRelay when we finish today's work, not now.

You

we will use qwen3.5:4b not the gemma4:e4b. tell me what is it's size. ollama is being downloaded

Agent

Recording the model switch and the finished housekeeping in project memory so the next session doesn't redo them.

The script now uses qwen3.5:4b by default and is ready to time.

  • Size: qwen3.5:4b is 3.4 GB (text + image, 256K context), from the ollama.com tags page. Gemma4:e4b was about 10 GB, so the download is faster and more of your 16 GB RAM stays free.
  • Thinking turned off: Qwen3.5 thinks before answering by default in Ollama, which could add minutes per letter on CPU. I added think: false to the request in scripts/time-letter.js and made qwen3.5:4b the default model. Syntax checks out; it hasn't been run yet.
  • Memory updated: project memory now records the model switch and that the repo cleanup is finished.

When Ollama finishes installing, tell me and I'll run:

ollama pull qwen3.5:4b
node scripts/time-letter.js samples/synthetic-kyc-notice.jpg hi # twice: first includes model load, second is real speed

You

[screenshot attached]
ollama download and install finished, it is showing this

Agent

Installed. Locating the binary and checking the server is up before pulling.

Ollama 0.35.1 running on localhost:11434. Pulling qwen3.5:4b (3.4 GB) in background.

The install worked: Ollama 0.35.1 is running on localhost:11434. You can ignore the "Run Ollama" window; I'm calling the binary directly.

qwen3.5:4b (3.4 GB) is now downloading in the background. When it finishes I'll run the timing script on the sample letter twice. The first run includes loading the model; the second shows the real per-letter time. You'll get both results plus the Hindi explanation.

Model pulled (3.4 GB, digest verified). Run 1 тАФ cold, includes model load:

Node 22 resolves localhost to IPv6 ::1, but Ollama listens only on IPv4 [REDACTED]. Switching the default URL to [REDACTED] (the backend will need the same).

Works. Cold run: 147.5s. The bottleneck is reading the image: 2391 image+prompt tokens take 95s. Generation speed is fine at 8.5 tok/s. Warm run next, to remove the 13s model load:

Step 1 numbers are in. Recording them in memory before the session upload:

Step 1 is done, so I'm saving the session to DEV as you asked. First I'll build a cleaned transcript from the local session log, scrubbing your email, machine paths and any secrets.

Uploading only the cleaned transcript, not the raw S3 upload, because the raw log would include your unscrubbed paths and email.

Scrubbed transcript is clean: 6 messages, no leaks found. It covers the session from your first prompt through the step 1 timing results. Submitting it to DEV as claude_code:

34 KB is too big to paste reliably into the submit call. DevRelay has an upload slot for that, so I'll upload the scrubbed file there instead of the raw log.

Two gaps in the cleaned transcript: some of my in-between explanations are missing (e.g. the "Misleading: 0.3s is the cache" note), and the 768px result got cut off. Checking why:

All four timings are now in, with no leaks. Pulling the two updated outputs so the submission matches the file exactly:

Saved. Recording the session ID in memory:

Step 1 is done and the session is saved to DEV. qwen3.5:4b is fast enough if we shrink the photo before sending it: about 50s per letter at 768px.

Timing on the synthetic bank letter (Hindi output):

Photo size Time per letter Image tokens
Full, 1240├Ч1754 ~130s (147s on the first run, which includes 13s model load) 2391
1024px long side 69s 982
768px long side 50s 654

Reading the image is the slow part. Writing the answer runs at about 9 tokens/s and adds 33тАУ40s whatever the photo size. A repeat run on the identical image took only 34s, but that's because Ollama cached the image; a new letter won't get that.

Plan for step 2 (recorded in memory):
- Shrink in the browser: resize photos to 768px on the long side before upload. Go to 1024px if real phone photos get misread.
- Keep the model loaded: set Ollama's keep_alive long so Grandpa doesn't wait for the 13s load.
- Stream the answer: show text as it arrives so the wait feels shorter.

Prompt problems for step 4:
- Missed: the warning that ATM and cheque payments will be blocked, and the doorstep-visit offer for seniors.
- Urgency: it said "medium" when it should probably be "high".
- Invented: at 1024px it made up an "online/mobile app" option.
- Word choice: at 768px it called the letter "рдкрддреНрд░рд┐рдХрд╛" (magazine).

Bug fixed: Node 22 sends localhost to the IPv6 address, but Ollama only listens on IPv4. The script now uses [REDACTED], and the backend will need the same.

DEV session: saved as an unpublished draft (ID 453): https://dev.to/agent_sessions/build-for-a-friend-step-1-a-local-vision-model-reads-grandpas-bank-letter-ax9nge. To embed it in your post: {% agent_session 453 %}.
- Scrubbed: your email, home folder paths and session IDs. No API keys appear in it.
- Trimmed: long outputs are shortened; the timing results are kept in full. I also removed terminal colour codes and shortened the error trace.
- Left out: the upload steps themselves, because they contained the scrub rules, including your email. A few of my short in-between notes weren't in the session log, so they're missing too.

Review it on DEV before publishing.

Repo:
- Committed: 01fe098, which ignores samples/. Nothing pushed.
- Not committed: scripts/time-letter.js and the cleaned-up package.json. I'll commit them with the backend code unless you want the script committed now.

Next is step 2, the Fastify /explain backend.

You

revert the last commit, and you do not do any git code movement or commits, I will do it myself

Agent

Recording this preference in memory and fixing the step 1 note that mentions the commit:

I've undone the commit with git reset HEAD~1, so it no longer exists. Your .gitignore change (ignoring samples/) is still there, just uncommitted, and main is back at 5b3d3d0. I'd committed locally only, so there's nothing to undo on GitHub.

From now on I won't commit, stage, reset, switch branches or push. I'll only edit files and use read-only git commands like status, log and diff. That preference is now saved in memory for future sessions.

Ready for you to commit:
- .gitignore: adds samples/
- scripts/time-letter.js: the timing script
- package.json and package-lock.json: Backboard removed

The saved DEV session (453) still shows the 01fe098 commit being made. Mention that it was reverted if you embed that part, or tell me and I'll trim it from the session.

You

lets go to next step

Agent

Now switching the timing script to the shared module, so step 4 prompt tuning changes the real app's prompt too:

package.json already lists fastify ^5.12.5, so it looks like you installed it. Checking, and making sure Ollama is running:

Fastify 5.12.5 is installed. Ollama is down, so I'm starting it in the background, then the server:

Model loads at startup, and bad requests are rejected. GET / returns 500 for now because the page comes in step 3. Now a real streamed request with the 768px letter:

Restarting the server, then testing that closing the page mid-request really stops Ollama. If it doesn't, the next letter would wait behind the abandoned one.

Rerunning the cancel test on the new server, using two image sizes that haven't been sent before so Ollama's cache can't skew the timing:

"Reading" now shows instantly (0.1s). Cancelling half works: request B still waited about 25s extra. Checking Ollama's log to see when the cancelled request actually stopped:

Step 2 is done: the backend explains a 768px letter in about 54s, streaming progress while it works. All tests passed except one known delay after cancelling, explained below.

What's built:
- src/explain.js holds the prompt, the agreed JSON schema and the Ollama call. The timing script now uses it too, so prompt tuning in step 4 changes the real app.
- src/server.js is a Fastify server that listens only on [REDACTED]:3000, so the letter never leaves the laptop. It loads the model at startup and keeps it in memory for 30 minutes.
- POST /explain takes { image, lang } and streams progress lines ("reading", then "writing"), then the result.
- GET / will serve public/index.html, which comes in step 3. Until then it returns a 500.
- package.json now has npm start and npm run time (the timing script).

Tested:
- Bad input: a wrong language or a non-image file gets a clear error back.
- Real letter: the 768px letter streamed correctly, with "reading" shown instantly and the full result in about 54s.
- Better answers: the stricter prompt now gives the consequence of ignoring the letter and the bank's phone number.
- Still weak: it still misses the doorstep-visit offer and rates urgency "medium". That's step 4.
- Cancelling: if the page closes mid-request, the Ollama request stops at once. But Ollama finishes reading the current image (about 20s) before it can start the next letter. That's inside Ollama; it only matters on an immediate retry.

Bug fixed: the "reading" message used to arrive 16s late because Ollama sends nothing until its first word. The server now sends it immediately.

On Windows, stopping npm start left the old server running on port 3000. Run node src/server.js directly when restarting.

Ollama and the new server are still running in the background for step 3.

Ready for you to commit: src/, scripts/time-letter.js, package.json, package-lock.json and .gitignore.

Next is step 3, the page: photo upload, resizing to 768px in the browser, big text, a Hindi/English toggle and live progress. Shall I go ahead?

You

yes

Agent

Building the page now. Design choices for your grandfather:
- One big button: choosing a photo starts the explanation straight away, no second click.
- Hindi by default: the toggle switches both the page labels and the explanation language.
- Big text, high contrast, Hindi fonts built into Windows: no web fonts, so the page makes zero internet requests, which supports the privacy story.
- Urgency and deadline first: urgency is a colored badge with words, and the deadline shows "N days left".
- Progress: the progress bar and seconds counter cover the ~50s wait.

One gap: the server's error messages are English-only, so in Hindi mode they'd show up in English. Making the page show its own translated error text instead:

It was my test's fault: sed turned \n into real line breaks inside a JS string. The app code is fine. Rebuilding the mock-up with Node, plus a phone-width (400px) shot:

Step 3 is written: the page is public/index.html. I checked it with screenshots in a headless browser, but nobody has clicked through it yet. That's the next thing to do.

What it does:
- One button: "рдЪрд┐рдЯреНрдареА рдХреА рдлреЛрдЯреЛ рдЪреБрдиреЗрдВ" opens the photo picker, and choosing a photo starts the explanation right away.
- Shrinks the photo to 768px in the browser before sending it.
- Language toggle (рд╣рд┐рдВрджреА / English): switches the page text, and if a letter is showing, explains it again in the new language (about 50s).
- Progress: a progress bar for "reading" then "writing", plus a seconds counter with "usually takes about 1 minute".
- Result: a colored urgency badge, the deadline box with "N days left", what the letter says, step-by-step actions, key details, and a "check with family if in doubt" note.
- Errors in the chosen language: for a bad photo, a failed reading, or the app not running.
- No internet requests at all. It uses Hindi fonts already on Windows.

Checked:
- The page loads from the server.
- The page's code has no syntax errors.
- Screenshots of the start screen and a mock-up result screen look right.

Your turn: both servers are still running, so open http://[REDACTED]:3000 and pick samples/synthetic-kyc-notice.jpg, or a real letter. Tell me what looks off.

For step 4 (prompt tuning):
- Language: the document type came back in English in Hindi mode.
- Doorstep visit: still missing.
- Urgency: still rated too low.
- Real letters: testing 3тАУ4 real ones will show whether 768px is sharp enough for phone photos.

New file ready for you to commit: public/index.html.

You

I added few more documents for testing

Agent

Five new documents: copyright notice, Election Commission letter, gau seva portal inauguration, general checkup report, pension notice. Looking at each first, so I know what the model should say:

All five are now at the size the browser would send. Running the Hindi baseline. This takes about 5 minutes, so it runs in the background:

While the baseline runs I'll leave the prompt alone, because each run loads it fresh. Once I see the results I'll make all the prompt fixes in one pass.

Code changes are done. Rerunning all six documents with the new prompt and photo size, in the background, about 7 minutes:

Round 2 is running: 6 documents at about 70s each. The running web server still has the old prompt loaded, so I'll restart it once the test finishes.

Saving the findings so far to memory:

Findings are in memory. The test is still running; I'll report as soon as all six are done (about 7 minutes).

My replacement turned \n into a real line break inside a quoted string. Fixing it:

All edits are in and the syntax is valid. Running round 3 on all six documents in the background (about 7 minutes):

Round 3 is running. I'll report when all six finish.

Round 3 is 2 of 6 done; the remaining 4 need about 5 more minutes. Both finished faster than round 2: copyright 53s, Election Commission 56s (was 84s). Here's what they say:

Round 3, 2 of 6 done (02:28 IST). The remaining four (gau seva, checkup, pension, KYC) need about 5 more minutes.

So far, partly better and partly worse:

  • Fixed: both answers are in Devanagari, both are faster, and the Election Commission note is correctly flagged as general information with low urgency.
  • New problem 1, today's date leaks into answers: the model wrote "5 рдЕрдХреНрдЯреВрдмрд░ 2026" (today) as a date in the documents. That's wrong: the Election Commission work started on 4.11.2025, not today. Adding today's date to the prompt caused this.
  • New problem 2, scary wording: for the copyright email, the model translated "claimed" as "рдЕрдкрд╣рд░рдг" (kidnapping). It also added an invented "your video may be removed". The higher temperature makes word choice looser.

My plan once the last four finish:
- Keep Devanagari and the 6-detail limit.
- Today's date: stop putting it in the prompt. Instead the page compares the deadline with today itself, which it already does ("N days left" / "date has passed"). That's more reliable than asking the model.
- Temperature: settle on 0.4 with the 900-token cap as the safety net against loops.

Then one final run, and step 4 is done.

Five of six are done; the KYC letter is still running. The three newest:

Round 3 finished. Starting the final round in the background, then looking at the KYC result from round 3:

The round 3 KYC result is the best so far:
- Urgency: high, correctly.
- Consequence: the ATM and cheque block, with pension still coming in.
- Doorstep visit: mentioned for the first time ("call and they'll come home to update KYC").
- Documents: all four listed.

It also listed "рдЖрдЬ рдХреА рддрд╛рд░реАрдЦ: 5-10-2026" as a fact from the letter, which confirms removing today's date was right.

The final round is running (about 7 minutes). After that I'll restart the web server so the page uses the final prompt.

The final round is 2 of 6 done (02:37 IST). The other four need about 5 more minutes. Both answers are the cleanest so far:

  • YouTube copyright: accurate and calm. "Just information, your account is fine, nothing to do", plus the dispute option. No scary words like "рдЕрдкрд╣рд░рдг" this time.
  • Election Commission: most numbers right now.
    • Right: 51 рдХрд░реЛрдбрд╝ voters (was "51 рд▓рд╛рдЦ"), the real dates (published 04.11.2025, work runs until 4 Dec 2025), Booth Level Officer (was "Blue Level"), helpline 1950, and it's flagged as a general notice.
    • Still off: "12 states and 3 UTs" (it's 9 + 3 = 12). The deadline went into the details rather than the deadline field, so the page won't show the "date has passed" warning for this one.

I'll restart the web server with the final prompt once all six finish.

Step 4 is done. The final prompt gave the right urgency on all six test documents, with no made-up threats or numbers, at 56тАУ91s per letter. The web server is restarted with it, so http://[REDACTED]:3000 is ready to try.

What changed in step 4:
- Photo size: capped at ~900k pixels instead of 768px on the long side. That fixed the gau seva letter, where the model had invented "Jamshedpur" and a punishment because the Hindi text was too small to read.
- Language: every answer in Devanagari Hindi. Earlier rounds drifted into English or Roman-letter Hindi.
- Prompt rules:
- Copy numbers and units exactly.
- Never invent anything.
- A printed or report date isn't a deadline.
- Say who a notice applies to.
- Flag general announcements.
- No medical advice.
- At most 6 details, preferring ones he needs to act on.
- Write out every step and every document to bring.
- Settings: temperature 0.4 with a 900-token cap. 0.2 made the model loop and break its output; 0.7 produced odd words like "рдЕрдкрд╣рд░рдг" (kidnapping) for "claimed".
- Today's date is no longer in the prompt. The model was copying it into answers as a date from the document. The page already works out "N days left" itself.

Still imperfect:
- An occasional odd word.
- The Election Commission note still says "12 states" instead of 9 states + 3 UTs.
- The deadline sometimes lands only in the details list, so the "date has passed" box doesn't show for that letter.

Worth saying honestly in the post. The four rounds, the invented "punishment" letter and the speed-versus-accuracy trade-off are all strong material.

Next is step 5, the optional read-aloud. I suggest swapping ElevenLabs for the browser's built-in speech. It works offline and keeps the "nothing leaves this computer" promise, which ElevenLabs would break. It needs a Hindi voice installed on Windows; I can check whether yours has one. Or we skip step 5 and go straight to the README and demo (step 6). It's 02:50, and the deadline is 12:29. Which do you prefer?

You

go to step 5 no elevenlabs for this project

Agent

No ElevenLabs. I'll use the browser's built-in speech, restricted to voices that run on the computer. Edge also offers "Online (Natural)" voices that send the text to Microsoft's servers, and the page will skip those. First, checking which Hindi voices Windows has installed:

The page is longer than when I wrote it (618 lines, was about 340), so you or something else has edited it. Reading the current version before changing anything:

You've redesigned the page (dark mode, steps, drag-and-drop, phase pills), so I'll build on your version and keep its style.

Now the text strings and the speech logic:

Checking the page's syntax, then listing which voices Edge reports and which of them run on this computer:

The layout works: with no voices, the "add a Hindi voice" note appears where the button would be. The card looks faded only because the screenshot caught its fade-in animation.

Step 5 is written: a "рд╕реБрдиреЗрдВ / Listen" button on the result, using only voices that run on the computer. It hasn't been heard yet. Headless Edge has no voices, and your Windows has no Hindi voice, so in Hindi mode you'll see the "add a voice" note instead of the button for now.

How it works:
- Button on the result: "рд╕реБрдиреЗрдВ / Listen" sits just under the urgency badge. Pressing it again ("рд░реЛрдХреЗрдВ / Stop") stops it.
- What it reads: urgency, document type, deadline with days left, what the letter says, what to do, and the key details. It speaks at 0.9├Ч speed, one sentence at a time.
- Only voices that run on this computer. Edge's "Online (Natural)" voices are skipped because they send the text to Microsoft.
- Stops when a new letter starts or the language changes.
- No voice installed: the button hides and a note explains how to add one.

To hear it in Hindi, add the free Hindi voice (your call):
Windows Settings тА║ Time & language тА║ Speech тА║ Add voices тА║ рд╣рд┐рдВрджреА (рднрд╛рд░рдд). Then reload http://[REDACTED]:3000.

English mode should work right away, since Windows already has the Indian-English voices Heera and Ravi.

Please try:
1. Explain a letter and press "рд╕реБрдиреЗрдВ / Listen".
2. Check it reads clearly.
3. Check "рд░реЛрдХреЗрдВ" stops it.

.env.example probably still has an ElevenLabs key line. Remove it if so, since ElevenLabs is out.

Ready for you to commit: public/index.html, src/explain.js.

Next is step 6: the README, demo video and screenshots. Shall I start on the README?

You

let us use the elevenlabs

Agent

Switching back to ElevenLabs is your call, and it's quick to wire up. Two things to settle first, because one of them could disqualify the entry:

  1. Rules: your memory notes say the app's AI must be open-weight, and ElevenLabs is a closed AI voice service. Checking the official challenge rules now.
  2. Privacy: the page promises "never sent over the internet". With ElevenLabs, the explanation text goes to their servers when he presses Listen, so that line has to change.

Installed into .agents/skills/, which is gitignored, so nothing extra gets committed. The skill loads properly next session, so I'm reading its text-to-speech guide directly now:

ELEVENLABS_API_KEY is set. Checking which voices on your ElevenLabs account suit Hindi. This reads the key from .env and never prints it:

The key is empty: ELEVENLABS_API_KEY exists in .env but has no value. Checking the line's length without showing its contents:

Server side is done. Now the page: ElevenLabs first, local voice as fallback, and the honest privacy wording. Checking the current speech code first, since you've been editing this file:

Page code is valid. Restarting the server with .env loading, then checking the new routes:

ElevenLabs is wired in as the main "Listen" voice, but it's untested. Your .env has an empty ELEVENLABS_API_KEY, so only the "no key" path has been checked. With no key, the page falls back to a voice installed on the computer.

How it works:
- Server: src/speak.js calls ElevenLabs directly (model eleven_multilingual_v2, which supports Hindi, at 0.9├Ч speed). The key stays on the server.
- Routes: POST /speak returns the MP3, and GET /config tells the page whether ElevenLabs is set up.
- Sent to ElevenLabs: only urgency, deadline, summary and steps, as you chose. Never the photo or the details list.
- Replay is free: pressing Listen again on the same letter reuses the audio instead of using more credits.
- Fallback: if there's no key, no internet or an error, it uses a local Windows voice. Those voices stay on the computer, so they can also read the details list.
- Privacy text: the page now says "the photo never goes online". A note under the Listen button says exactly what goes to ElevenLabs. You'll want the same honesty in the post.
- Prize: the rules only require open-source AI at the core, which Qwen via Ollama is. "Best Use of ElevenLabs" is a $100 prize category, so the entry now qualifies for it.
- Start command: npm start now loads .env automatically.
- Official skill: ElevenLabs' skill is installed in .agents/skills/, which is gitignored. Full loading needs a new session.
- .env.example: rewritten to list only the ElevenLabs and Ollama settings. I removed the unused Backboard, Render and Tinker lines.

Your turn:
1. Paste your key into .env: ELEVENLABS_API_KEY=sk_.... ElevenLabs dashboard тА║ Developers тА║ API Keys.
2. Tell me when it's done. I'll restart the server and test a real Hindi clip.
3. Optional: the default voice, George, is British. For Hindi, an Indian voice from the ElevenLabs Voice Library will sound more natural. Add one to your account and put its ID in .env as ELEVENLABS_VOICE_ID=....

Ready for you to commit: src/speak.js, src/server.js, public/index.html, package.json, .env.example.

Next is step 6: README, screenshots and demo video.

You

api key added in .env

Agent

The key looks well formed (51 characters, sk_ prefix, no stray quotes or spaces). Restarting the server so it picks up the key:

ElevenLabs works with your key: the KYC explanation in Hindi came back as a 553 KB MP3 in 7.5s. I can't listen to audio, so you need to judge how it sounds.

  • Test clip: samples/elevenlabs-test-kyc-hi.mp3 is the KYC letter read in Hindi (gitignored). Check the Hindi pronunciation and the pace.
  • Wait label: during the ~7s it takes to make the voice, the button now shows "рдЖрд╡рд╛рдЬрд╝ рдмрди рд░рд╣реА рд╣реИтАж (рд░реЛрдХреЗрдВ)". It switches to "рд░реЛрдХреЗрдВ" once playback starts. The page is read fresh on every load, so a browser refresh picks this up.
  • Full flow: the server is running with the key. On http://[REDACTED]:3000: choose a letter, wait for the explanation, press "рд╕реБрдиреЗрдВ", replay it, then switch to English.

If George's British accent sounds off in Hindi, pick an Indian voice in the ElevenLabs Voice Library and put its ID in .env as ELEVENLABS_VOICE_ID. I'll restart the server to apply it.

Ready for you to commit: src/speak.js, src/server.js, public/index.html, package.json, .env.example.

Next is step 6: README, screenshots and demo video. It's 03:15, about 9 hours before the deadline.

You

voice is fine, and it is working perfectly fine. before going to next step. let's make the folder structure professional

Agent

I'll check the current layout and what git sees, since the ElevenLabs skill install may have added folders:

The split boundaries are clear. Splitting the page with a script that cuts at exact markers, so nothing gets retyped by hand:

The page is split into five files, all with valid syntax. Checking that no code still refers to something that moved to another file:

Frontend done: the theme script stays inline in <head> on purpose, so dark mode doesn't flash white on load. Now the server side. Reading the current prompt file so it moves over unchanged:

Writing the new server layout. The prompt and schema move over by cutting the exact lines out of the old file, so the tested prompt stays byte-for-byte identical:

The prompt moved unchanged; the only differences are the two export keywords. Now config, services, routes and the server entry:

Now the timing script imports, removing the old files, the banner move, LICENSE, package.json and .gitignore:

npm pkg set saved private as the string "true" instead of a boolean. Fixing that and adding the LICENSE:

Everything parses. Restarting the server on the new layout and testing every route, including that crafted paths can't escape public/:

All routes work, ElevenLabs is enabled, and all four path-escape attempts get a 404. Last check: the page's JavaScript modules actually run in a browser. If they do, the text labels get filled in:

Edge's DOM dump returned nothing; it's been unreliable all night. Using a screenshot instead, since those have worked:

Probably leftover headless Edge processes from earlier runs blocking new ones. Closing only the ones I started (they have my scratch profile folder in their command line), not your normal browser:

The page renders from the new files. The styles load, the scripts run (all labels filled in), and the updated privacy line shows. Final layout and memory update:

The project now uses the layered structure with an MIT license. The server is running on it and every route passed. The page renders and its scripts run, but I could only check the start screen in a headless browser; a real letter hasn't been explained through the new files yet.

build-for-a-friend/
тФЬтФАтФА public/
тФВ   тФЬтФАтФА index.html
тФВ   тФЬтФАтФА css/styles.css
тФВ   тФФтФАтФА js/
тФВ       тФЬтФАтФА app.js           upload, resize, progress, result
тФВ       тФЬтФАтФА i18n.js          Hindi/English text
тФВ       тФФтФАтФА speech.js        Listen button (ElevenLabs, then local voice)
тФЬтФАтФА src/
тФВ   тФЬтФАтФА server.js            sets up Fastify and registers the routes
тФВ   тФЬтФАтФА config.js            every env setting in one place
тФВ   тФЬтФАтФА prompts/explain-document.js   prompt + schema (moved over unchanged)
тФВ   тФЬтФАтФА services/ollama.js, elevenlabs.js
тФВ   тФФтФАтФА routes/explain.js, speak.js, pages.js
тФЬтФАтФА scripts/time-letter.js
тФЬтФАтФА docs/images/banner.webp  (README link updated)
тФЬтФАтФА .env.example ┬╖ LICENSE (MIT) ┬╖ package.json ┬╖ README.md
тФФтФАтФА samples/                 (gitignored)

Worth knowing:
- Prompt file: src/prompts/explain-document.js holds the prompt exactly as tested, so the post can link to it.
- Serving the page: pages.js only serves files that exist in public/, with no extra dependency. I tried four ways of reaching files outside it (like ../.env) and all were blocked.
- package.json: MIT license, author, private: true, Node 22.9 or newer required, and keywords.
- .gitignore: now covers .claude/ and skills-lock.json, which the ElevenLabs skill install added. They're tooling, not app code.

Checked:
- The server starts with npm start, with ElevenLabs enabled and the model loaded.
- Every page file loads with the right type, and bad uploads are still rejected.
- The start screen renders correctly in headless Edge.

Git, for you to sort out when you commit:
- Staged but deleted: you'd staged src/explain.js, which no longer exists.
- Old banner: hfw1c.webp shows as deleted; it now lives at docs/images/banner.webp.
- New files and folders: LICENSE, docs/, public/css/, public/js/, scripts/, src/config.js, src/prompts/, src/routes/, src/services/.
- Modified: README.md, .gitignore, .env.example, package.json, public/index.html, src/server.js.

Please run one real letter at http://[REDACTED]:3000, including "рд╕реБрдиреЗрдВ", to confirm nothing broke for you. Then step 6: README, screenshots and demo video. It's about 03:30.

You

let's do step 6

Agent

Screenshots first, since the README needs them.

For honest screenshots, the result screens will show real model output. I'll replay the actual answers from the test rounds through the real page code: a small separate demo server in my scratch folder fakes only the network responses, and nothing in the project changes. There's no English result for the KYC letter yet, so I'm running one now (about 60s):

The demo server replays saved results through the real page and fakes only the network calls. Starting it and taking the Hindi shots while the English run finishes:

Screenshots keep failing even with plain settings. Running Edge directly with its error output visible:

The screenshots were being written; Edge's launcher just exits a moment before the file lands, so my check ran too early. Adding a short wait to the helper and taking the set again:

The CSS and JS didn't load. That's a bug in my demo server: on Windows, join() returns backslash paths, so my folder check rejected every file. Fixing it:

Styles load now, but the screenshot is taken before the replayed result arrives. My earlier "failures" with the wait flag were only the file-check timing, so I'm restoring it:

At 600px the page has its proper right margin, so the clipping at 500px was a headless Edge quirk, not a layout bug. I'll skip the phone shot. Checking on the final-prompt Hindi run:

Still running. Ollama is sharing the CPU with the screenshot runs. While it finishes, I'll write the README, starting from your existing one:

The page wording is switched over. Now the README:

A pointless line in my edit script stopped it before it changed anything. Removing it and covering the remaining product mentions:

The remaining "letter" mentions are about actual letters (the story, the sample, the test documents), so they stay. Retaking the four screenshots with the new name:

Both screenshots show the new name. Checking the GitHub remote so the README's clone command has the real URL instead of a placeholder:

The app is now Document Explainer / рджрд╕реНрддрд╛рд╡реЗрдЬрд╝ рд╕рдордЭрд╛рдУ, and the README and screenshots for step 6 are done. The demo video is left for you to record.

Rename:
- App: page title, header, buttons and messages now say "document" / "рджрд╕реНрддрд╛рд╡реЗрдЬрд╝" in both languages. The Hindi grammar is adjusted for masculine рджрд╕реНрддрд╛рд╡реЗрдЬрд╝ ("рдЖрдкрдХрд╛ рджрд╕реНрддрд╛рд╡реЗрдЬрд╝тАж рдкрдврд╝рд╛ рдЬрд╛рддрд╛ рд╣реИ", "рджреВрд╕рд░рд╛ рджрд╕реНрддрд╛рд╡реЗрдЬрд╝ рд╕рдордЭреЗрдВ").
- README and package.json: updated too. "Letter" stays only where it means actual letters: the story about your grandfather, the sample letter and the test documents.

Screenshots in docs/images/:
- start.png, progress.png, result-hindi.png, result-english-dark.png, plus sample-letter.jpg (the made-up bank letter).
- Real output only: the results are the model's actual answers to the synthetic letter with the final prompt (59s in Hindi, 78s in English). I replayed them through the real page code, so no personal data appears anywhere.

README (README.md): this replaces your challenge notes (deadline, prize list), since the README is for repo visitors.
- The pitch, why and features: what it does, why you built it for your grandfather, and the main features.
- How it works: a diagram GitHub renders automatically, plus "why open-source AI".
- Setup and settings: setup steps (the clone command uses your real repo URL) and the settings table.
- "What it took to make the answers trustworthy": each failure from the four tuning rounds and its fix. This is also strong material for the post.
- Limits, project layout, credits and MIT license.

Your turn: the demo video (about 90 seconds)
1. Start Ollama and the server: npm start, then open http://[REDACTED]:3000.
2. Record with Win+Alt+R (Xbox Game Bar) or Clipchamp.
3. Shot list:
- Start screen in Hindi (3s). Point at the privacy line.
- Pick docs/images/sample-letter.jpg and let the progress run. Cut or speed up the middle of the wait, but say "about 1 minute on a laptop CPU".
- Result: scroll slowly past the urgency badge, deadline "10 рджрд┐рди рдмрд╛рдХреА", summary and steps.
- Press рд╕реБрдиреЗрдВ and let the ElevenLabs voice play about 10s.
- Switch to English, show the English result, then the dark mode toggle.
- Optional: one real document from your grandfather with personal details blurred, or his reaction.
4. Upload it to YouTube (unlisted is fine) or a similar site so the DEV post can embed it.

Ready for you to commit: README.md, docs/images/ (5 new images), public/index.html, public/js/i18n.js, package.json.

Next is step 7, the DEV post. I can start drafting it on DEV now while you record. Shall I?

You

let's rename to this name: Padh Ke Batao. i thnk it will be good one

Agent

Good name. "рдкрдврд╝ рдХреЗ рдмрддрд╛рдУ" means "read it and tell me", which is exactly what your grandfather says when he hands someone a letter. Renaming:
- Hindi title: рдкрдврд╝ рдХреЗ рдмрддрд╛рдУ
- English title: Padh Ke Batao, keeping the name itself rather than translating it
- README: add a line explaining what the name means

Then retaking the screenshots.

Retaking the screenshots with the new name:

The app is now Padh Ke Batao / рдкрдврд╝ рдХреЗ рдмрддрд╛рдУ.

  • Page: the Hindi title is "рдкрдврд╝ рдХреЗ рдмрддрд╛рдУ", the English title is "Padh Ke Batao", and the browser tab shows both.
  • README: the heading is renamed, with a line explaining what the name means: Hindi for "read it and tell me", what you say when you hand someone a letter you can't fully read. I first wrote "that's what my grandfather says", but you never told me that, so I took it out. If he really says it, it's a lovely opening line for the DEV post.
  • Screenshots: all four in docs/images/ are retaken with the new name.
  • Not renamed: the GitHub repo (build-for-a-friend) and the npm package name. Renaming the repo to padh-ke-batao is up to you, on GitHub under Settings тА║ Repository name. If you do, I'll update the clone URL in the README.

Ready for you to commit: public/js/i18n.js, public/index.html, public/css/styles.css, README.md, docs/images/*.png.

Next: record the demo video (shot list in my previous message), and I'll draft the DEV post. Shall I start the post now?

You

repo is renamed

Agent

Checking what the repo is called now, rather than guessing:

The repo is now Rahullkumr/padh-ke-batao. Updating the README clone commands, and the package.json name to match:

package-lock.json still carries the old name in its two top-level name fields. Updating just those:

The repo is now Rahullkumr/padh-ke-batao, and the project files point to it:
- README: the clone commands now use git clone https://github.com/Rahullkumr/padh-ke-batao.git and cd padh-ke-batao.
- package.json and package-lock.json: the package name is now padh-ke-batao. No old name is left in either file or the README.

Your local git remote still points at the old URL. GitHub redirects it, so pushing still works, but you can update it with:

git remote set-url origin https://github.com/Rahullkumr/padh-ke-batao.git

Your local folder is still called build-for-a-friend. That's harmless; rename it whenever you like, with the server stopped.

Ready for you to commit: README.md, package.json, package-lock.json, plus the earlier rename and screenshot changes.

Next: you record the demo video and I draft the DEV post. Shall I start the post?

You

@docs/images/result-hindi.png image is cut at the top, so regenrate it and replace it

Agent

The page's own "scroll to the result" runs after my scroll-back-to-top in the demo setup, so the header ends up cut off. A timing tweak won't fix that reliably. I'll turn off scrollIntoView in the demo page only (the app is untouched), then retake both result shots:

I regenerated docs/images/result-hindi.png, and it now shows the full header with nothing cut at the top. The English dark one (docs/images/result-english-dark.png) had the same problem, so I replaced it too.

The cause was my screenshot setup, not the app. When a result arrives, the page scrolls down to it, which cut off the header in the screenshot. I turned off that auto-scroll for screenshots only; the app itself still scrolls to the result as before.

Next: record the demo video (shot list two messages up), and I draft the DEV post. Shall I start the post?

You

why you are using optional wherever elevenlabs it there?

Agent

I've been calling ElevenLabs "optional" for two reasons:
1. It's technically true. The app runs without an ElevenLabs key: the Listen button falls back to a voice installed on the computer. Someone who clones the repo without a key still gets a working app.
2. The rules. The challenge requires open-source AI at the core. Calling ElevenLabs an add-on made it clear the core (Qwen in Ollama) is open.

But it undersells ElevenLabs. You're entering "Best Use of ElevenLabs", and it's the voice your grandfather actually hears. Here's where the word appears:

I'll reword it so ElevenLabs reads as the voice of the app, with the local-voice fallback mentioned once as a safety net rather than labelling the whole feature "optional":

"Optional" is gone. ElevenLabs is now presented as the app's voice everywhere, and the local-voice fallback is mentioned once as a safety net.

What changed:
- README diagram: ElevenLabs is a solid arrow, labelled "Listen: summary + steps only".
- "Works offline": now says reading and explaining need no internet, the ElevenLabs voice does, and without a connection Listen switches to a local voice.
- Setup heading: "ElevenLabs voice", no longer "Optional: тАж".
- Settings line: "Everything except the ElevenLabs key has a default".
- "Built with": ElevenLabs is listed as the "Hindi and English voice".
- .env.example, src/server.js, src/routes/speak.js: comments reworded the same way.

The open-source requirement is still clearly met: Qwen in Ollama does all the reading and explaining. I've noted this framing for the DEV post too.

Ready for you to commit: README.md, .env.example, src/server.js, src/routes/speak.js.

You

in the mermaid diagram optional is still visible

Agent

The README file on disk no longer has "optional" anywhere. The diagram's last line now reads:

  B -->|Listen: summary + steps only| E[ElevenLabs voice]

A search of the whole project (excluding node_modules and the agent tooling) also finds no "optional".

So what you're seeing is probably an older copy:
- VS Code preview: close and reopen the Markdown preview. If the README tab shows an unsaved dot, that open copy is older than the file on disk. Revert it with File тА║ Revert File.
- GitHub: it shows the last version you pushed, so the change appears there after you commit and push README.md.

If it still says "optional" after reopening, send me a screenshot of where you see it.

You

I have special skill: archify installed for architecture/other diagrams, can you use that instead of this mermaid diagram

Agent

Using archify. Architecture is the right type here: components plus a "this laptop" vs "internet" boundary. Reading its schema and one example first:

Source-file links need a pinned public commit, and this code isn't pushed yet. Removing the sources fields, which is the fix the validator offers:

Two labels sit on top of their boxes. Applying the validator's suggested fix, labelDy +24, to both:

Showcase validation passed: all 9 checks, 0 errors, 0 warnings. That spec is now frozen. Running the skill's one-time update check, then delivering the HTML:

Delivery passed: 9/9 checks, exit 0. Running archify's browser check on the delivered HTML. It also captures screenshots, which I need for the README image:

The viewer marks its controls with no-print, so I can hide them with one CSS rule. Then a 2├Ч capture at 1440├Ч900:

Version 2 passes showcase validation too (9/9, 0 errors, 0 warnings). Delivering it and re-running the browser check:

Both checks pass, and text is now larger (smallest label 9.0px, up from 7.4px). Capturing clean 2├Ч images in light and dark:

Both versions are cleanly cropped. Replacing the Mermaid block in the README with a <picture> element, so GitHub shows the light or dark image to match the reader's theme:

The Mermaid diagram in the README is replaced with an archify architecture diagram. GitHub shows the light or dark version to match each reader's theme.

What it shows:
- Main path, left to right: Document photo тЖТ Browser page (resize to ~900k px) тЖТ POST /explain тЖТ Fastify server ([REDACTED]:3000) тЖТ image + prompt тЖТ Ollama (qwen3.5:4b, with the Ollama logo).
- Listen branch: the Listen button sends POST /speak to the server, which sends text only over HTTPS to ElevenLabs. A dashed arrow shows the fallback to the local OS voice.
- The boundary: "This laptop: the photo never leaves" around everything local, and a separate "Internet" box around ElevenLabs.
- Three summary cards: what stays on the laptop, what goes to ElevenLabs, and how streaming works.

Checks:
- Archify validation: 9/9 checks passed with 0 errors and 0 warnings.
- Archify browser check: no overflow, and labels readable at every size from 1440├Ч900 to 2048├Ч1320.
- My own look: I checked the final images by eye.
- Layout fix: my first version had a big empty band at the top, so I moved everything up. That also made the smallest text larger (7.4px to 9px).

Files:
- docs/images/architecture-light.png and architecture-dark.png, at 2├Ч resolution for sharp text.
- docs/architecture.archify.json: the source, so the diagram can be regenerated later.
- README.md: a <picture> element replaces the Mermaid block, with alt text for screen readers and a credit line for archify.

Limits:
- No clickable file links: archify can link each box to its source file, but only with a pinned public commit. Your latest code isn't pushed yet, so I left the links out.
- No SVG: archify's viewer can export one, but only through its Export button in a browser, which I can't click from here. PNG is what the README needs anyway.

Ready for you to commit: README.md, docs/architecture.archify.json, docs/images/architecture-light.png, docs/images/architecture-dark.png.

You

[screenshot attached]
in archify image the heading should be one line (marked in red arrow)

Agent

The title wraps because archify's viewer limits its width to leave room for the toolbar. I already hide the toolbar in the README capture, so I'll lift that width limit in the same capture-only stylesheet:

Found it: .header reserves padding-right: 29.5rem for the toolbar. Setting it to 0 in the capture copy, plus white-space: nowrap on the title, then recapturing both themes:

The title is on one line now. Removing the reserved space made the layout wider, so I'll adjust the crop to match:

The diagram title is now on one line in both the light and dark images. The README picks them up automatically, since the file names (architecture-light.png, architecture-dark.png) are unchanged.

The diagram content and archify validation are unchanged. This only affects how the image is captured: the width archify reserves for its toolbar is removed in the copy I screenshot, and I widened the crop to fit. I checked the new light image and nothing is cut off at the edges.

Ready for you to commit: docs/images/architecture-light.png, docs/images/architecture-dark.png.

You

do: 1. DEV post draft

Agent

Checking two things before drafting: whether you already have a draft for this on DEV (so I don't create a duplicate), and the challenge's exact submission template. Loading the tools:

The draft is saved. Checking whether your commit is pushed, since the post's images load from GitHub:

Written for: DEV readers and the challenge judges, who score writing quality most heavily.

Your DEV post is saved as an unpublished draft: https://dev.to/rahullkumr/padh-ke-batao-an-ai-that-reads-my-grandfathers-letters-to-him-on-a-laptop-with-no-gpu-14o4-temp-slug-4672643/edit

The draft is marked as written with some AI help ("some AI"). That's the honest setting, since I drafted it.

Title: Padh Ke Batao: an AI that reads my grandfather's letters to him, on a laptop with no GPU

What's in it:
- Format: it follows the official template, with the tags devchallenge, weekendchallenge, hf26challenge and ai.
- Story: the four tuning rounds told with real numbers:
- 147s down to 50s: shrinking the photo, because reading the image was the slow part.
- The invented threat: a harmless Hindi letter turned into "you may be punished", plus a fake phone number, until the resolution was fixed.
- Language drift: English and Roman-letter Hindi creeping into Hindi mode.
- Temperature: "рдЕрдкрд╣рд░рдг" for "claimed" at a high setting, broken output at a low one.
- Honest about limits: the remaining small mistakes, such as "Sunrise" becoming "рд╕реВрд░рдЬ".
- ElevenLabs: presented as the app's voice, with exactly what is and isn't sent. It's entered under "Best Use of ElevenLabs".
- Development vs runtime: it states that Claude Code was used for building only, and the app itself uses no closed AI.
- Images: your push is done, and I checked that the images load from GitHub.

Four TODOs only you can fill, marked in bold brackets so they can't slip through:
1. Opening: a real detail about your grandfather (what you call him, a letter that started this).
2. Demo: the video link.
3. Handing It Over: his reaction, ideally his own words in Hindi with a translation, plus a photo if he's comfortable. This earns the bonus points.
4. Agent session: I'll replace the session embed with this full session right before you publish.

Before publishing, preview it on DEV and read it once aloud. If anything sounds unlike you, change it, since judges reward an authentic voice. When the TODOs are filled, tell me and I'll upload this session and update the embed. Then you publish, before 12:29 IST.

You

Prize categories: "Best Use of ElevenLabs".

Writing quality counts most in judging, so it helps to have a few real details from you. Answer whatever you can now and I'll leave the rest as clearly marked blanks:

What do you call him (Nanaji, Dadaji, something else), and which language does he read best? ==> Nana
Was there a real moment that made you build this, like a letter he worried about or a deadline someone nearly missed? he got SIR notice but was not able to read it, so I created this solution. This solution can be usful for so many people in real world
Does he actually say "padh ke batao"? yes

Agent

That's a great detail. The Election Commission SIR press note was one of your six test documents, so the story connects directly: the notice he couldn't read is the same voter-roll drive the app was tested on. Fetching the current draft first, in case you've edited it on DEV:

The draft is updated. The README's name line can now be personal too, since you confirmed Nana says it:

The DEV draft now opens with your real story, and the README's name line matches.

What changed in the draft:
- Opening: "My Nana (my mother's father)тАж". Then the SIR notice he couldn't read, the kind where missing a step can keep your name off the voter list. Then him asking, as he always does, "padh ke batao", which is where the name comes from.
- New line: Nana is far from the only one. In millions of Indian homes an official letter waits on a shelf until someone has time to read it, so the app is for anyone in that spot.
- Nana throughout: he's "Nana" in the body; the title keeps "grandfather" so readers anywhere understand it.
- Test connection: the Election Commission test document is now "the press note about the same SIR drive".
- README (README.md): the name line now reads "It's what my Nana says whenever he hands someone a letter he can't read."

Open the draft: https://dev.to/rahullkumr/padh-ke-batao-an-ai-that-reads-my-grandfathers-letters-to-him-on-a-laptop-with-no-gpu-14o4-temp-slug-4672643/edit

Please check the SIR paragraph is accurate for his notice. If it was specifically the enumeration form a booth officer brought, you could say so, which makes it even more concrete.

Still left:
1. Demo video (you): record it and paste the link into the Demo section.
2. Handing It Over (you): let Nana try it, ideally on his real SIR notice. His reaction in his own words is the strongest part the post can have.
3. Session (me): before you publish, I save this full session and swap the embed.

Ready for you to commit: README.md.

You

Handing It Over (you): let Nana try it, ideally on his real SIR notice. His reaction in his own words is the strongest part the post can have.

he loved it specially the Hindi speech, as he speak Hindi mostly

Agent

That's a perfect detail for the post, and it supports the ElevenLabs entry directly: the voice is the part he liked most. I'll fill in the Handing It Over section using only what you've told me, with no invented quotes or details. Fetching the current draft first in case you've edited it:

The draft's Handing It Over section is filled in, using only what you told me:

I showed it to Nana, and he loved it. The part he liked most was the Hindi voice. He speaks Hindi most of the time, so hearing the explanation out loud, in his own language, is what made it click for him. He didn't have to read anything, not even the screen.

That changed how I see the app. I'd treated the Listen button as an add-on; for Nana it's the main way in. It's also why I'm comfortable with the one thing that leaves the laptop: a few sentences of summary sent to ElevenLabs, in exchange for a voice he actually enjoys listening to.

The ElevenLabs entry under Prize Categories now ends with "It was also Nana's favourite part."

Optional, but it makes this section stronger: if you remember anything he actually said, even one short Hindi line like "рдЕрдм рдХрд┐рд╕реА рд╕реЗ рдкрдврд╝рд╡рд╛рдирд╛ рдирд╣реАрдВ рдкрдбрд╝реЗрдЧрд╛", add it with a translation. A photo of him using it, if he's comfortable, helps too. I won't make up a quote for him.

Draft: https://dev.to/rahullkumr/padh-ke-batao-an-ai-that-reads-my-grandfathers-letters-to-him-on-a-laptop-with-no-gpu-14o4-temp-slug-4672643/edit

Left before you publish (deadline 12:29 IST):
1. Demo video (you): record it, upload it, and replace the Demo TODO with the link.
2. Session (me): when you're ready to publish, tell me. I'll save this full session to DEV and swap the embed.
3. Final read (you): preview the post on DEV, check that no "[TODO" is left, then publish.

You

Padh Ke Batao: an AI that reads my grandfather's letters to him, on a laptop with no GPU

i think it should be documents instead of letters

Agent

I'll change only the title, through DEV's API, so the body stays exactly as it is:

The response was too large to display, so I'm pulling out just the key fields to confirm the change worked and the post is still a draft:

The update went through with the new title. The response doesn't show the disclosure field, so I'm reading it back from DEV to confirm:

I've also saved that lesson in memory: for any future edits to this post, I'll only use the DevRelay update tool, which keeps the disclosure setting.

You

when the document is being processed we should not show the file choosing section

Agent

Syntax is valid. Checking it visually. This also changes how the progress screen looks, so docs/images/progress.png (used in the README and the post) needs retaking anyway:

The photo picker now disappears while a document is being processed. Only the progress card shows: thumbnail, progress bar, reading/explaining steps and the timer. When the result or an error appears, the picker comes back in its compact form at the top.

I checked both states with screenshots.
- What changed in code: three lines in public/js/app.js. The picker is hidden when processing starts and shown again in showResult and showError.
- Edge cases: dropping a photo onto the page mid-processing does nothing, because the drop target is hidden too. Switching language mid-processing still restarts the explanation in the new language, as before.
- Screenshot: I retook docs/images/progress.png so the README and the DEV post match the new behaviour. Its URL is the same, so the post updates once you push.

On the result screen, the top "рджрд╕реНрддрд╛рд╡реЗрдЬрд╝ рдХреА рдлреЛрдЯреЛ рдЪреБрдиреЗрдВ" button and the bottom "рджреВрд╕рд░рд╛ рджрд╕реНрддрд╛рд╡реЗрдЬрд╝ рд╕рдордЭреЗрдВ" button do the same thing. Should I hide the top one on the result screen as well, so there's only one clear "next document" button?

Ready for you to commit: public/js/app.js, docs/images/progress.png, plus the earlier README.md change. The running server picks up the page change on a browser refresh, so your demo recording can use it straight away.

You

[screenshot attached] remove these sections also, red arrow pointing

Agent

Removing the "рдкрдврд╝рдирд╛ / рд╕рдордЭрд╛рдирд╛" step pills. The status line above the bar ("рдЪрд┐рдЯреНрдареА рдкрдврд╝ рд░рд╣реЗ рд╣реИрдВтАж" / "рдЖрд╕рд╛рди рднрд╛рд╖рд╛ рдореЗрдВ рд▓рд┐рдЦ рд░рд╣реЗ рд╣реИрдВтАж") already says the same thing. I'll delete them from the HTML, the JS that updates them, their CSS, and their text strings, so nothing is left behind:

All the step-pill code is gone and the scripts still parse; pulse was only used by the pills. Retaking the progress screenshot to check the look and update the README image:

The "рдкрдврд╝рдирд╛ / рд╕рдордЭрд╛рдирд╛" pills are gone. The progress card now shows only the thumbnail, the status line, the bar and the timer. I checked it in light and dark mode.

I removed them from four places, so nothing unused is left:
- public/index.html: the pill list.
- public/js/app.js: the two lines that switched them between active and done.
- public/css/styles.css: their styles and the pulse animation, which only they used.
- public/js/i18n.js: their Hindi and English text.

The status line above the bar still tells you the stage: "рджрд╕реНрддрд╛рд╡реЗрдЬрд╝ рдкрдврд╝ рд░рд╣реЗ рд╣реИрдВтАж" (reading), then "рдЖрд╕рд╛рди рднрд╛рд╖рд╛ рдореЗрдВ рд▓рд┐рдЦ рд░рд╣реЗ рд╣реИрдВтАж" (writing).

I also retook docs/images/progress.png, so the README and the DEV post show the new card once you push.

Refresh http://[REDACTED]:3000 to see it before recording the demo; no server restart is needed.

Commit message, if you want one:
```
fix: simplify the processing screen

  • Hide the photo picker while a document is being read; it returns (compact) with the result or an error
  • Remove the reading/explaining step pills; the status line above the progress bar already shows the stage
  • Update the progress screenshot

Co-Authored-By: Claude Opus 5.5 <[REDACTED]>
```

Ready for you to commit: public/js/app.js, public/index.html, public/css/styles.css, public/js/i18n.js, docs/images/progress.png, plus the earlier README.md change.

You

https://youtu.be/VO9MM8JvKCQ also add it in the readme

Agent

The draft matches my last version, so you haven't made editor changes that I would overwrite. I'll replace the Demo TODO with the YouTube embed, keeping the AI disclosure at some_ai.

The video is now in the post's Demo section. Next I'll add it to the README, using the thumbnail as a clickable preview, since GitHub can't play videos inline.

Prize Categories

  • Best Use of ElevenLabs. ElevenLabs is the voice of the app: the рд╕реБрдиреЗрдВ / Listen button reads each explanation aloud in natural Hindi or English, at a slower pace for an elderly listener. Only the summary and steps are sent, never the photo or the details list, and the audio is cached so replays are free. It was also Nana's favourite part.

Top comments (0)