This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Arjun is a composite of friends we all have. He's generous, he's first to say "I'll pay my share", and when you remind him he says "haan bhai, kal bhej dunga" ("yeah, I'll send it tomorrow"). He means it. Then the reminder sinks under forty messages in a group chat, and nobody wants to be the one who asks twice.
It's never about the money. His share of a dinner someone else paid for, the ₹500 a friend lent him, his part of the flat's electricity and broadband: the intention survives, the reminder doesn't.
PaidYet is a Telegram bot that runs on my laptop and does the remembering for him:
- Arjun sends a bill photo, forwards a message, or types "Rahul ko 500 dene hai Friday tak" ("I owe Rahul 500 by Friday").
- Gemma, running locally, reads it. The bot replies with what it understood and the whole reminder plan, and he taps Save.
- A Temporal workflow reminds him on that plan until he taps Paid. Snooze postpones; it never pulls a reminder earlier.
- When I know he owes something, like our trip split, I add it for him as the admin, and every message says who added it.
He's on Android, he already lives in Telegram, and he writes in Hinglish. So there's nothing to install and nothing new to learn.
Demo
PaidYet is a private bot that runs on my laptop and handles people's bills, so there's no public instance to click. Here's the flow, captured from the live bot. To run your own copy, clone the repo and run uv run poe up; the README has the setup.
All the bills in these screenshots are synthetic: fictional companies, fake numbers.
Code
prateek11rai
/
paidyet
Telegram bot that reads bills and IOUs with local Gemma (Ollama) and keeps reminding until you pay. Durable reminders on Temporal, agent traces in Sentry. Built for the DEV Hacktoberfest 2026 Weekend Challenge.
PaidYet
Keeps reminding your friend until they've actually paid.
PaidYet is a Telegram bot that runs on your laptop, for the friend who always says "haan bhai, kal bhej dunga" ("yeah, I'll send it tomorrow") and then forgets. Send it a bill photo or a line like "Rahul ko 500 dene hai Friday tak". Gemma, running locally in Ollama, reads it, and you tap Save. Then a durable Temporal workflow keeps reminding you until you tap Paid. Sentry traces every step without seeing a single bill.
What it is
A private reminder bot for money you owe: your share of a dinner, a trip someone else paid for, a friend's loan, the flat's electricity and broadband. Reading a bill happens on your machine. The reminders survive the laptop sleeping, restarting or crashing. An admin (you, if you run it for friends) can add dues for a friend, and…
How I Built It
Telegram ──▶ bot ──▶ Temporal: ReminderWorkflow, one per due
├─ activity: read ──▶ Gemma 4 in Ollama (127.0.0.1) ──▶ code checks the answer
├─ activity: delete the photo / forget the text
├─ wait for Save or Fix (signals)
└─ durable timers: remind on the plan (Paid / Snooze signals, status query)
Sentry: one trace from the Telegram update through every activity; Gemma calls as gen_ai spans
It's Python, built on python-telegram-bot, Temporal, Ollama running gemma4:e4b, and Sentry. It runs on an Apple M4 with 16 GB. uv run poe up starts Ollama (bound to 127.0.0.1), a Temporal dev server and the bot, and Ctrl-C stops all three.
Gemma reads, code decides
Gemma fills a JSON schema through Ollama's structured output, at temperature 0 with thinking off. Then plain code takes over. It parses the dates in Asia/Kolkata ("Friday tak", "kal", "3 din mein" and "in 2 minutes" are all parsed by code, not the model), rejects past due dates and odd amounts, and decides whether the draft can be saved at all.
I made four synthetic bills, each with a trap: a bill date next to the due date, a late-fee amount, a "next invoice" date, and an insurance notice with a grace-period end. I also wrote six Hinglish and English IOU lines. Here's how my iterations scored on those 10 samples:
| Version | Due date right | Silently wrong due date | Warm photo |
|---|---|---|---|
| Baseline prompt | 8/10 | 2 | 4.8 s |
| Thinking on | 8/10 | 2 | ~33 s |
| Model lists every date with its label | 8–9/10 | 1 ("12 Oct" became 12 Dec) | 8–11 s |
| Gemma quotes the deadline, code parses it, one hinted retry | 9/10 | 0 | ~7 s |
The final version scores 10/10 on amount and kind and 9/10 on payee. The one due date it still misses is the insurance notice: Gemma keeps reaching for the grace-period date. But now it fails safely. The bot says "The date I found looks like the end of a grace period, not the due date" and offers only Fix, no Save. A silently wrong due date is the one failure a payment reminder can't afford; a blocked one costs a tap.
On the live bot, a bill photo took 7.4 s (1,098 tokens in, 99 out), a typed correction 3.4 s, and a Hinglish line 3 to 12 s. Gemma 4 e4b holds about 6.6 GB of RAM while loaded (5.5 GB of weights plus a 1 GB vision projector), so PaidYet keeps it loaded for 30 minutes after each read instead of paying a ~20 s cold load every time.
This matches what Taranpreet Kaur describes in Why an LLM Alone Cannot Do Invoice Extraction Yet: models show "silent confidence", returning a plausible value instead of nothing, so you want "model for reading, code for checking". Gemma reported confidence: 1.0 on the grace-period date.
Temporal makes the reminder durable
Every due is one ReminderWorkflow, started the moment a message arrives:
- Reading the bill is an activity with retries. A slow model load or a bad JSON reply just gets another try, up to three.
-
The input is deleted in a
finally. The photo goes on success and after the final failed retry. What Arjun typed never enters Temporal's history: the workflow carries a random key, and the text waits in memory until it's read. -
Save, Fix, Paid and Snooze are signals;
statusis a query. Before signalling, the bot asks the workflow who added a due (before Save) and who owes it (before Paid). - Reminders are durable timers on a plan the confirm message spells out:
| Due | Reminders |
|---|---|
| A date ("Friday tak", a bill due 12 Oct) | 10 AM three days before (if it's at least 3 days away), 7 PM the evening before, 10 AM and 7 PM on the day, then 10 AM daily while overdue |
| A time at least an hour away | 30 minutes before and at the deadline, then 10 AM daily |
| Under 10 minutes (demo mode) | Every N minutes, 5 times, then 10 AM daily |
Nothing goes out between 10 PM and 8 AM, except demos. A reminder that has to beat the deadline moves earlier, to 9:30 PM. After 30 days overdue, PaidYet sends one last message, and the due then waits quietly in /due. A long-overdue due continues-as-new, so its history stays short.
My laptop sleeps and restarts, so I tested exactly that. I saved "Rahul ko 500 dene hai in 3 minutes", then pressed Ctrl-C on the whole stack, Temporal included, before the reminder was due:
| Time (IST) | What happened |
|---|---|
| 02:55:39 | Saved; durable timer set for 02:57:21 |
| 02:56:05 | Ctrl-C: bot and worker down |
| 02:57:21 | Temporal records the timer firing, with no worker to act on it |
| 02:57:25 | Temporal stopped too; the pending work exists only in .data/temporal.db
|
| 02:57:40 |
uv run poe up again |
| 02:57:45.6 | The reminder arrives: 24.6 s late, not lost |
There are 138 unit tests (about 6 s). They run the workflow in Temporal's time-skipping test server, so "nudge daily for a month" takes milliseconds. Five integration tests run the real stack (Gemma, Temporal, Telegram) in about 2 minutes.
Sentry shows the agent's work, not the bill
Each read is an invoke_agent span. Inside it are a gen_ai.chat span per Gemma call, carrying token counts and how Ollama's time split between loading the model, reading the prompt and generating, and execute_tool spans for the two code checks (review_answer, validate_draft). So the trace shows exactly where code overruled the model. Temporal interceptors carry one trace from the Telegram update into every activity, including a reminder that fires days later.
What the timings showed me:
-
The first message after idle was slow, and it wasn't Gemma's fault. One cold call took 29.5 s, and 22.2 s of that was Ollama loading the model; the read itself took about 7 s. The fix was a warm-up at start-up and a 30-minute
keep_alive. -
The hinted retry is visible. A read where code asked Gemma again shows two chat spans and
retry_reason: no_due_date. It costs a few seconds on the bills that need it and nothing on IOUs. - Cost: ₹0. Tokens are counted for latency, not billing.
Sentry never sees the bill. Spans carry model names, token counts, timings and outcome categories; never prompts, replies, amounts or names. A unit test runs a full read through Sentry configured exactly like the app, against an in-memory transport, and checks that none of that leaks. A scrubber also removes anything shaped like a bot token.
Why Does Open Innovation Matter?
Because these are bills. A bill photo has a name, an address, an account number and what someone owes. With an open-weight model on my laptop, no AI company reads Arjun's electricity bill, keeps it, or bills me for reading it. A reminder about ₹500 shouldn't cost API credits, and here each read is a few seconds of my laptop's time.
Open also meant I could change how the agent behaves. I turned thinking off when it cost 6–8x the latency for no accuracy gain, pinned temperature to 0, and a smaller machine can switch to gemma4:e2b with one variable. Anyone can run their own copy: a friend can self-host it with their own BotFather token and be their own admin.
One caveat, stated plainly: Telegram bot chats are not end-to-end encrypted, so the photo Arjun sends also stays in his Telegram chat, on Telegram's servers. My claim isn't "it never leaves the laptop". It's that no model provider ever gets it. On my side, the photo exists only in a temp folder under a random name until Gemma has read it, and what's stored afterwards is the confirmed title, payee, amount, due date and kind.
What I Got Wrong
-
I trusted the model with the date. Version one read the insurance notice's grace period end as the due date, at
confidence: 1.0. Asked plainly to list the dates, Gemma reads "Renew by: 20-10-2026" perfectly. The mistake was in choosing, so choosing moved to code. - I assumed thinking would help. It made a photo take ~33 s instead of ~5 s and fixed nothing.
- I made the model do more work, and it got worse. Asking it to list every date with its label turned "12 Oct" into 12 December, a silent error. I dropped it within the hour.
- My first reminder schedule would have nagged at 3 AM. An insurance renewal three weeks out got its first reminder the night before, and "in 5 hours" repeated every 5 hours around the clock. A code review caught it, and the schedule above replaced it.
- I quoted the bill into my own database. An early "blocked" message read "The date I found is labelled 'Grace period ends 19-11-2026'", and that message is stored in Temporal's history. It now names a category instead.
- One slow Telegram response killed the bot at start-up. python-telegram-bot defaults to zero retries for its polling bootstrap. It now retries until it gets through.
- Ctrl-C left Ollama and Temporal running. My task runner escalates a Ctrl-C to SIGKILL after 1.6 s, and a Temporal worker with live workflows takes about that long to stop, so my cleanup died with it. The fix was to have the runner hand the process over instead of supervising it.
The full log, with every number and dead end, is in docs/BUILD_LOG.md.
My Agent Session
I built PaidYet with Claude Code as my pair programmer, with DevRelay pulling the challenge rules live. The session is curated into three parts: thinking and long tool output are trimmed, and local paths and IDs are scrubbed.
Part 1: the kickoff, and getting Gemma to read bills I could trust
PaidYet: build session kickoff
We're building PaidYet tonight and submitting it tonight. It's a personal project on my personal GitHub (prateek11rai). We're in <local path> a fresh git init on main with no commits. It currently contains:
- .claude/skills/devrelay-* and .claude/devrelay-rules.md (DevRelay);
- docs/img/logo.svg, Lucide's "badge-indian-rupee" (₹) recoloured green, with its ISC notice in docs/img/logo-LICENSE.txt.
Links: what we're working towards
- Challenge page (rules, deadline, judging, FAQ, the source of truth): https://dev.to/challenges/hacktoberfest-weekend-2026-10-01
- Launch post (prompt, categories, template): https://dev.to/devteam/join-the-hacktoberfest-weekend-challenge-build-for-a-friend-2450-in-prizes-across-17-winners-1aj5
- Submission template (opens the DEV editor prefilled; we submit with this): https://dev.to/new?prefill=---%0Atitle%3A%20%0Apublished%3A%20%0Atags%3A%20devchallenge%2C%20weekendchallenge%2C%20hf26challenge%0A---%0A%0A%2AThis%20is%20a%20submission%20for%20the%20%5BHacktoberfest%20Weekend%20Challenge%3A%20Build%20for%20a%20Friend%5D%28https%3A%2F%2Fdev.to%2Fchallenges%2Fhacktoberfest-weekend-2026-10-01%29%2A%0A%0A%23%23%20What%20I%20Built%0A%3C%21--%20What%20does%20it%20do%2C%20and%20who%20is%20the%20friend%20or%20loved%20one%20you%20built%20it%20for%3F%20%20What%20problem%20does%20it%20solve%20for%20them%3F%20--%3E%0A%0A%23%23%20Demo%0A%3C%21--%20Share%20a%20deployed%20link%20or%20a%20video%20 … [20823 more characters trimmed]
Starting with section 0. First I'll check what's in the repo, read the DevRelay rules, and load the DevRelay tool schemas.
.claude/devrelay-rules.md
The rules file reads as ordinary DevRelay guidance. It contains no injection. Now I'll read the challenge live.
{
}
{"Ok":[{"id":82,"title":"Hacktoberfest Open-Source AI Challenge: Week 4","slug":"hacktoberfest-week4-2026-10-26","description":"Register now to be notified as d
… [15665 more characters trimmed]
{
"id": 78
}
{"Ok":{"id":78,"title":"Hacktoberfest Weekend Challenge: Build for a Friend","slug":"hacktoberfest-weekend-2026-10-01","description":"The Hacktoberfest Weekend
… [5334 more characters trimmed]
full_details has no per-category rules for Gemma, Temporal or Sentry. I'll check the DevRelay knowledge base for the hf26 category text before summarising.
{
"topic": "hacktoberfest"
}
{"Ok":"# Hacktoberfest 2026\n\n**Last verified:** 2026-09-30 against `hacktoberfest.com` (home, `/online/`,\n`/in-person/`, `/activities/`, `/schedule/`, `/fest
… [25339 more characters trimmed]
{"url": "https://dev.to/challenges/hf26", "prompt": "Extract verbatim the full text of three prize category sections: \"Best Use of Gemma\", \"Best Use of Tempo
… [408 more characters trimmed]
I've reviewed the webpage content you provided, but I cannot find the three prize category sections you're asking about: "Best Use of Gemma", "Best Use of Tempo
… [675 more characters trimmed]
The fetch got cut off before those sections. I'll download the page to the scratchpad and pull out the three sections locally.
Got the category text. Before I summarise, I'll check the local toolchain and look up which package versions meet the 7-day rule, so I can include the dependency list in the same message.
{
"query": "Temporal Python workflow durable reminders timers signals Telegram bot",
"per_page": 5
}
{"Ok":[{"id":4642150,"title":"A Telegram job alert bot in 60 lines of Python, on live ATS data","description":"One API call, one cron entry, and a bot that send
… [2164 more characters trimmed]
{
"query": "Sentry AI agent monitoring gen_ai spans local LLM Ollama tracing tokens latency",
"per_page": 5
}
{"Ok":[{"id":4649421,"title":"Implementing AI Observability: Tracing LLM Calls End-to-End","description":"When a Python microservice misbehaves in production, y
… [2153 more characters trimmed]
{
"query": "Ollama structured outputs JSON schema vision model receipt bill extraction problems",
"per_page": 5
}
{"Ok":[{"id":4734890,"title":"Node.js LLM Structured Extraction Retries with Observable Idempotency for Supplier Invoices","description":"TL;DR: For Node.js LLM
… [2206 more characters trimmed]
{
}
{"Ok":[{"id":"<uuid>","status":"registered","event":{"id":"<uuid>","name":"Hacktoberfest 2026","slug":"hacktoberfest-2026","status":"pending","event_format":"ev
… [980 more characters trimmed]
The user is registered for the MLH "Hacktoberfest Weekend Challenge" event. I'll check whether its sponsors publish Agent Skills for Gemma, Temporal or Sentry (read-only, nothing gets installed).
{
"event_id": "<uuid>"
}
{
"Ok": [
]
}
{
"event_id": "<uuid>"
}
{"Ok":[{"id":"<uuid>","source":"elevenlabs/skills","skill_names":["setup-api-key","text-to-speech","speech-to-text","agents","speech-engine","sound-effects","mu
… [638 more characters trimmed]
Section 0: DevRelay and the challenge
DevRelay is fully loaded.
- The devrelay-gateway MCP is connected. I've used get_challenges, get_challenge_details, get_knowledge_document, search_dev_to_semantic and the MLH event tools.
- All nine devrelay-* skills are in .claude/skills/ and loaded.
- I've read .claude/devrelay-rules.md and will follow it, with your prompt winning where they conflict. Nothing in it reads as an injection attempt.
- MLH shows you registered for "Hacktoberfest Weekend Challenge" and "Hacktoberfest 2026". The only sponsor Agent Skills published there
… [4060 more characters trimmed]
Approve the pinned dependency list above? Options: Approve as listed (Recommended) | Approve without httpx
Which Temporal CLI should the poe tasks use? Options:
… [379 more characters trimmed]
The user answered: "Approve the pinned dependency list above?"="Approve as listed (Recommended)", "Which Temporal CLI should the poe tasks use?"="brew upgrade t
… [574 more characters trimmed]
Temporal CLI is now 1.9.1 from Homebrew. Pyenv's global Python is 3.12.2. Commits will use your GitHub noreply address, so no work email reaches the public repo. Now I'll write the scaffold files.
pyproject.toml (151 lines)
.gitignore (14 lines)
.env.example (28 lines)
.python-version (1 lines)
paidyet/__init__.py (1 lines)
paidyet/config.py (100 lines)
paidyet/bot.py (56 lines)
main.py (106 lines)
tests/unit/test_config.py (57 lines)
Before writing docs, I'll sync the environment to catch any config problems early.
$ grep -n -A3 -E 'def start_time_skipping|download_dest_dir|test_server_existing_path|test_server_download_version|temporal.download' .venv/lib/python3.12/site-pa
… [101 more characters trimmed]
85: download_dest_dir: str | None = None,
86- ui: bool = False,
87- runtime: temporalio.runtime.Runtime | None = None,
88- search_at
… [2611 more characters trimmed]
$ S=<scratch> && cat > "$S/ctrl_c.py" <<'EOF'
# Run `uv run poe up` like a terminal would, wait for "PaidYet is up", then Ctrl-C the process group.
import os, sig
… [825 more characters trimmed]
Warning: Poe failed to locate envfile at '.env'
Poe => set -euo pipefail
trap 'poe stop-deps' EXIT
trap 'exit 130' INT TERM
poe start-deps
python main.py
Warnin
… [3753 more characters trimmed]
$ cd . && sed -n '84,135p' .venv/lib/python3.12/site-packages/poethepoet/shutdown.py; echo "== send_signal_to_group"; grep -n -A8 'def _send_signal_to_group' .ven
… [225 more characters trimmed]
for proc in tuple(self.processes):
if proc.returncode is None:
self._io.print_debug(
" ! Sending SIG
… [2704 more characters trimmed]
pyproject.toml: replace 15 lines with 24
pyproject.toml: replace 3 lines with 4
$ S=<scratch> && uv run python "$S/ctrl_c.py" 180; echo "== after: pid files"; ls .data/run/; echo "== processes"; pgrep -fl 'ollama serve|temporal server' || ech
… [21 more characters trimmed]
Warning: Poe failed to locate envfile at '.env'
Warning: Poe failed to locate envfile at '.env'
ollama: started on 127.0.0.1:11434 (pid 61424)
ollama: model
… [725 more characters trimmed]
$ S=<scratch> && (OLLAMA_HOST=127.0.0.1:11434 nohup ollama serve >"$S/ext-ollama.log" 2>&1 & echo $! > "$S/ext-ollama.pid"); sleep 3; echo "external ollama pid $(
… [328 more characters trimmed]
external ollama pid 61714
ollama: reusing the instance already serving 127.0.0.1:11434 (poe will not stop it)
ollama: model gemma4:e4b ready
temporal: start
… [153 more characters trimmed]
LICENSE (21 lines)
CLAUDE.md (27 lines)
README.md (43 lines)
docs/BUILD_LOG.md (56 lines)
Writing step 1's docs is done. Now I'll stage and check what's about to be committed (no .env, .data or DevRelay files), then commit, create the private repo and push.
$ S=<scratch> mkdir -p "$S"; cd "$S"
cat > base.css <<'EOF'
body{margin:0;background:#e9e6df;font-family:-apple-system,Helvetica,Arial,sans-serif;color:#222}
.pag
… [5715 more characters trimmed]
Command did not complete within its 120s timeout and was moved to the background (ID: bnht5bjc3). Output is being written to: <scratch> You will be notified whe
… [276 more characters trimmed]
Headless Chrome is hanging. I'll check which renders finished.
All four rendered. Some Chrome helper processes from the scratch profiles are still alive, so I'll stop those, then look at a bill to check it's legible.
<scratch>
The bill looks realistic and every distractor is legible. Some Chrome helpers survived pkill, so I'll check what they are.
New PIDs keep appearing because the first backgrounded render loop is still running and starts a new Chrome each time I kill one. I'll stop that background task.
paidyet/extract.py (312 lines)
main.py should use the shared sweep now instead of its own copy:
main.py: replace 11 lines with 1
main.py: replace 3 lines with 4
main.py: replace 2 lines with 2
Now a scratch harness that runs every sample through the real Gemma call and the validator, timing each case. The first call includes loading the model, so I'll report cold and warm times separately.
$ S=<scratch> && cat > "$S/accuracy.py" <<'EOF'
"""Run every sample through Gemma + validation; print per-field accuracy and latency."""
import asyncio, json, sys
… [2405 more characters trimmed]
Traceback (most recent call last):
File "<scratch>", line 6, in <module>
from paidyet.extract import ask_gemma, validate
ModuleNotFoundError: No module na
… [13 more characters trimmed]
$ S=<scratch> && PYTHONPATH=. uv run python "$S/accuracy.py" gemma4:e4b 2>&1 | tee "$S/acc-e4b.txt"
electricity.png 29.5s load=22.2s in= 631 out= 72 bill 'Electricity Bill' 'Sahyadri Power Distribution Ltd.' 1240.0 2026-10-09T00:
… [2118 more characters trimmed]
paidyet/extract.py: replace 7 lines with 8
paidyet/extract.py: replace 6 lines with 11
paidyet/extract.py: replace 9 lines with 10
$ S=<scratch> && sed -i '' -e 's/^MODEL = .*/MODEL = sys.argv[1] if len(sys.argv) > 1 else "gemma4:e4b"\nTHINK = len(sys.argv) > 2 and sys.argv[2] == "think"/' -e
… [321 more characters trimmed]
=== v2, think off
electricity.png 7.7s load= 0.1s in= 753 out=116 bill 'Electricity Bill' 'Sahyadri Power Distribution Ltd.' 1240.
… [2365 more characters trimmed]
$ S=<scratch> && echo "=== v2, think on" && PYTHONPATH=. uv run python "$S/accuracy.py" gemma4:e4b think 2>&1 | tee "$S/acc-e4b-v2-think.txt"
=== v2, think on
electricity.png 35.8s load= 0.0s in= 755 out=792 bill 'Electricity Bill' 'Sahyadri Power Distribution Ltd.' 1240.0
… [2315 more characters trimmed]
paidyet/extract.py: replace 2 lines with 4
paidyet/extract.py: replace 3 lines with 5
paidyet/extract.py: replace 6 lines with 10
paidyet/extract.py: replace 4 lines with 49
paidyet/extract.py: replace 7 lines with 12
paidyet/extract.py: replace 4 lines with 3
Still on step 2: I'm re-measuring accuracy with the new self-check, where code reviews Gemma's answer and asks once more with a hint if a bill's date looks wrong.
$ S=<scratch> && cat > "$S/accuracy2.py" <<'EOF'
"""Accuracy and latency of extract.read() (ask, code review, at most one hinted retry, validate)."""
import async
… [2223 more characters trimmed]
=== run 1
Traceback (most recent call last):
File "<scratch>", line 6, in <module>
from paidyet.extract import read
File "paidyet/extract.py", line 212
… [553 more characters trimmed]
My edit left the tail of a section-header comment stuck onto a line. Fixing it:
paidyet/extract.py
paidyet/extract.py: replace 3 lines with 4
$ S=<scratch> && for i in 1 2; do echo "=== run $i"; PYTHONPATH=. uv run python "$S/accuracy2.py" gemma4:e4b 2>&1 | tee "$S/acc-e4b-v3-run$i.txt"; done
=== run 1
electricity.png 10.3s calls=2 bill 'Sahyadri Power Distribution Ltd.' 1240.0 2026-10-09T00:00 problems=[] OK
broadband.png
… [3038 more characters trimmed]
paidyet/extract.py: replace 8 lines with 16
paidyet/extract.py: replace 2 lines with 4
Now the code side: labelled_deadline() picks the single date whose printed label says it's a deadline, and that choice wins over the model's own due_date.
paidyet/extract.py: replace 18 lines with 45
paidyet/extract.py: replace 5 lines with 6
$ S=<scratch> && PYTHONPATH=. uv run python "$S/accuracy2.py" gemma4:e4b 2>&1 | tee "$S/acc-e4b-v4.txt"
electricity.png 7.1s calls=1 bill 'Sahyadri Power Distribution Ltd.' 1240.0 2026-10-09T00:00 problems=[] OK
broadband.png
… [1460 more characters trimmed]
paidyet/extract.py: replace 2 lines with 3
$ S=<scratch> && PYTHONPATH=. uv run python "$S/accuracy2.py" gemma4:e4b 2>&1 | tee "$S/acc-e4b-v5.txt"
electricity.png 10.9s calls=1 bill 'Sahyadri Power Distribution Ltd.' 1240.0 2026-10-09T00:00 problems=[] OK
broadband.png
… [1434 more characters trimmed]
paidyet/extract.py: replace 16 lines with 8
paidyet/extract.py: replace 3 lines with 0
paidyet/extract.py: replace 20 lines with 32
paidyet/extract.py: replace 12 lines with 13
paidyet/extract.py: replace 1 lines with 1
paidyet/extract.py: replace 1 lines with 1
paidyet/extract.py: replace 3 lines with 3
$ S=<scratch> && for i in 1 2; do echo "=== run $i"; PYTHONPATH=. uv run python "$S/accuracy2.py" gemma4:e4b 2>&1 | tee "$S/acc-e4b-v6-run$i.txt"; done
=== run 1
electricity.png 10.7s calls=2 bill 'Sahyadri Power Distribution Ltd.' 1240.0 2026-10-09T00:00 problems=[] OK
broadband.png
… [3080 more characters trimmed]
tests/unit/test_extract.py (267 lines)
$ uv run poe test 2>&1 | grep -v Warning | tail -25
Poe => pytest tests/unit
============================= test session starts ==============================
platform darwin -- Python 3.12.2, pytest-9.1.1, pluggy
… [545 more characters trimmed]
All 79 unit tests pass. Now I'll add the step-2 numbers and dead ends to BUILD_LOG, then commit and push.
Part 2: the Temporal workflow, and a code review that changed the reminder schedule
paidyet/extract.py: replace 1 lines with 18
paidyet/workflows.py (345 lines)
I left a pointless _Unused dataclass at the end just to justify an import. Removing it, along with the import:
paidyet/workflows.py: replace 6 lines with 1
paidyet/workflows.py: replace 2 lines with 2
paidyet/workflows.py: replace 13 lines with 13
paidyet/workflows.py: replace 2 lines with 7
The snooze handling has an edge case: a snooze that arrives just as a reminder fires could be lost. I'll restructure the loop to apply any new snooze at the top of each pass.
paidyet/workflows.py: replace 19 lines with 17
Now store.py, the SQLite read model that /due will use. It holds only the confirmed fields and IDs.
paidyet/store.py (96 lines)
That _row helper came out garbled (a dead if False else branch). Cleaning it up:
paidyet/store.py: replace 6 lines with 4
paidyet/bot.py (426 lines)
Same mistake again: a stray assert to justify an unused import. Removing both:
paidyet/bot.py: replace 5 lines with 1
paidyet/bot.py: replace 1 lines with 1
paidyet/activities.py (204 lines)
The cleanup in fix_draft is convoluted ("draft" in locals()). I'll make it explicit: forget the correction text after a successful read, or on the final failed attempt.
paidyet/activities.py: replace 9 lines with 11
paidyet/activities.py: replace 1 lines with 10
paidyet/workflows.py: replace 6 lines with 9
Now wiring the worker and the bot together in main.py. I'm also adding a background model warm-up at start, so the first message doesn't pay the ~22 s cold load.
main.py
main.py: replace 23 lines with 47
main.py: replace 6 lines with 11
tests/unit/test_workflow.py (291 lines)
$ uv run pytest tests/unit/test_workflow.py -x -q 2>&1 | tail -40; echo "== cache"; ls -la .pytest_cache/temporal 2>/dev/null | head; ls "${TMPDIR:-/tmp}" | grep
… [63 more characters trimmed]
Pytest: 0 passed, 1 failed
Failures:
1. [FAIL] test_save_remind_snooze_remind_paid
tests/unit/test_workflow.py:156: in test_save_remind_snooze_remind_paid
… [277 more characters trimmed]
tests/unit/test_workflow.py: replace 1 lines with 9
tests/unit/test_workflow.py: replace 2 lines with 2
$ uv run pytest tests/unit/test_workflow.py -q 2>&1 | tail -30
Command did not complete within its 600s timeout and was moved to the background (ID: bs4cgv1c4). Output is being written to: <scratch> You will be notified whe
… [165 more characters trimmed]
The run hung past 10 minutes, so something is blocking. I'll check how far it got.
$ pkill -f temporal-test-server; for t in test_save_remind_snooze_remind_paid test_flaky_reads_are_retried test_final_read_failure_still_deletes_the_input test_in
… [372 more characters trimmed]
=== test_save_remind_snooze_remind_paid
[2m2026-10-03T20:00:43.756227Z[0m [33m WARN[0m [2mtemporalio_sdk_core::worker[0m[2m:[0m Temporal Server 1.16.0 o
… [1047 more characters trimmed]
Run one at a time, nothing hangs, but three fail. I'll get the full tracebacks.
$ for t in test_save_remind_snooze_remind_paid test_flaky_reads_are_retried test_final_read_failure_still_deletes_the_input; do echo "=== $t"; uv run python -m py
… [168 more characters trimmed]
=== test_save_remind_snooze_remind_paid
tests/unit/test_workflow.py:174: in test_save_remind_snooze_remind_paid
assert fakes.payloads("edit_reminder")
… [622 more characters trimmed]
paidyet/workflows.py: replace 14 lines with 15
tests/unit/test_workflow.py: replace 2 lines with 3
tests/unit/test_workflow.py: replace 4 lines with 6
$ pkill -f temporal-test-server; for t in test_save_remind_snooze_remind_paid test_flaky_reads_are_retried test_final_read_failure_still_deletes_the_input test_fi
… [650 more characters trimmed]
test_save_remind_snooze_remind_paid 1 passed in 0.96s
test_flaky_reads_are_retried 1 passed in 0.3
… [494 more characters trimmed]
$ python3 - <<'EOF'
import re
p = "tests/unit/test_workflow.py"
s = open(p).read()
s, n = re.subn(r" worker, handle = await start\((env, fakes[^)]*)\)\n asy
… [1079 more characters trimmed]
replaced 9
9
collected 88 items
tests/unit/test_config.py ............. [ 14%]
tests/unit/test_extract.py ....................
… [274 more characters trimmed]
tests/unit/test_bot.py (109 lines)
$ uv run poe test 2>&1 | grep -v 'WARN\|Warning' | tail -15
tests/unit/test_workflow.py ......... [100%]
=================================== FAILURES ==================================
… [871 more characters trimmed]
A real bug: .capitalize() lowercased the rest of the line, giving "was due fri 9 oct". It should uppercase only the first letter.
paidyet/bot.py: replace 1 lines with 2
$ uv run poe test 2>&1 | grep -v 'WARN\|Warning' | tail -3 && echo "=== smoke: poe up without .env (worker must pass sandbox validation)" && S=<scratch> && uv run
… [152 more characters trimmed]
tests/unit/test_workflow.py ......... [100%]
============================= 106 passed in 2.25s =============================
… [1043 more characters trimmed]
poe up now starts the worker, which means the workflow passed Temporal's sandbox validation, and shutdown is still clean. All 106 unit tests pass. Logging step 3 and committing:
Code review from Prateek's planning session. A second Claude instance read the whole repo at f61a137.
This review comes BEFORE any live testing. Don't start the stack, and don't run the integration tests or anything else live yet.
Work through the points below, run the unit tests, commit, and then report back. We'll do real-keys testing after that.
Keep as is (no churn needed):
- privacy handling (keys instead of text, photo discard in finally, start-up sweep);
- the who-may-tap-what checks via the status query;
- pid-file stop-deps;
- loopback guards;
- Sentry scrubbing;
- parameterised SQL.
== A. Reminder timing. This is a product decision, so confirm the proposed schedule with Prateek before implementing it.
What the code does today (workflows.py first_reminder / next_nudge / nudge_interval, extract.py parse_relative):
- Date dues ("Friday tak", a bill due 12 Oct):
- a reminder at 7 PM the evening before;
- then 10 AM on the due day;
- then 10 AM daily until Paid.
There's nothing earlier, however far away the due date is: an insurance renewal due in 3 weeks gets its first reminder the night before.
- Timed dues exist only for "in N seconds/minutes/hours"; clock times like "Friday 5 PM" aren't supported.
- They fire exactly AT the deadline, with no lead time.
- They then repeat every (due − sent_at), capped at a day, around the clock. Real "in 5 hours" nudges every 5 h, including 3 AM. Real "in 5 minutes" nudges every 5 min forever.
- So t
… [3338 more characters trimmed]
Points A–C are a review I'll act on. You've asked me to confirm A (the reminder schedule) with you before implementing it, and some of B depends on that schedule. So the schedule question comes first; everything else gets implemented right after. Nothing runs live.
My take on A: I agree with all five points. There are three gaps the proposal leaves open, and I've folded my suggested answers into the options below.
Approve the proposed reminder schedule (A1–A5), with these gaps filled: timed dues 10 min–1 h away get one reminder at the deadline and then daily 10 AM; demo d
… [464 more characters trimmed]
Your questions have been answered: "Approve the proposed reminder schedule (A1–A5), with these gaps filled: timed dues 10 min–1 h away get one reminder at the d
… [907 more characters trimmed]
Approved, including the gap-fillers, with pre-deadline quiet-hour reminders moved earlier to 21:30. I'll reread the current workflows.py and bot.py before changing the schedule code.
paidyet/workflows.py
I'll check how much RAM Gemma uses while resident (item 10). The model blob size on disk is a fair proxy, and reading it doesn't need the stack.
paidyet/workflows.py (517 lines)
paidyet/workflows.py: replace 10 lines with 10
paidyet/workflows.py: replace 3 lines with 3
paidyet/workflows.py: replace 8 lines with 9
paidyet/workflows.py: replace 7 lines with 7
Now activities.py: split mark_paid (SQLite, bounded retries) from show_paid (Telegram), treat Telegram's BadRequest as final, and give the "added" and "last" messages only a ✅ Paid button.
paidyet/activities.py
$ python3 - <<'EOF'
p = "paidyet/activities.py"
s = open(p).read()
s = s.replace('''def _gone(what: str) -> ApplicationError:''', '''def _refused(e: Exception) -
… [3796 more characters trimmed]
31:def _refused(e: Exception) -> ApplicationError:
71: self.show_paid,
150: raise _refused(e) from e
162: raise _refused(e) fro
… [211 more characters trimmed]
paidyet/bot.py
$ sed -i '' 's/ return f"every morning until it'"'"'s paid"/ return "every morning until it'"'"'s paid"/' paidyet/bot.py && grep -n "def capture" -A
… [29 more characters trimmed]
371:async def capture(update: Update, context: ContextTypes.DEFAULT_TYPE) -> None:
372- msg, user, settings = update.effective_message, update.effective_user
… [1541 more characters trimmed]
$ python3 - <<'EOF'
p = "tests/unit/test_bot.py"
s = open(p).read()
s = s[: s.index("# --- schedule rules (pure functions in workflows.py) ---")]
s = s.replace("f
… [7856 more characters trimmed]
31 passed in 0.14s
All 31 schedule and message tests pass on the first run. Now the workflow tests. I'll reread the current file before editing.
tests/unit/test_workflow.py
$ python3 - <<'EOF'
p = "tests/unit/test_workflow.py"
s = open(p).read()
def rep(old, new, count=1):
global s
assert s.count(old) == count, (s.count(old)
… [13015 more characters trimmed]
# Do call (this throws an RPCError on failure)
raise ValueError(f"Unrecognized service {service}")
E temporal_sdk_bridge.RPCError: (1, 'Timeou
… [1002 more characters trimmed]
An RPC "Timeout expired" at line 387, inside one of the new tests. I'll find which one.
$ sed -n '380,392p' tests/unit/test_workflow.py; uv run pytest tests/unit/test_workflow.py -q -p no:logging --tb=no 2>&1 | grep -E 'PASSED|FAILED|passed|failed' |
… [99 more characters trimmed]
Exit code 1
async def test_sqlite_failures_are_bounded_and_surfaced(env):
now = await now_ist(env)
fakes = Fakes(reads=[draft(now + timedelta(days=5))],
… [644 more characters trimmed]
$ uv run python -m pytest tests/unit/test_workflow.py -q -rf --tb=line -p no:logging 2>&1 | grep -v WARN | tail -8
Failures:
E temporalio.service.RPCError: Timeout expired
raise RuntimeError("database is locked")
RuntimeError: database is locked
2. [FAIL] t
… [232 more characters trimmed]
$ grep -nE "^(E | |tests/)|Error|assert" <local path> Support/rtk/tee/<id>_pytest.log | grep -v 'site-packages\|^[0-9]*: [a-z_]*\.' | grep -E 'test_workflow
… [27 more characters trimmed]
3:E temporal_sdk_bridge.RPCError: (1, 'Timeout expired', b'')
6:E temporalio.service.RPCError: Timeout expired
$ for t in test_sqlite_failures_are_bounded_and_surfaced test_after_30_days_overdue_it_says_so_once_and_waits_in_due; do echo "=== $t"; uv run python -m pytest "t
… [105 more characters trimmed]
=== test_sqlite_failures_are_bounded_and_surfaced
Pytest: 1 passed
=== test_after_30_days_overdue_it_says_so_once_and_waits_in_due
Pytest: 1 passed
Both pass on their own, so something earlier in the suite leaks into them. The likeliest culprit is the test that runs just before them. I'll check pairs.
$ T=tests/unit/test_workflow.py; for prev in test_a_snooze_never_moves_the_next_reminder_earlier test_one_sentry_trace_from_the_telegram_update_to_every_activity
… [234 more characters trimmed]
test_a_snooze_never_moves_the_next_reminder_earlier + sqlite Pytest: 1 passed, 1 failed
test_one_sentry_trace_from_the_telegram_update_to_every_activi
… [116 more characters trimmed]
$ cp tests/unit/test_workflow.py tests/unit/test_zz_debug.py && python3 - <<'EOF'
p = "tests/unit/test_zz_debug.py"
s = open(p).read()
s = s.replace('''
… [783 more characters trimmed]
BEFORE 1
AFTER 5
1 passed in 0.90s
$ uv run python -m pytest "tests/unit/test_workflow.py::test_sqlite_failures_are_bounded_and_surfaced" "tests/unit/test_workflow.py::test_flaky_reads_are_retried"
… [77 more characters trimmed]
1:Pytest: 2 passed
$ /usr/bin/time -p uv run python -m pytest "tests/unit/test_workflow.py::test_sqlite_failures_are_bounded_and_surfaced" -q -p no:logging -p no:cacheprovider 2>&1
… [306 more characters trimmed]
2:1 passed in 0.32s
3:real 0.70
3:E temporal_sdk_bridge.RPCError: (1, 'Timeout expired', b'')
6:E temporalio.service.RPCError: Timeout expired
10:Completing
… [1590 more characters trimmed]
$ python3 - <<'EOF'
p = "tests/unit/test_workflow.py"
s = open(p).read()
old = ''' finally:
with contextlib.suppress(Exception): # by ID, so a
… [1051 more characters trimmed]
============================= 137 passed in 6.09s ==============================
All 137 pass. Since that was a timing bug, I'll run the suite twice more to check it's stable, not lucky.
$ for i in 1 2; do uv run poe test 2>&1 | grep -E 'passed|failed' | tail -1; done
============================= 137 passed in 5.62s ==============================
============================= 137 passed in 5.73s =============================
… [1 more characters trimmed]
Updating README, the post and BUILD_LOG to the new schedule. The README also gets the keep-alive RAM trade-off, /cancel, "Snooze only postpones", and the 30-day stop.
Now the BUILD_LOG entry for this round:
$ cat >> docs/BUILD_LOG.md <<'EOF'
## Sun 4 Oct, ~03:00: code review before live testing
A second Claude instance reviewed the repo at `f61a137`. Prateek confir
… [5215 more characters trimmed]
============================= 137 passed in 6.15s ==============================
M README.md
M docs/BUILD_LOG.md
M docs/devto-post.md
M main.py
M paidyet/a
… [244 more characters trimmed]
I worked through all ten review points and committed them as add5481. 137 unit tests pass, stable over three runs. I didn't start the stack or run the integration tests; I only checked that they still collect.
A. Reminder schedule (built as you approved it)
- Date dues: 10 AM three days before (only if saved at least 3 days ahead), 7 PM the evening before, 10 AM and 7 PM on the day, then 10 AM daily while overdue.
-
Timed dues:
- at least 1 hour away: 30 minutes before, then at the deadline;
- 10 minutes to 1 hour away: once, at the deadline;
- under 10 minutes (demo): … [3265 more characters trimmed]
Part 3: live testing in Telegram, and killing the stack mid-wait
env ready. The review fixes at add5481 look good. Prateek has filled in .env. Don't read or print it; rely on the app's own start-up log lines ("not set" / "enabled") to confirm what's present.
Now run the live testing, in this order:
1. Start the stack with uv run poe up, as a background process you can watch and stop. Confirm from the logs:
- Ollama and Temporal are up;
- the model is loaded;
- Sentry is enabled;
- Telegram is polling, with the bot's @username shown.
Then tell Prateek the bot's username, so he can open it in Telegram.
2. Run uv run poe test (unit), then uv run poe test-integration against the running stack. Integration tests message only the admin, prefixed "[test]", so Prateek should see those arrive. Fix any failures with small commits, then re-run.
3. Live run with Prateek in Telegram, covering the section 12 checklist:
- bill photo → confirm → reminder → Snooze → Paid. Use a demo due like "in 2 minutes" so it finishes quickly;
- a Hinglish IOU text;
- /due;
- the admin's /remind for arjun (only if a second account is in ALLOWED_USERS; otherwise say so and skip it);
- an unknown user being turned away;
- the kill-and-restart proof (stop the app mid-wait, restart it, and the reminder still arrives late);
- .data/tmp/ empty afterwards;
- Sentry traces arriving with no bill contents.
How to work with Prateek:
- Whenever you need him to act on his phone, stop and end your turn with a short "👉 Your
… [609 more characters trimmed]
Starting the stack. I'll launch uv run poe up in its own process session, logging to a scratch file, so I can watch it and later stop it with a real Ctrl-C (SIGINT to the group), which lets the cleanup trap run.
$ S=<scratch> && cat > $S/launch_up.py <<'EOF'
"""Start `uv run poe up` in its own session (like a terminal tab), logging to a file; record its process group."""
… [662 more characters trimmed]
nothing running yet
started, process group 55180
ollama: started on 127.0.0.1:11434 (pid 55211)
ollama: model gemma4:e4b ready
temporal: started on 127.0.0.
… [13162 more characters trimmed]
$ pgrep -fl 'ollama serve|temporal server|python main.py' || echo "all stopped (cleanup ran)"; ls .data/run/; for i in 1 2 3; do curl -sS -m 15 -o /dev/null -w "a
… [118 more characters trimmed]
all stopped (cleanup ran)
(empty)
api.telegram.org: HTTP 302, connect 0.168822s, total 0.516032s
api.telegram.org: HTTP 302, connect 0.174461s, total 0.523979s
… [62 more characters trimmed]
$ python3 - <<'EOF'
p = "main.py"
s = open(p).read()
old = " await app.updater.start_polling()\n"
new = " # Keep retrying the bootstrap: one
… [1348 more characters trimmed]
============================= 137 passed in 6.18s ==============================
started, process group 56854
ollama: started on 127.0.0.1:11434 (pid 56874)
o
… [596 more characters trimmed]
$ S=<scratch> for _ in $(seq 60); do grep -q "Ollama:" $S/up.log && break; sleep 1; done; grep "Ollama:" $S/up.log; echo "=== integration"; /usr/bin/time -p uv ru
… [137 more characters trimmed]
2026-10-04 02:31:16,169 INFO paidyet: Ollama: gemma4:e4b loaded in 0 ms
=== integration
E AssertionError: timed out; last status: Status(state='sc
… [543 more characters trimmed]
$ python3 - <<'EOF'
def edit(path, old, new):
s = open(path).read()
assert s.count(old) == 1, (path, old[:60])
open(path, "w").write(s.replace(old, ne
… [2530 more characters trimmed]
============================= 137 passed in 6.17s ==============================
d826acc Set next_at before sending a reminder; time the warm-up
Committed (d826acc). Restarting the stack to load the new workflow code, using a real Ctrl-C so the cleanup runs, then re-running the integration tests.
$ S=<scratch> kill -INT -- -$(cat $S/up.pgid); for _ in $(seq 30); do pgrep -f 'python main.py|temporal server|ollama serve' >/dev/null || break; sleep 1; done; p
… [462 more characters trimmed]
stopped cleanly
2026-10-04 02:35:25,451 INFO temporalio.worker._worker: Beginning worker shutdown, will wait 0:00:00 before cancelling activities
ollama: stoppe
… [559 more characters trimmed]
All 5 integration tests pass against the real stack (2 min 18 s), using real Gemma, Temporal, SQLite and Telegram. You should have seen [test] messages arrive in your chat. The model warmed up in 7 s.
So far, live testing has found and fixed three things (three small commits):
- Start-up crash on one slow Telegram response. The polling bootstrap now retries until it succeeds, with longer timeouts.
- A status race around sending a reminder. The next reminder is now computed before sending, so a Snooze is never checked against a stale time.
- A misleading warm-up log line ("
… [927 more characters trimmed]
I sent it from the admin chat. I sent an electricity bill photo and then saved it. Check if the records have been saved and check if Temporal is awaiting now on the reminders to the workflow. Also tell me which endpoint I should check on localhost so that I can see my Temporal workflow and be able to attach screenshots later in the article when we require
I'll check three places: the app log for the read, the SQLite record, and the workflow in Temporal (its status and the pending timer). I'll print only the fields we need, with no Telegram IDs.
Everything is saved correctly, and Temporal is durably waiting on the next reminder.
What happened with your photo:
- Read: Gemma took one call, about 7.4 s. Code accepted the result: Electricity Bill · ₹1,240 · Sahyadri Power Distribution Ltd. · due Fri 9 Oct.
- Photo deleted: discard_input ran right after the read, and .data/tmp/ is empty. An earlier check of mine said "1 file", but that was a shell-display quirk; Python confirms 0 files.
- Saved to SQLite: one row with the confirmed fields only, paid_at empty.
- Temporal: workflow r-<id> is RUNNING in state
… [1636 more characters trimmed]
And said it is due in 2 minutes and then saved it as well.
I'll check the Fix, the save and the timer on the new workflow, and whether the first reminder has gone out.
The Fix worked. Gemma applied "due in 2 minutes" in 3.4 s, making it due 02:45:19, and the first reminder is scheduled for that exact moment. I'll watch for it to go out:
The reminder went out at 02:45:19, right on time; the workflow says 1 sent and the message is posted. The next demo reminder is due at 02:47:19.
👉 Your turn (admin account), on the ⏰ reminder message:
1. Tap 😴 Snooze → Tomorrow. You should get a toast "Snoozed until Mon 5 Oct, 10 AM", and the reminder should change to add "😴 Snoozed until Mon 5 Oct, 10 AM", still with Paid / Snooze. If a second ⏰ arrives before you tap (they repeat every 2 min until snoozed), use the newest one; the older one loses its buttons.
2. Then tap ✅ Paid on that same message. It should become "✅ Paid
… [131 more characters trimmed]
I tapped Pay and I think it is all done. Confirm it once and then we will be done with the flow.
I'll confirm from Temporal's history (did the snooze and the paid signal both land, and did the workflow finish?), SQLite and the temp folder.
$ echo "=== workflow history (IST)"; temporal workflow show -w r-<id> --address 127.0.0.1:7233 -o json 2>/dev/null | python3 -c '
import json,sys
from datetime im
… [1441 more characters trimmed]
=== workflow history (IST)
02:43:08 activity read_input
02:43:08 activity discard_input
02:43:09 activity show_confirm
02:43:09 timer 604800s
02:43:19 SIGNAL fi
… [877 more characters trimmed]
Paid is confirmed end to end:
- Workflow r-<id> is COMPLETED, result "paid". Its history: read → discard → confirm → fix → updated confirm → save → a 107 s timer → reminder sent at 02:45:20 → paid at 02:45:26 → SQLite marked paid → Telegram message edited to "✅ Paid".
- SQLite: paid_at is set and next_at cleared, so it's gone from /due. The Tuesday electricity reminder is still pending, as it should be.
- .data/tmp: 0 files.
One gap: no snooze signal is in the history. Paid came 6 s after the reminder, so Snooze hasn't been tested live yet. That leaves one
… [679 more characters trimmed]
We are not testing the other user one. We can save that piece but we can assume that it will all work fine because it worked fine for us in the admin profile.
Understood: /remind and the unknown-user check are skipped live. I'll mark them that way in the report. Both are unit-tested, but I won't claim they were verified live. I'll check whether the Hinglish IOU has come in yet.
Live testing is done; /remind and the unknown-user check stay skipped. If the Hinglish IOU / Snooze check you were on has landed, report it in one line. Then work through the steps below in order. Stop at each 👉 point and wait for Prateek.
1. Two last checks
- Kill-and-restart proof: do it yourself. Use a pending demo reminder, or ask Prateek once for a fresh "Rahul ko 500 dene hai in 3 minutes". Stop the app mid-wait with a real Ctrl-C, wait until the due time has passed, then start it again. Confirm from Temporal's history that the reminder went out late and wasn't lost, and note the timings for BUILD_LOG and the post.
- Sentry: check that traces arrived. Tell Prateek exactly which view to open and which trace (by operation name and time) so he can confirm by eye that there's no bill text. He takes the screenshots.
[… the remaining steps of this plan are omitted from the curated session]
First, whether the Hinglish IOU landed:
The Hinglish IOU never landed: there's no new read in the log or SQLite. So Snooze is still untested by tap. It's covered by the live integration test, which sent the snooze signal directly, plus the unit tests.
For the kill-and-restart proof I need a pending demo reminder. There isn't one, and I can't create one myself, because typed text only lives in the running app's memory. So I'm asking you once.
👉 Your turn (admin account)
1. Type: Rahul ko 500 dene hai in 3 minutes
2. Tap ✅ Save, and tell me right away.
Then I'll press Ctrl-C on the whole stack (app and Temporal) whi
… [226 more characters trimmed]
sent and saved, go ahead
Finding the new workflow and its due time, then stopping the whole stack with a real Ctrl-C:
$ S=<scratch> pgrep -fl 'python main.py|temporal server|ollama serve' | cut -c1-80; echo "=== last log lines (no tracebacks)"; grep -v '^\s' $S/up.log | grep -v '
… [113 more characters trimmed]
60105 ollama serve
60123 temporal server start-dev --db-filename .data/temporal.d
60162 python main.py
=== last log lines (no tracebacks)
3:temporal: started on
… [1726 more characters trimmed]
It just arrived: the read finished at 02:55:35. Gemma took 12.5 s this time, so my first check ran a few seconds too early. The Save may not have landed yet. Checking the workflow state:
The Hinglish IOU landed and is saved: Rahul · ₹500 · due 02:57:21, state scheduled, 0 reminders sent. Pressing Ctrl-C on the whole stack now, with the timer about 85 s from firing:
$ S=<scratch> uv run poe stop-deps 2>&1 | tail -2; pgrep -fl 'python main.py|temporal server|ollama serve' || echo "everything down at $(date +%H:%M:%S), before t
… [46 more characters trimmed]
ollama: stopped (pid 60105)
temporal: stopped (pid 60123)
everything down at 02:57:25, before the 02:57:21 due time
$ S=<scratch> while [ "$(date +%H%M%S)" \< "025830" ]; do sleep 2; done; echo "restart at $(date +%H:%M:%S)" | tee -a $S/kr-times.txt; : > $S/up.log; python3 $S/l
… [924 more characters trimmed]
(eval):1: condition expected: <
restart at 02:57:40
temporal: started on 127.0.0.1:7233, UI at http://127.0.0.1:8233 (pid 78752)
2026-10-04 02:57:44,765 INFO pa
… [252 more characters trimmed]
$ temporal workflow show -w r-<id> --address 127.0.0.1:7233 -o json 2>/dev/null | python3 -c '
import json,sys
from datetime import datetime, timedelta
d=json.loa
… [1052 more characters trimmed]
3:02:55:35.294 activity read_input
6:02:55:35.330 activity discard_input
9:02:55:36.108 activity show_confirm
12:02:55:36.119 timer started 604800s
13:02:55:38.
… [471 more characters trimmed]
It's poe itself; uv isn't involved. Reading poe's shutdown logic to see what escalates:
$ sed -n '1,60p;135,200p' .venv/lib/python3.12/site-packages/poethepoet/shutdown.py | grep -v '^\s*$'
1:from __future__ import annotations
3:import asyncio
4:import contextlib
5:import os
6:import signal
7:import subprocess
8:import sys
9:from typing import TYPE
… [4257 more characters trimmed]
$ S=<scratch> R=.; cd $S/sigtest && python3 - <<'EOF'
import re, pathlib
p = pathlib.Path("pyproject.toml"); s = p.read_text()
new_up = """[tool.poe.tasks.up]
cmd
… [521 more characters trimmed]
run 1 (uv run poe up, Ctrl-C to the group)
Poe => bash -c 'set -uo pipefail
# poe hands this process over (use_exec), so nothing SIGKILLs a slow shutdown.
… [1403 more characters trimmed]
Before testing the fix on the real stack, which means restarting it, I'll check whether you've already tapped Snooze/Paid on the pending reminder:
$ temporal workflow show -w r-<id> --address 127.0.0.1:7233 -o json 2>/dev/null | python3 -c '
import json,sys
from datetime import datetime, timedelta
d=json.loa
… [848 more characters trimmed]
02:55:38 SIGNAL save
02:57:45 activity send_reminder
02:59:22 activity send_reminder
now 03:01:16
$ S=<scratch>
stop_old() { kill -INT -- -$(cat $S/up.pgid) 2>/dev/null; for _ in $(seq 20); do pgrep -f 'python main.py' >/dev/null || break; sleep
… [961 more characters trimmed]
1. old stack fully down
2026-10-04 03:01:38,383 INFO paidyet: Telegram: polling as @<bot>
2026-10-04 03:01:38,383 INFO paidyet: PaidYet is up. Ctrl-C to stop.
2
… [397 more characters trimmed]
$ cat >> docs/BUILD_LOG.md <<'EOF'
## Sun 4 Oct, ~02:30–03:05: live testing (real Telegram, Sentry, Temporal, Gemma)
**Integration tests: 5/5 pass** against the
… [4142 more characters trimmed]
============================= 137 passed in 6.49s ==============================
652eb61 poe up: hand the process over so Ctrl-C can't SIGKILL the cleanup
Prize Categories
- Best Use of Gemma: Gemma 4 (e4b) runs locally in Ollama and reads bill photos and Hinglish IOUs through structured output. Code checks every answer and asks Gemma again, once, when it spots a known mistake.
-
Best Use of Temporal: One durable workflow per due. The Gemma read is a retried activity and the input is deleted in a
finally. Save, Fix, Paid and Snooze are signals, status is a query, and reminders survive a full stop and restart (24.6 s late, not lost). -
Best Use of Sentry Agent Tracing:
invoke_agent,gen_ai.chatandexecute_toolspans carry tokens and Ollama's load, prompt and generation timings, in one trace from Telegram through Temporal into every activity, with no bill contents in Sentry, enforced by a test.






Top comments (0)