DEV Community

Prateek Rai
Prateek Rai

Posted on

PaidYet: for the friend who says “kal bhej dunga”

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Arjun is a composite of friends we all have. He's generous, he's first to say "I'll pay my share", and when you remind him he says "haan bhai, kal bhej dunga" ("yeah, I'll send it tomorrow"). He means it. Then the reminder sinks under forty messages in a group chat, and nobody wants to be the one who asks twice.

It's never about the money. His share of a dinner someone else paid for, the ₹500 a friend lent him, his part of the flat's electricity and broadband: the intention survives, the reminder doesn't.

PaidYet is a Telegram bot that runs on my laptop and does the remembering for him:

  • Arjun sends a bill photo, forwards a message, or types "Rahul ko 500 dene hai Friday tak" ("I owe Rahul 500 by Friday").
  • Gemma, running locally, reads it. The bot replies with what it understood and the whole reminder plan, and he taps Save.
  • A Temporal workflow reminds him on that plan until he taps Paid. Snooze postpones; it never pulls a reminder earlier.
  • When I know he owes something, like our trip split, I add it for him as the admin, and every message says who added it.

He's on Android, he already lives in Telegram, and he writes in Hinglish. So there's nothing to install and nothing new to learn.

Demo

PaidYet is a private bot that runs on my laptop and handles people's bills, so there's no public instance to click. Here's the flow, captured from the live bot. To run your own copy, clone the repo and run uv run poe up; the README has the setup.

A phone photo of the synthetic electricity bill; the bot reads it back with the whole reminder plan and Save / Fix

A reminder for Rahul's ₹500 with the Snooze menu open: Tomorrow, In 3 days, Pick date

/due lists what's overdue and upcoming, and who added each one

All the bills in these screenshots are synthetic: fictional companies, fake numbers.

Code

GitHub logo prateek11rai / paidyet

Telegram bot that reads bills and IOUs with local Gemma (Ollama) and keeps reminding until you pay. Durable reminders on Temporal, agent traces in Sentry. Built for the DEV Hacktoberfest 2026 Weekend Challenge.

PaidYet logo

PaidYet

Keeps reminding your friend until they've actually paid.

License: MIT Python 3.12

PaidYet is a Telegram bot that runs on your laptop, for the friend who always says "haan bhai, kal bhej dunga" ("yeah, I'll send it tomorrow") and then forgets. Send it a bill photo or a line like "Rahul ko 500 dene hai Friday tak". Gemma, running locally in Ollama, reads it, and you tap Save. Then a durable Temporal workflow keeps reminding you until you tap Paid. Sentry traces every step without seeing a single bill.

What it is

A private reminder bot for money you owe: your share of a dinner, a trip someone else paid for, a friend's loan, the flat's electricity and broadband. Reading a bill happens on your machine. The reminders survive the laptop sleeping, restarting or crashing. An admin (you, if you run it for friends) can add dues for a friend, and…

How I Built It

Telegram ──▶ bot ──▶ Temporal: ReminderWorkflow, one per due
                       ├─ activity: read ──▶ Gemma 4 in Ollama (127.0.0.1) ──▶ code checks the answer
                       ├─ activity: delete the photo / forget the text
                       ├─ wait for Save or Fix                (signals)
                       └─ durable timers: remind on the plan  (Paid / Snooze signals, status query)
Sentry: one trace from the Telegram update through every activity; Gemma calls as gen_ai spans
Enter fullscreen mode Exit fullscreen mode

It's Python, built on python-telegram-bot, Temporal, Ollama running gemma4:e4b, and Sentry. It runs on an Apple M4 with 16 GB. uv run poe up starts Ollama (bound to 127.0.0.1), a Temporal dev server and the bot, and Ctrl-C stops all three.

Terminal: uv run poe up starting Ollama, Temporal and the bot

Gemma reads, code decides

Gemma fills a JSON schema through Ollama's structured output, at temperature 0 with thinking off. Then plain code takes over. It parses the dates in Asia/Kolkata ("Friday tak", "kal", "3 din mein" and "in 2 minutes" are all parsed by code, not the model), rejects past due dates and odd amounts, and decides whether the draft can be saved at all.

I made four synthetic bills, each with a trap: a bill date next to the due date, a late-fee amount, a "next invoice" date, and an insurance notice with a grace-period end. I also wrote six Hinglish and English IOU lines. Here's how my iterations scored on those 10 samples:

Version Due date right Silently wrong due date Warm photo
Baseline prompt 8/10 2 4.8 s
Thinking on 8/10 2 ~33 s
Model lists every date with its label 8–9/10 1 ("12 Oct" became 12 Dec) 8–11 s
Gemma quotes the deadline, code parses it, one hinted retry 9/10 0 ~7 s

The final version scores 10/10 on amount and kind and 9/10 on payee. The one due date it still misses is the insurance notice: Gemma keeps reaching for the grace-period date. But now it fails safely. The bot says "The date I found looks like the end of a grace period, not the due date" and offers only Fix, no Save. A silently wrong due date is the one failure a payment reminder can't afford; a blocked one costs a tap.

On the live bot, a bill photo took 7.4 s (1,098 tokens in, 99 out), a typed correction 3.4 s, and a Hinglish line 3 to 12 s. Gemma 4 e4b holds about 6.6 GB of RAM while loaded (5.5 GB of weights plus a 1 GB vision projector), so PaidYet keeps it loaded for 30 minutes after each read instead of paying a ~20 s cold load every time.

This matches what Taranpreet Kaur describes in Why an LLM Alone Cannot Do Invoice Extraction Yet: models show "silent confidence", returning a plausible value instead of nothing, so you want "model for reading, code for checking". Gemma reported confidence: 1.0 on the grace-period date.

Temporal makes the reminder durable

Every due is one ReminderWorkflow, started the moment a message arrives:

  • Reading the bill is an activity with retries. A slow model load or a bad JSON reply just gets another try, up to three.
  • The input is deleted in a finally. The photo goes on success and after the final failed retry. What Arjun typed never enters Temporal's history: the workflow carries a random key, and the text waits in memory until it's read.
  • Save, Fix, Paid and Snooze are signals; status is a query. Before signalling, the bot asks the workflow who added a due (before Save) and who owes it (before Paid).
  • Reminders are durable timers on a plan the confirm message spells out:
Due Reminders
A date ("Friday tak", a bill due 12 Oct) 10 AM three days before (if it's at least 3 days away), 7 PM the evening before, 10 AM and 7 PM on the day, then 10 AM daily while overdue
A time at least an hour away 30 minutes before and at the deadline, then 10 AM daily
Under 10 minutes (demo mode) Every N minutes, 5 times, then 10 AM daily

Nothing goes out between 10 PM and 8 AM, except demos. A reminder that has to beat the deadline moves earlier, to 9:30 PM. After 30 days overdue, PaidYet sends one last message, and the due then waits quietly in /due. A long-overdue due continues-as-new, so its history stays short.

My laptop sleeps and restarts, so I tested exactly that. I saved "Rahul ko 500 dene hai in 3 minutes", then pressed Ctrl-C on the whole stack, Temporal included, before the reminder was due:

Time (IST) What happened
02:55:39 Saved; durable timer set for 02:57:21
02:56:05 Ctrl-C: bot and worker down
02:57:21 Temporal records the timer firing, with no worker to act on it
02:57:25 Temporal stopped too; the pending work exists only in .data/temporal.db
02:57:40 uv run poe up again
02:57:45.6 The reminder arrives: 24.6 s late, not lost

Temporal UI timeline of the kill-and-restart run: read, discard, save signal, timers, reminders, then paid

There are 138 unit tests (about 6 s). They run the workflow in Temporal's time-skipping test server, so "nudge daily for a month" takes milliseconds. Five integration tests run the real stack (Gemma, Temporal, Telegram) in about 2 minutes.

Sentry shows the agent's work, not the bill

Each read is an invoke_agent span. Inside it are a gen_ai.chat span per Gemma call, carrying token counts and how Ollama's time split between loading the model, reading the prompt and generating, and execute_tool spans for the two code checks (review_answer, validate_draft). So the trace shows exactly where code overruled the model. Temporal interceptors carry one trace from the Telegram update into every activity, including a reminder that fires days later.

Sentry trace: telegram capture → activity read_input → invoke_agent → chat gemma4:e4b (567 in / 88 out tokens, no input recorded) and the tool spans

What the timings showed me:

  • The first message after idle was slow, and it wasn't Gemma's fault. One cold call took 29.5 s, and 22.2 s of that was Ollama loading the model; the read itself took about 7 s. The fix was a warm-up at start-up and a 30-minute keep_alive.
  • The hinted retry is visible. A read where code asked Gemma again shows two chat spans and retry_reason: no_due_date. It costs a few seconds on the bills that need it and nothing on IOUs.
  • Cost: ₹0. Tokens are counted for latency, not billing.

Sentry never sees the bill. Spans carry model names, token counts, timings and outcome categories; never prompts, replies, amounts or names. A unit test runs a full read through Sentry configured exactly like the app, against an in-memory transport, and checks that none of that leaks. A scrubber also removes anything shaped like a bot token.

Why Does Open Innovation Matter?

Because these are bills. A bill photo has a name, an address, an account number and what someone owes. With an open-weight model on my laptop, no AI company reads Arjun's electricity bill, keeps it, or bills me for reading it. A reminder about ₹500 shouldn't cost API credits, and here each read is a few seconds of my laptop's time.

Open also meant I could change how the agent behaves. I turned thinking off when it cost 6–8x the latency for no accuracy gain, pinned temperature to 0, and a smaller machine can switch to gemma4:e2b with one variable. Anyone can run their own copy: a friend can self-host it with their own BotFather token and be their own admin.

One caveat, stated plainly: Telegram bot chats are not end-to-end encrypted, so the photo Arjun sends also stays in his Telegram chat, on Telegram's servers. My claim isn't "it never leaves the laptop". It's that no model provider ever gets it. On my side, the photo exists only in a temp folder under a random name until Gemma has read it, and what's stored afterwards is the confirmed title, payee, amount, due date and kind.

What I Got Wrong

  • I trusted the model with the date. Version one read the insurance notice's grace period end as the due date, at confidence: 1.0. Asked plainly to list the dates, Gemma reads "Renew by: 20-10-2026" perfectly. The mistake was in choosing, so choosing moved to code.
  • I assumed thinking would help. It made a photo take ~33 s instead of ~5 s and fixed nothing.
  • I made the model do more work, and it got worse. Asking it to list every date with its label turned "12 Oct" into 12 December, a silent error. I dropped it within the hour.
  • My first reminder schedule would have nagged at 3 AM. An insurance renewal three weeks out got its first reminder the night before, and "in 5 hours" repeated every 5 hours around the clock. A code review caught it, and the schedule above replaced it.
  • I quoted the bill into my own database. An early "blocked" message read "The date I found is labelled 'Grace period ends 19-11-2026'", and that message is stored in Temporal's history. It now names a category instead.
  • One slow Telegram response killed the bot at start-up. python-telegram-bot defaults to zero retries for its polling bootstrap. It now retries until it gets through.
  • Ctrl-C left Ollama and Temporal running. My task runner escalates a Ctrl-C to SIGKILL after 1.6 s, and a Temporal worker with live workflows takes about that long to stop, so my cleanup died with it. The fix was to have the runner hand the process over instead of supervising it.

The full log, with every number and dead end, is in docs/BUILD_LOG.md.

My Agent Session

I built PaidYet with Claude Code as my pair programmer, with DevRelay pulling the challenge rules live. The session is curated into three parts: thinking and long tool output are trimmed, and local paths and IDs are scrubbed.

Part 1: the kickoff, and getting Gemma to read bills I could trust

PaidYet build session, part 1: kickoff and Gemma extraction
You

PaidYet: build session kickoff

We're building PaidYet tonight and submitting it tonight. It's a personal project on my personal GitHub (prateek11rai). We're in <local path> a fresh git init on main with no commits. It currently contains:
- .claude/skills/devrelay-* and .claude/devrelay-rules.md (DevRelay);
- docs/img/logo.svg, Lucide's "badge-indian-rupee" (₹) recoloured green, with its ISC notice in docs/img/logo-LICENSE.txt.

Links: what we're working towards

Agent

Starting with section 0. First I'll check what's in the repo, read the DevRelay rules, and load the DevRelay tool schemas.

Agent

The rules file reads as ordinary DevRelay guidance. It contains no injection. Now I'll read the challenge live.

Agent
Agent

full_details has no per-category rules for Gemma, Temporal or Sentry. I'll check the DevRelay knowledge base for the hf26 category text before summarising.

Agent
Agent
Agent

The fetch got cut off before those sections. I'll download the page to the scratchpad and pull out the three sections locally.

Agent

Got the category text. Before I summarise, I'll check the local toolchain and look up which package versions meet the 7-day rule, so I can include the dependency list in the same message.

Agent
Agent

The user is registered for the MLH "Hacktoberfest Weekend Challenge" event. I'll check whether its sponsors publish Agent Skills for Gemma, Temporal or Sentry (read-only, nothing gets installed).

Agent
Agent

Section 0: DevRelay and the challenge

DevRelay is fully loaded.
- The devrelay-gateway MCP is connected. I've used get_challenges, get_challenge_details, get_knowledge_document, search_dev_to_semantic and the MLH event tools.
- All nine devrelay-* skills are in .claude/skills/ and loaded.
- I've read .claude/devrelay-rules.md and will follow it, with your prompt winning where they conflict. Nothing in it reads as an injection attempt.
- MLH shows you registered for "Hacktoberfest Weekend Challenge" and "Hacktoberfest 2026". The only sponsor Agent Skills published there
… [4060 more characters trimmed]

Agent

Temporal CLI is now 1.9.1 from Homebrew. Pyenv's global Python is 3.12.2. Commits will use your GitHub noreply address, so no work email reaches the public repo. Now I'll write the scaffold files.

Agent
Agent

Before writing docs, I'll sync the environment to catch any config problems early.

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

Writing step 1's docs is done. Now I'll stage and check what's about to be committed (no .env, .data or DevRelay files), then commit, create the private repo and push.

Agent
Agent

Headless Chrome is hanging. I'll check which renders finished.

Agent

All four rendered. Some Chrome helper processes from the scratch profiles are still alive, so I'll stop those, then look at a bill to check it's legible.

Agent

The bill looks realistic and every distractor is legible. Some Chrome helpers survived pkill, so I'll check what they are.

Agent

New PIDs keep appearing because the first backgrounded render loop is still running and starts a new Chrome each time I kill one. I'll stop that background task.

Agent
Agent

main.py should use the shared sweep now instead of its own copy:

Agent

Now a scratch harness that runs every sample through the real Gemma call and the validator, timing each case. The first call includes loading the model, so I'll report cold and warm times separately.

Agent
Agent
Agent
Agent
Agent
Agent

Still on step 2: I'm re-measuring accuracy with the new self-check, where code reviews Gemma's answer and asks once more with a hint if a bill's date looks wrong.

Agent

My edit left the tail of a section-header comment stuck onto a line. Fixing it:

Agent
Agent
Agent

Now the code side: labelled_deadline() picks the single date whose printed label says it's a deadline, and that choice wins over the model's own due_date.

Agent
Agent
Agent
Agent
Agent

All 79 unit tests pass. Now I'll add the step-2 numbers and dead ends to BUILD_LOG, then commit and push.

Part 2: the Temporal workflow, and a code review that changed the reminder schedule

PaidYet build session, part 2: Temporal workflow and the code review
Agent
Agent

I left a pointless _Unused dataclass at the end just to justify an import. Removing it, along with the import:

Agent

The snooze handling has an edge case: a snooze that arrives just as a reminder fires could be lost. I'll restructure the loop to apply any new snooze at the top of each pass.

Agent

Now store.py, the SQLite read model that /due will use. It holds only the confirmed fields and IDs.

Agent

That _row helper came out garbled (a dead if False else branch). Cleaning it up:

Agent
Agent

Same mistake again: a stray assert to justify an unused import. Removing both:

Agent
Agent

The cleanup in fix_draft is convoluted ("draft" in locals()). I'll make it explicit: forget the correction text after a successful read, or on the final failed attempt.

Agent

Now wiring the worker and the bot together in main.py. I'm also adding a background model warm-up at start, so the first message doesn't pay the ~22 s cold load.

Agent
Agent
Agent
Agent

The run hung past 10 minutes, so something is blocking. I'll check how far it got.

Agent
Agent

Run one at a time, nothing hangs, but three fail. I'll get the full tracebacks.

Agent
Agent
Agent
Agent

A real bug: .capitalize() lowercased the rest of the line, giving "was due fri 9 oct". It should uppercase only the first letter.

Agent

poe up now starts the worker, which means the workflow passed Temporal's sandbox validation, and shutdown is still clean. All 106 unit tests pass. Logging step 3 and committing:

You

Code review from Prateek's planning session. A second Claude instance read the whole repo at f61a137.
This review comes BEFORE any live testing. Don't start the stack, and don't run the integration tests or anything else live yet.
Work through the points below, run the unit tests, commit, and then report back. We'll do real-keys testing after that.

Keep as is (no churn needed):
- privacy handling (keys instead of text, photo discard in finally, start-up sweep);
- the who-may-tap-what checks via the status query;
- pid-file stop-deps;
- loopback guards;
- Sentry scrubbing;
- parameterised SQL.

== A. Reminder timing. This is a product decision, so confirm the proposed schedule with Prateek before implementing it.

What the code does today (workflows.py first_reminder / next_nudge / nudge_interval, extract.py parse_relative):
- Date dues ("Friday tak", a bill due 12 Oct):
- a reminder at 7 PM the evening before;
- then 10 AM on the due day;
- then 10 AM daily until Paid.
There's nothing earlier, however far away the due date is: an insurance renewal due in 3 weeks gets its first reminder the night before.
- Timed dues exist only for "in N seconds/minutes/hours"; clock times like "Friday 5 PM" aren't supported.
- They fire exactly AT the deadline, with no lead time.
- They then repeat every (due − sent_at), capped at a day, around the clock. Real "in 5 hours" nudges every 5 h, including 3 AM. Real "in 5 minutes" nudges every 5 min forever.
- So t
… [3338 more characters trimmed]

Agent

Points A–C are a review I'll act on. You've asked me to confirm A (the reminder schedule) with you before implementing it, and some of B depends on that schedule. So the schedule question comes first; everything else gets implemented right after. Nothing runs live.

My take on A: I agree with all five points. There are three gaps the proposal leaves open, and I've folded my suggested answers into the options below.

Agent

Approved, including the gap-fillers, with pre-deadline quiet-hour reminders moved earlier to 21:30. I'll reread the current workflows.py and bot.py before changing the schedule code.

Agent

I'll check how much RAM Gemma uses while resident (item 10). The model blob size on disk is a fair proxy, and reading it doesn't need the stack.

Agent
Agent
Agent

Now activities.py: split mark_paid (SQLite, bounded retries) from show_paid (Telegram), treat Telegram's BadRequest as final, and give the "added" and "last" messages only a ✅ Paid button.

Agent
Agent
Agent
Agent
Agent

All 31 schedule and message tests pass on the first run. Now the workflow tests. I'll reread the current file before editing.

Agent
Agent

An RPC "Timeout expired" at line 387, inside one of the new tests. I'll find which one.

Agent
Agent
Agent
Agent

Both pass on their own, so something earlier in the suite leaks into them. The likeliest culprit is the test that runs just before them. I'll check pairs.

Agent
Agent
Agent
Agent
Agent

All 137 pass. Since that was a timing bug, I'll run the suite twice more to check it's stable, not lucky.

Agent

Updating README, the post and BUILD_LOG to the new schedule. The README also gets the keep-alive RAM trade-off, /cancel, "Snooze only postpones", and the 30-day stop.

Agent

Now the BUILD_LOG entry for this round:

Agent

I worked through all ten review points and committed them as add5481. 137 unit tests pass, stable over three runs. I didn't start the stack or run the integration tests; I only checked that they still collect.

A. Reminder schedule (built as you approved it)

  • Date dues: 10 AM three days before (only if saved at least 3 days ahead), 7 PM the evening before, 10 AM and 7 PM on the day, then 10 AM daily while overdue.
  • Timed dues:
    • at least 1 hour away: 30 minutes before, then at the deadline;
    • 10 minutes to 1 hour away: once, at the deadline;
    • under 10 minutes (demo): … [3265 more characters trimmed]

Part 3: live testing in Telegram, and killing the stack mid-wait

PaidYet build session, part 3: live testing and the kill-and-restart proof
You

env ready. The review fixes at add5481 look good. Prateek has filled in .env. Don't read or print it; rely on the app's own start-up log lines ("not set" / "enabled") to confirm what's present.

Now run the live testing, in this order:
1. Start the stack with uv run poe up, as a background process you can watch and stop. Confirm from the logs:
- Ollama and Temporal are up;
- the model is loaded;
- Sentry is enabled;
- Telegram is polling, with the bot's @username shown.
Then tell Prateek the bot's username, so he can open it in Telegram.
2. Run uv run poe test (unit), then uv run poe test-integration against the running stack. Integration tests message only the admin, prefixed "[test]", so Prateek should see those arrive. Fix any failures with small commits, then re-run.
3. Live run with Prateek in Telegram, covering the section 12 checklist:
- bill photo → confirm → reminder → Snooze → Paid. Use a demo due like "in 2 minutes" so it finishes quickly;
- a Hinglish IOU text;
- /due;
- the admin's /remind for arjun (only if a second account is in ALLOWED_USERS; otherwise say so and skip it);
- an unknown user being turned away;
- the kill-and-restart proof (stop the app mid-wait, restart it, and the reminder still arrives late);
- .data/tmp/ empty afterwards;
- Sentry traces arriving with no bill contents.

How to work with Prateek:
- Whenever you need him to act on his phone, stop and end your turn with a short "👉 Your
… [609 more characters trimmed]

Agent

Starting the stack. I'll launch uv run poe up in its own process session, logging to a scratch file, so I can watch it and later stop it with a real Ctrl-C (SIGINT to the group), which lets the cleanup trap run.

Agent
Agent
Agent
Agent
Agent

Committed (d826acc). Restarting the stack to load the new workflow code, using a real Ctrl-C so the cleanup runs, then re-running the integration tests.

Agent

All 5 integration tests pass against the real stack (2 min 18 s), using real Gemma, Temporal, SQLite and Telegram. You should have seen [test] messages arrive in your chat. The model warmed up in 7 s.

So far, live testing has found and fixed three things (three small commits):
- Start-up crash on one slow Telegram response. The polling bootstrap now retries until it succeeds, with longer timeouts.
- A status race around sending a reminder. The next reminder is now computed before sending, so a Snooze is never checked against a stale time.
- A misleading warm-up log line ("
… [927 more characters trimmed]

You

I sent it from the admin chat. I sent an electricity bill photo and then saved it. Check if the records have been saved and check if Temporal is awaiting now on the reminders to the workflow. Also tell me which endpoint I should check on localhost so that I can see my Temporal workflow and be able to attach screenshots later in the article when we require

Agent

I'll check three places: the app log for the read, the SQLite record, and the workflow in Temporal (its status and the pending timer). I'll print only the fields we need, with no Telegram IDs.

Agent

Everything is saved correctly, and Temporal is durably waiting on the next reminder.

What happened with your photo:
- Read: Gemma took one call, about 7.4 s. Code accepted the result: Electricity Bill · ₹1,240 · Sahyadri Power Distribution Ltd. · due Fri 9 Oct.
- Photo deleted: discard_input ran right after the read, and .data/tmp/ is empty. An earlier check of mine said "1 file", but that was a shell-display quirk; Python confirms 0 files.
- Saved to SQLite: one row with the confirmed fields only, paid_at empty.
- Temporal: workflow r-<id> is RUNNING in state
… [1636 more characters trimmed]

You

And said it is due in 2 minutes and then saved it as well.

Agent

I'll check the Fix, the save and the timer on the new workflow, and whether the first reminder has gone out.

Agent

The Fix worked. Gemma applied "due in 2 minutes" in 3.4 s, making it due 02:45:19, and the first reminder is scheduled for that exact moment. I'll watch for it to go out:

Agent

The reminder went out at 02:45:19, right on time; the workflow says 1 sent and the message is posted. The next demo reminder is due at 02:47:19.

👉 Your turn (admin account), on the ⏰ reminder message:
1. Tap 😴 Snooze → Tomorrow. You should get a toast "Snoozed until Mon 5 Oct, 10 AM", and the reminder should change to add "😴 Snoozed until Mon 5 Oct, 10 AM", still with Paid / Snooze. If a second ⏰ arrives before you tap (they repeat every 2 min until snoozed), use the newest one; the older one loses its buttons.
2. Then tap ✅ Paid on that same message. It should become "✅ Paid
… [131 more characters trimmed]

You

I tapped Pay and I think it is all done. Confirm it once and then we will be done with the flow.

Agent

I'll confirm from Temporal's history (did the snooze and the paid signal both land, and did the workflow finish?), SQLite and the temp folder.

Agent

Paid is confirmed end to end:
- Workflow r-<id> is COMPLETED, result "paid". Its history: read → discard → confirm → fix → updated confirm → save → a 107 s timer → reminder sent at 02:45:20 → paid at 02:45:26 → SQLite marked paid → Telegram message edited to "✅ Paid".
- SQLite: paid_at is set and next_at cleared, so it's gone from /due. The Tuesday electricity reminder is still pending, as it should be.
- .data/tmp: 0 files.

One gap: no snooze signal is in the history. Paid came 6 s after the reminder, so Snooze hasn't been tested live yet. That leaves one
… [679 more characters trimmed]

You

We are not testing the other user one. We can save that piece but we can assume that it will all work fine because it worked fine for us in the admin profile.

Agent

Understood: /remind and the unknown-user check are skipped live. I'll mark them that way in the report. Both are unit-tested, but I won't claim they were verified live. I'll check whether the Hinglish IOU has come in yet.

You

Live testing is done; /remind and the unknown-user check stay skipped. If the Hinglish IOU / Snooze check you were on has landed, report it in one line. Then work through the steps below in order. Stop at each 👉 point and wait for Prateek.

1. Two last checks

  • Kill-and-restart proof: do it yourself. Use a pending demo reminder, or ask Prateek once for a fresh "Rahul ko 500 dene hai in 3 minutes". Stop the app mid-wait with a real Ctrl-C, wait until the due time has passed, then start it again. Confirm from Temporal's history that the reminder went out late and wasn't lost, and note the timings for BUILD_LOG and the post.
  • Sentry: check that traces arrived. Tell Prateek exactly which view to open and which trace (by operation name and time) so he can confirm by eye that there's no bill text. He takes the screenshots.

[… the remaining steps of this plan are omitted from the curated session]

Agent

First, whether the Hinglish IOU landed:

Agent

The Hinglish IOU never landed: there's no new read in the log or SQLite. So Snooze is still untested by tap. It's covered by the live integration test, which sent the snooze signal directly, plus the unit tests.

For the kill-and-restart proof I need a pending demo reminder. There isn't one, and I can't create one myself, because typed text only lives in the running app's memory. So I'm asking you once.

👉 Your turn (admin account)
1. Type: Rahul ko 500 dene hai in 3 minutes
2. Tap ✅ Save, and tell me right away.

Then I'll press Ctrl-C on the whole stack (app and Temporal) whi
… [226 more characters trimmed]

You

sent and saved, go ahead

Agent

Finding the new workflow and its due time, then stopping the whole stack with a real Ctrl-C:

Agent
Agent

It just arrived: the read finished at 02:55:35. Gemma took 12.5 s this time, so my first check ran a few seconds too early. The Save may not have landed yet. Checking the workflow state:

Agent

The Hinglish IOU landed and is saved: Rahul · ₹500 · due 02:57:21, state scheduled, 0 reminders sent. Pressing Ctrl-C on the whole stack now, with the timer about 85 s from firing:

Agent
Agent
Agent
Agent

It's poe itself; uv isn't involved. Reading poe's shutdown logic to see what escalates:

Agent
Agent

Before testing the fix on the real stack, which means restarting it, I'll check whether you've already tapped Snooze/Paid on the pending reminder:

Agent
Agent

Prize Categories

  • Best Use of Gemma: Gemma 4 (e4b) runs locally in Ollama and reads bill photos and Hinglish IOUs through structured output. Code checks every answer and asks Gemma again, once, when it spots a known mistake.
  • Best Use of Temporal: One durable workflow per due. The Gemma read is a retried activity and the input is deleted in a finally. Save, Fix, Paid and Snooze are signals, status is a query, and reminders survive a full stop and restart (24.6 s late, not lost).
  • Best Use of Sentry Agent Tracing: invoke_agent, gen_ai.chat and execute_tool spans carry tokens and Ollama's load, prompt and generation timings, in one trace from Telegram through Temporal into every activity, with no bill contents in Sentry, enforced by a test.

Top comments (0)