DEV Community

ANIRUDDHA  ADAK
ANIRUDDHA ADAK Subscriber

Posted on

Griha: the chore ledger where the fairness argument is auditable, and TabPFN runs in your tab

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

Entry for the Hacktoberfest Weekend Challenge — Build for a Friend (#hf26challenge)
Submitted in the Best Use of TabPFN (Prior Labs) category.

The person this is for

Someone in a shared home who has quietly become the household's project manager.

Not the manager by title — nobody appointed them. They are the one who remembers the geyser vent needs cleaning, who notices the milk is out, who ends up asking "did you take the recycling out?" in a tone that makes everyone hate them. Then, when the work actually gets done, nobody can say whether it was fair. The argument is always the same argument, and it is always unwinnable, because the evidence lives in people's heads and it is different in each one.

Live: https://griha.vercel.app
Code: https://github.com/aniruddhaadak80/griha


Watch live in action :)

Griha is a chore ledger for that person. Three things, and only three things, it does:

  1. Claim a chore without asking twice. It is on the board, you take it, it is yours. Nobody has to assign it.
  2. See the arithmetic, not a verdict. Not "Ishita contributed 41%" — the actual completed minutes, her capacity weighting, how many chores of each category she has, how many are overdue, and what the weekend and public holidays did to the numbers.
  3. Export something checkable. An ICS rota, a CSV, a Markdown summary, or the raw JSON including the audit chain — so the export can be verified rather than trusted.

Why open matters here

The obvious feature of this app is the "who is likely to flake" prediction. That is precisely the feature I most wanted to keep off someone's server.

It runs TabPFN v2 — Prior Labs' open weights — through the WebTabPFN runtime, inside the tab. Not a hosted endpoint. The features never leave the browser, the completions never leave the browser, and the probabilities never leave the browser. The weights are fetched once and cached.

That matters more than it sounds. A closed API for this feature would mean shipping a record of who in a family is failing to take out the recycling to a server neither they nor I control, where it becomes a training input or a log line or an incident. That is a genuinely different product, and the honest answer is that I would not have built this one.

Concretely, where the open approach beat a closed one:

  • Cost. It costs nothing to run. There is no per-request metering on a chore ledger, which is the kind of thing that quietly decides whether a small household tool is viable at all.
  • It works offline. Once the weights are cached, the panel runs with the network off. A chore app that stops working because a vendor had an incident is not a chore app.
  • The failure modes are inspectable. When it breaks, it tells me the backend, the precision, and the actual error string. I could read WebTabPFN's source and find out why.

The part I'm happiest about

The fairness number is computed by one pure function with an injected clock, computeFairness, versioned as griha-fairness/2026.10.1. Every term that moves the score is shown to the user with its own contribution.

And every mutation is chained:

seal_n = SHA-384( UTF-8(seal_n-1) || canonicalJson(event_n) )
Enter fullscreen mode Exit fullscreen mode

Replay it and you get the same seals or you get told exactly where the chain breaks. The ordering follows a stored sequence number, never a timestamp.

This started as an integrity feature and turned into the feature I care about most. When a fairness score is the thing two people are arguing about, "trust me" is a bad foundation. Being able to hand over the JSON and say check it yourself is a different social object entirely.

Two bugs the browser found that reading the code would not have

I want to be specific about these, because both were invisible to me until something ran.

The model would not start on most machines. WebTabPFN exposes hasWebGpu(), and I used it to pick a backend. It is a feature check. It happily returns true in headless Chromium where there is no usable GPU adapter at all, and then load() fails with Failed to get GPU adapter. WebTabPFN does not fall back on its own. So the panel — the headline feature — was dead in headless Chromium, Safari and Firefox. It now attempts webgpu/int4 and then falls back to wasm/int8 explicitly, and if neither starts it shows you both failures rather than one.

The demo board produced a model that could not exist. Once it started, it said classification requires at least two classes. That was not a model problem. My seeded chores had no assignee, and a negative training row requires a chore that was assigned and then passed its due date unfinished — so every row in the table had the same label. A classifier is correct to refuse that.

The fix was not to loosen the check. It was to make the seed honest: the first-run household now contains two genuine lapses, assigned, overdue, never done. Any real four-week household history contains chores that got missed. A spotless seed record would have looked tidier and made the ML panel inoperative on the first screen anyone saw.

The panel also states in plain words when a table has one class, instead of surfacing the model's opaque error to a user who has never heard of a classifier.

What I verified, and how

I would rather show the receipts than assert it works.

  • 109 unit and integration tests. The persistence tests run against PGlite — actual Postgres compiled to WASM, not a mock, because the bugs that layer is prone to (a reserved word, a LIMIT/OFFSET type mismatch, a JSONB cast) only surface against a real engine.
  • 91 live checks against the deployed instance, including that /api/health reports neon-postgres with a successful SELECT 1.
  • 26 browser journeys, 13 on desktop and 13 on a Pixel 7 viewport, against production — including a real mutating MCP call read back through REST, and a share link that works with no session at all.
  • TabPFN actually fitting in the browser and returning live probabilities on 20 real labelled rows. It reports UNCERTAIN at 0.41–0.45 rather than dressing a near-coin-flip up as a verdict. Twenty rows is a very small table for a prior-fitted network and the UI says so.
  • Green CI on both jobs.

Two of those gates earned their keep. The live verifier caught a real audit-ordering bug in production that I had missed by reading. And CI failed on every commit including the first, because Next 16 generates PageProps and RouteContext into .next/types — so tsc could never pass on a clean checkout. Local typecheck had been passing only because a previous build left those files behind. npm run typecheck now runs next typegen first, and I verified it by deleting .next and typechecking cold.

Honest limitations

  • This is a PWA, not a native app. It installs to a home screen and runs offline, but I did not ship an APK or an IPA, and there is no store listing.
  • Cross-device sync is not solved. Your household lives in one browser profile. I used a per-session household precisely so I never had to build accounts, which also means there is no "log in on your phone and see the same board" — yet. There is a read-only share link, which is a real but partial answer.
  • The predictions are weak, and that is partly the point. With tens of rows they hover near 0.5. The deterministic engine is the part I would trust; the ML panel is an honest second opinion that says when it does not know.
  • Weather and holidays degrade gracefully, not magically. Live Open-Meteo and Nager.Date with sealed dated fallbacks, and the UI labels which one you are looking at.

Where it goes next

Households are per-session because accounts were the wrong thing to build first. The chain is already household-scoped, so shared boards and real multi-device sync are an additive change rather than a rewrite. The audit chain could also stand alone as a general provenance primitive — it does not care that the thing being sealed is a chore.

A note on the "build for a friend" part: the household in the demo (Aarav, Ishita, Nani) is seeded sample data, not real people, and I would rather say that plainly than invent a testimonial. The problem above is a real recurring one, but if you have actually handed this to someone and know what they said, that belongs here — and it belongs here in their words, not mine.


Built with PriorLabs-TabPFN, as the model licence requires. Next.js 16, Neon Postgres, and a lot of arguing about fairness arithmetic.

Top comments (0)