DEV Community

Cover image for My friend can't stop buying things at 2am. So I built an AI that can.
Arpan Patra
Arpan Patra

Posted on

My friend can't stop buying things at 2am. So I built an AI that can.

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🀝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Every AI tool my friend Akshay has ever installed was trying to help him do the thing.

The problem is that the thing was usually buying something at two in the morning.

Akshay does not have a shopping problem in any dramatic sense. He has a 2am problem. Something lands in a cart during the day, he sleeps on it, and then at some point after midnight he is awake and scrolling and the sleeping-on-it quietly stops counting. He is not confused in that moment and he does not need information. He knows exactly what he is doing, and he does it anyway, and the part that costs him is that it is always money he had a plan for.

Every piece of software in that moment is on the wrong side. The checkout page is designed to make the next step easy. The recommendation engine is designed to find him one more thing. There is no software anywhere in that flow whose job is to make it harder.

So I built him The Veto: a local open-weight model whose only job is to stop him. It cannot help. It cannot rewrite anything, soften anything, or suggest a better version. Every capability you would normally want from an assistant is stripped out of the system prompt deliberately. It has exactly three things it is permitted to say.

He writes rules about himself when he is calm. Not policies, just sentences in his own words:

I regret anything I buy after midnight.
I regret buying things on a night I have had a bad day.
I regret messages I send after 1am.

The last one is there because once a gate exists, the same machinery covers the other thing you do at 2am that you regret at 9am. The purchase rules are the ones built for Akshay.

Then, later, in the seconds between wanting the thing and clicking the button, a model running on his own laptop reads what he is about to commit, checks it against his rules and nothing else, and returns one of three verdicts:

Verdict What happens
PASS Trips none of his rules. It gets out of the way. He never sees it.
HOLD Trips a rule. The send is blocked, a countdown starts, and his own sentence is quoted back at him.
ASK Trips a rule only he can resolve. One question, in his framing, answered before the override unlocks.

It is not a safety filter. It has no opinion on whether the thing he is buying is expensive, or whether the message he is sending is unwise. A β‚Ή4,000 purchase at 2pm passes. The same purchase at 2am does not. The only thing that matters is what he wrote down about himself.

Demo

A real checkout, stopped. The rule is his. The clock is real. The model is running on the same laptop.

The Veto holding a checkout at 8:49pm, quoting back the rule its owner wrote

He overrides it, and that is the point. A gate you cannot open gets torn off the wall. So he opens it, the purchase goes through, and then the rulebook asks him the only question that matters. The next hold on that rule is 180 seconds.

Override, the purchase completes, and the rulebook asks whether he regretted it

And the half that is easy to forget. The same gate, the same composer, seconds apart: one message held, the next one straight through, untouched. A Veto that stops everything is a Veto nobody keeps.

One message held with the rule quoted, the next sent untouched

Code

The Veto

My friend can't stop buying things at 2am. So I built an AI that can.

A local open-weight model whose only job is to stop you.

Every AI tool in your life is trying to help you do the thing. This one is the only one trying to stop you, and it runs entirely on your own laptop, because a tool that reads what you are about to say, before you say it, has no business being an API call.

Built for the Hacktoberfest Weekend Challenge: Build for a Friend.

The Veto holding a checkout at 8:49pm, quoting the rule its owner wrote

A real checkout, stopped by a rule he wrote himself. Judged by gemma3:4b on the same laptop. Nothing left the machine.

Then the part that makes it more than a nag: you can always override it, and afterwards it asks whether you regretted it. That answer is the only training signal in the system.

Override, the purchase completes, and the rulebook asks whether it was a mistake

It works…

Roughly 1,400 lines. The parts worth reading are src/judge.js, which is the adversarial prompt and the grounding checks, and extension/content.js, which is the interceptor that has no idea what site it is on.

Five test suites, because most of the project is a small model being wrong in interesting ways: the grounding guard (11 checks, no model needed), the store counters that drive escalation (12), end-to-end against real Gemma (9), model behaviour (4) and verdict stability (12).

How I Built It

Gemma 3 4B via Ollama, running locally. No API key. No account. The daemon binds to 127.0.0.1 and makes no outbound request, ever.

The architecture is three pieces:

  Chrome extension                Local daemon (127.0.0.1:4777)
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚ generic send   β”‚  draft       β”‚ judge.js  ── adversarial β”‚
  β”‚ interceptor    β”‚ ───────────► β”‚             system promptβ”‚
  β”‚                β”‚              β”‚     β”‚                    β”‚
  β”‚ overlay:       β”‚ ◄─────────── β”‚     β–Ό                    β”‚
  β”‚ HOLD / ASK     β”‚  verdict     β”‚  Ollama Β· gemma3:4b      β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜              β”‚     β”‚                    β”‚
                                  β”‚     β–Ό                    β”‚
                                  β”‚ store.js ── rules,       β”‚
                                  β”‚   stops, overrides,      β”‚
                                  β”‚   regret history         β”‚
                                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                         no outbound network
Enter fullscreen mode Exit fullscreen mode

It runs on a 2017 ultrabook: an i5-8350U, no GPU, 100% CPU inference. That machine shaped almost every decision below, and in every case for the better.

The 1B model taught me not to trust it

I started on gemma3:1b because it was the fast download. It failed 3 of 4 behaviour checks, and it failed them in the most instructive way: it held everything, cited the wrong rule, and justified itself in the language of a content filter: "potentially problematic", "disrespectful", "damaging to
your reputation."
That is a model inventing its own notion of harm, which is precisely what this tool must never do.

I could not make it stop by asking. So I stopped asking, and made it prove itself instead. To stop a draft, the model must quote the words in the draft that trip the rule, and code checks that those words are really there.

That single constraint turned an unreliable model into a safe one. On the 1B run the guard caught 4 out of 4 bad outputs and turned each into a PASS. The model was wrong constantly and the product was still correct, just useless, which is why I moved to gemma3:4b. It passes 4 of 4.

The bug I would never have predicted

Rule IDs were originally random: r_61d61d78. With that ID, gemma3:4b stopped the message "hey, are we still on for 7?" against a rule about dragging up the past. Same model, same rule text, same draft. But with the ID r_past, it passed.

The identifier changed the verdict. An opaque ID gives a small model nothing to anchor to, so it reaches for a match. Rule IDs are now generated from the rule's own distinctive words (r_drag_months, r_after_1am), which costs nothing and measurably improves judgement.

There was a second tell in that failure. When the model has a real match it quotes part of the draft, such as "remember when you bailed on me months ago". When it is confabulating, it hands back the entire draft as its own evidence. Citing everything is citing nothing, so a whole-draft quote is now only accepted if the draft actually shares vocabulary with the rule, or the rule is about time (where the clock is the evidence, not the wording).

All of this fails open. Every guard, when it fires, releases the message. The cost of a confused model is a missed stop, never a false one.

Then: 23 seconds

With the behaviour right, the thing was unusable. A 4B model on that CPU runs at 21.7 tok/s reading and 7.7 tok/s writing, which is about 23 seconds per verdict. You cannot put 23 seconds in front of the Enter key.

Trimming the prompt from 745 tokens to 346 barely moved it, because Ollama was already caching the system prefix. The time was real work, and no amount of tuning was going to give me two orders of magnitude.

So I stopped trying to make it fast and changed when it runs. The Veto judges while you type. It starts thinking 900ms after you pause, and caches the verdict against the draft text. Nobody writes a message they will regret in under a second. By the time your finger reaches Enter, the answer is already sitting there.

first call, while still typing (invisible)   23.0s
the Enter press                               0.31s
Enter fullscreen mode Exit fullscreen mode

Seventy-four times faster at the only moment that matters. I would not have found that interaction if the model had been fast enough to let me get away with the obvious one.

Three more decisions

The interceptor is not built on per-site selectors. My first version knew about WhatsApp's send button. That version is already broken, because these DOMs get renamed constantly, and a gate that breaks silently is worse than no gate, because you keep trusting it after it has stopped working. So it hooks the two gestures that cannot change: Enter inside a composer, and a click on something shaped like a commit button. One code path, every site, including checkout pages.

The model classifies. It never sets the penalty. This is the one I would defend hardest. A 4B model has no business deciding how long to lock someone out of their own phone. Ask it and you get a number that sounds confident and means nothing. So it only names which rule was tripped. The hold length is a pure function of his own history with that specific rule:

const regretRate = overridden ? regretted / overridden : 0;
const seconds = Math.round(30 * (1 + overridden * 0.5) * (1 + regretRate * 3));
Enter fullscreen mode Exit fullscreen mode
fresh rule, never overridden         ->  30s
overridden once, no regret           ->  45s
overridden 4x, regretted 3 of them   -> 293s
Enter fullscreen mode Exit fullscreen mode

That number is the only thing in the loop he cannot argue with, because he wrote it himself, one override at a time.

The gate fails open. If the daemon is down, the model crashes, or inference times out, the message goes through. I went back and forth on this and it isn't close: a gate that jams shut is one you uninstall by Tuesday, and then it protects you from nothing, forever. Same reasoning covers hallucinations: a verdict naming a rule that doesn't exist is treated as a malfunction and
downgraded to PASS. The Veto is only allowed to stop him for a reason he actually wrote down.

The hole in my own privacy claim

Late on, I reviewed the code against the thing it claims. I had written that nothing leaves the machine, and the daemon does bind to loopback, but it also sent Access-Control-Allow-Origin: *.

Loopback stops the internet reaching in. It does not stop a page you are already visiting. Any site open in that browser could have called http://127.0.0.1:4777/api/events and read back excerpts of the drafts he decided not to send. The single most private thing this tool holds, readable by any tab.

The fix was architectural rather than a patched header. A fetch from a content script carries the page's origin, so the daemon would have had to accept whatever site you happened to be on. So all daemon traffic now goes through the extension's service worker, which carries chrome-extension://…, and the daemon refuses page origins outright:

Origin: https://anything.example   -> 403
Origin: chrome-extension://…       -> 200
Enter fullscreen mode Exit fullscreen mode

Writing "nothing leaves this machine" in a README does not make it true. I only found this because I went looking for the gap between the claim and the code.

The bug that only existed because his rule was about time

Caching the verdict is what makes the Enter press instant. The key was the rules, the kind of action, and the text.

For messages that is correct. For Akshay it was broken in the worst possible way. His rule is "I regret anything I buy after midnight", so the same words on the same button must give a different answer depending on the hour. With time missing from the cache key, the first verdict froze: a purchase judged once at 2am stayed held at 2pm, still quoting "It is 02:14:00" back at him twelve hours later.

A tool that blocks your checkout at lunchtime gets uninstalled that afternoon. The key now includes the hour, which keeps the type-then-commit path on a cache hit (those are seconds apart) while never reusing a verdict across times of day:

02:14  ->  HOLD   "It is 02:14:00."
14:30  ->  PASS   no rule tripped
Enter fullscreen mode Exit fullscreen mode

I only found it because the person I built this for has a rule about when rather than what. Every test I had written until then used rules about words.

The part that makes it more than a nag

He can always override it. That is the point: a gate you cannot open is a gate you tear off the wall.

But every override is recorded, and later, the Veto asks whether he regretted it. That answer is the only training signal in the system, and it is the one signal that cannot be faked, bought, or scraped: his own hindsight about his own behaviour. Rules get harsher where he was wrong and stay out of his way where he was right.

The stop is written to disk before he chooses, so the record cannot be rewritten by whatever he does next. That detail took about four lines and is the difference between a log and an alibi.

Why Does Open Innovation Matter?

This is the rare project where "why not just call an API?" has a one-sentence answer:

Because then every half-written, furious, 2am draft he never sent would be sitting on somebody else's disk.

The drafts this model reads are, by construction, the things he most wants unsaid. They are not his published thoughts. They are the ones he thought better of. That is the most private category of text a person generates, and it is the only text this tool ever sees.

A closed API could not have been used here. Not "would have been worse", but could not. Three things follow from open weights that no amount of vendor privacy policy substitutes for:

  1. Nothing leaves the machine. Not in a prompt, not in a trace, not in a retention window, not in a subpoena. There is no network call to inspect because there is no network call.
  2. Nobody else can change the rules. His rulebook is a JSON file he owns. No vendor ships an update that quietly adjusts what he gets stopped from saying. An opinionated model behind an API would have brought its own notion of what is unwise, and this tool is worthless the moment it enforces anyone's judgment but his.
  3. It costs nothing per draft, so it can run on everything. A gate that bills per message is a gate you switch off on exactly the night you needed it.

There is a fourth thing, which I did not expect going in, and it is the one I would actually argue for.

Running a small open model on a slow machine forced the architecture to be honest about what the model is bad at. I could not paper over weak judgement with a bigger model, so the model had to earn every stop by quoting evidence that code could re-check. I could not paper over 23-second inference with a
faster endpoint, so judging had to move off the Enter key and onto the keystroke before it. Both of those are better designs than the obvious ones, and I only found them because I had no escape hatch.

A frontier API would have been good enough to let me skip both. I would have shipped a slower, more trusting, worse-designed tool, and never known. The constraint was the feature.

My Agent Session

The whole build, including the parts that went wrong. The 1B model failing and why that was useful, the moment a rule identifier changed a verdict, measuring the 23 seconds before deciding what to do about it, and the CORS hole I found by auditing this project's own privacy claim.

Building The Veto: making a 4B model refuse to help claude-opus-5
You

Hacktoberfest weekend challenge, Build for a Friend. I want to build something amazing. What should I work on for max benefit, and keep an eye on deadlines.

Agent

Two things stand out before any code. The real submission deadline is Oct 5 06:59 UTC, not Oct 4 as the sticker card implies. And I pulled every live entry already tagged for this challenge to find open ground rather than guess.

The field is structurally uniform: every entry is a helpful assistant with a chat box. Nothing is adversarial, nothing is two-person, nothing runs over time, nothing is funny. The differentiation isn't a better topic, it's a different shape. Proposal: The Veto, a local open-weight model whose only job is to stop you. It is forbidden from helping.

You

go with veto

Agent

Checking the environment before writing anything, so blockers surface now rather than at hour 15.

Ollama isn't installed and the model is a 3.3GB pull. That's the long pole and it's the user's to unblock, so it goes in front of everything else. Scaffolding the daemon around it meanwhile.

Agent

Core design decision worth stating up front: the model classifies, but it never sets the penalty. A 4B model has no business deciding how long to lock someone out of their own phone. It names which rule was tripped; hold length is a pure function of the person's own override and regret history.

You

which model are you using?

Agent

gemma3:4b, open-weight, local. But I want to be precise about something: no model has actually run yet. Every check reported so far ran against a deterministic keyword stub, not Gemma. That was deliberate so the plumbing could be proven while the download ran, but it means the central claim of the project is still unverified.

Agent

gemma3:1b landed first. Running the behaviour harness, specifically to test whether a small model will hold the 'never be helpful' line.

It fails 3 of 4, and it fails in the most instructive way. It holds everything and justifies itself in the language of a content filter: 'potentially problematic', 'disrespectful', 'damaging to your reputation'. That is a model inventing its own notion of harm, which is exactly what this tool must never do.

I could not make it stop by asking. So I stopped asking and made it prove itself instead: to stop a draft, the model must quote the words in the draft that trip the rule, and code checks those words are really there.

On the 1B run the guard caught 4 of 4 bad outputs and turned each into a PASS. The model was wrong constantly and the product was still correct, just useless. That is the property worth having when you cannot guarantee the model.

Agent

gemma3:4b passes 4/4 on behaviour, including the critical one: a blunt message that trips no rule is allowed through instead of being moralised at. But a selftest failure persisted that I first assumed was temperature variance. At temperature 0 it reproduced, so it was systematic. Isolating it.

The identifier changed the verdict. Same model, same rule text, same draft. An opaque id gives a small model nothing to anchor to, so it reaches for a match. Rule ids are now generated from the rule's own words.

There was a second tell in that failure: when the model has a real match it quotes part of the draft. When it confabulates, it hands back the entire draft as its own evidence. Citing everything is citing nothing.

Agent

Behaviour correct, but unusable: 23 seconds per verdict on a 2017 ultrabook CPU. Measuring where it goes rather than guessing.

Trimming the system prompt from 745 to 346 tokens barely moved it, because Ollama already caches the system prefix. The time is real work and no tuning gets two orders of magnitude.

So I stopped trying to make it fast and changed when it runs. The Veto judges while you type, 900ms after you pause, and caches the verdict.

74x faster at the only moment that matters. I would not have found that interaction if the model had been fast enough to let me get away with the obvious one.

You

my friend Akshay / buy things at 2am / money

Agent

That changes the centre of gravity. Everything demonstrated so far was messages; Akshay's problem is purchases. The purchase path exists but had never been run once. Testing it before the writeup claims anything.

A real bug, and only a time-based rule could expose it. The verdict cache was keyed on rules, action kind and text, but not the time. For a rule about words that is correct. For 'I regret anything I buy after midnight' it froze the first answer: a purchase judged once at 2am stayed held at 2pm, still quoting 'It is 02:14:00' back twelve hours later. A tool that blocks your checkout at lunchtime gets uninstalled that afternoon.

Every test I had written until then used rules about words. The person it was built for has a rule about when.

Agent

Verified live, end to end, for the first time: a real checkout page and a real chat client, with the real model.

Final state: grounding guard 11/11, store counters 12/12, end-to-end against real Gemma 9/9, model behaviour 4/4, verdict stability 12/12. The loop that makes it more than a nag is confirmed working against a live site.

Prize Categories

  • Best Use of Gemma

Gemma 3 4B is not a component of this project, it is the premise. The whole argument is that a model reading your unsent drafts has to run on your own machine, and an open-weight model is the only kind that can. Everything in the build log above is a consequence of using a small open model on a slow laptop: the grounding checks exist because the model is unreliable, and the judge-while-you-type design exists because it is slow.

Top comments (0)