AI was going to increase our productivity. It is also apparently going to replaces the human race before I turn 24, but let's stick to the productivity part, because that part came true.
We genuinely can do ten hours of work in four now. I've watched it happen. An agent chews through a refactor that would have eaten my afternoon, and it's done before my coffee goes cold.
So here's the question nobody seems to have asked: what happened to the other six hours?
They did not become a walk. They did not become sunlight, or a conversation with a person, or a nap. They became more work. The bar moved. Ten hours of output in four hours became the new baseline, and then we filled the remaining six with more of it. We took a productivity windfall and spent all of it on more productivity.
And the cruel detail: with an agent, you're sitting more still than before. You're not even typing. You're watching a model work, waiting for a diff, reviewing output. Thirty minutes can pass where the most athletic thing you did was scroll.
I built BreakTension because I was doing exactly this and my own break reminders were useless.
๐ github.com/DnyaneshU/BreakTension
The problem with every break reminder I've installed
They're dismissible. A notification slides in, I click it away without reading it, and I keep sitting. I knew I should get up. The reminder was never the missing piece.
What I lacked was a reason I couldn't dismiss in a quarter of a second.
So BreakTension removes the dismiss button. After 35 minutes of active screen time, every browser tab goes dark. No "later," no X in the corner. There's a task โ touch a leaf and show two fingers โ and a QR code.
You scan it, you walk outside, you take the photo. Gemma, running locally on your own laptop, decides whether you actually went outside or just leaned toward the window. If it passes, your browser comes back.
And because the agent-driven sitting is the worst kind, there's a VS Code extension too, sharing the same timer. Switching from Chrome to your editor is not a break.
How it works
โโ Chrome Extension (MV3) โโโโโโโโโโโ
โ lock overlay + QR โ
โโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโ โโ VS Code Extension โโ
โ Node verifier server โโโโโโโค modal reminder โ
โ owns ONE shared 35-min budget โ โโโโโโโโโโโโโโโโโโโโโโโ
โโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโ
โ localhost:11434
โโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Ollama ยท gemma4:e4b (local) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โฒ
โ https://<your-lan-ip>:8787/capture?t=<token>
โโโโโโโดโโโโโโโโโโโ
โ Phone browser โ live camera only, no file picker
โโโโโโโโโโโโโโโโโโ
One laptop, one phone, same Wi-Fi. No datacenter, no account, no API key, no internet connection at any point.
The trust boundary matters: the extension never unlocks because the phone asked it to. The phone can only submit a photo against a single-use token the server minted. Only the server โ which called Gemma โ can emit the unlock. A forged message from a web page can't skip the capture step.
Why open models aren't a cost optimization here
This isn't a project where an open model was the cheaper option. It's one where a closed model would have been the wrong thing to build.
The data is the most invasive imaginable. To verify a break, the system needs photographs of your hands, your street, your garden โ taken at intervals that reveal precisely when you're away from your desk. Uploading that continuously, all day, to someone else's server is indefensible. Running Gemma locally isn't a privacy feature bolted on afterward. It's the only version of this idea I was willing to run on my own machine.
It works where "go outside" actually takes you. Basements. Trails. The dead patch behind my building. A cloud round-trip fails exactly where the product is supposed to succeed.
Zero marginal cost bought a better design. At ~15 breaks a day, a paid vision API would make this an expensive habit โ and I'd have built a cheaper, weaker check to compensate. Because inference is free, I could afford to ask the model four separate questions per photo instead of one.
That turned out to be the difference between catching the central cheat and missing it entirely.
The finding that changed everything
I assumed the hard part would be recognising "outdoors." It wasn't. The hard part was a prompt engineering failure I only caught because I tested the attack instead of the happy path.
The attack: lean over, photograph the garden through the window, never leave the chair. If that works, BreakTension is just a slow dismissible reminder with extra steps.
My first prompt asked Gemma all six questions at once โ is it outdoors, is it through glass, is it a screen, was the task done, how confident are you, and why. It read beautifully.
It detected the window photo 0 times out of 5.
Asking the same model the same question alone, in one short sentence:
Where is the camera standing: outdoors in the open air, or indoors
looking out through a window? Look for a window frame, glazing bars,
or reflections on glass.
3 out of 3 detections. 0 false positives on genuine outdoor photos.
The model could always see the window frame. My long prompt was drowning the signal.
| long combined prompt | four short prompts | |
|---|---|---|
| Window photo detected | 0 / 5 | 3 / 3 |
| False positives on real outdoor photos | โ | 0 / 3 |
| Time per judgement | 45s+ (timed out) | 3โ24s |
The verifier now asks four short questions in sequence โ outdoors? through glass? a screen? task done? โ short-circuiting on the first failure. Each prompt is kept under 420 characters, because the thing that broke it was length.
A second measurement in the same vein: gemma4:e4b has a thinking mode, and leaving it on cost ~24 seconds per judgement versus ~3 seconds off, with no accuracy gain. When someone is standing in a car park waiting to be let back into their browser, that's not a tuning detail.
Then I tested it on real photographs
Everything above was measured on synthetic fixtures โ gradients generated by a Python script. Those prove the model discriminates. They don't prove it copes with the noise and lighting of a real camera.
So I ran real photos through it, including the two that matter:
| Photo | Expected | Result | Rejected at | Time |
|---|---|---|---|---|
| Hikers on rocks, backlit | pass | โ all gates | task only | 12.2s |
| Field and trees | pass | โ all gates | task only | 13.2s |
| Reading under a tree | pass | โ all gates | task only | 13.0s |
| Interior window, trees beyond | reject | โ caught | glass |
28.8s |
| Desk, chair, lamp | reject | โ caught | outdoor |
3.4s |
Zero false rejections. Zero false acceptances.
The window result is worth reading closely. Gemma said the scene was outdoors โ and then caught that the camera wasn't:
"The image clearly shows a window frame and glass panes. Go all the way outside."
Separating "the view is outside" from "the photographer is outside" is the entire product.
And the short-circuit pays for itself: the indoor photo rejects in 3.4 seconds on one model call, the window photo in two, while only a genuinely outdoor photo costs all four. The most common cheat is the cheapest to catch.
The bugs I only found by actually running it
The code was finished and verified before I loaded the extension in Chrome for the first time. Four things broke immediately:
-
The extension reported "verifier not running" while the server sat there answering requests. An MV3 service worker can't click through a self-signed certificate warning the way a tab can โ its
fetchjust fails, silently. - It could sit "locked" with no overlay on screen and a dead button.
-
F5 in VS Code did nothing โ I'd never written a
launch.json. - The extension host timed out waiting for a debugger that never attached.
Every one is an integration or packaging failure. Verifying the code was correct told me nothing about whether the thing was runnable. Those are different questions, and I'd only been asking the first one.
The cert fix is the interesting one. The certificate exists so the phone can open its camera โ getUserMedia refuses a non-secure origin over the LAN. But that cert is self-signed, and local clients choked on it. The server now also listens on plain HTTP bound to 127.0.0.1, because localhost is a secure context by specification. Phone gets TLS, extensions get loopback, nothing is exposed to the network.
The escape hatches, and why they're limited
A tool that locks your browser at the wrong moment gets uninstalled, and an uninstalled tool prevents exactly zero minutes of sitting. So there has to be a way out. But every way out is also a way to cheat:
| Hatch | Limit | Why |
|---|---|---|
| Snooze | 5 min, once per break | Long enough to finish a thought, too short to be an escape |
| Emergency skip | 2 per day | A bad afternoon shouldn't become a bad week |
| In a meeting | 1 hour, self-expiring | See below |
| Quiet hours | configurable | 3am must never demand you go outside |
Skipping asks twice. The second prompt names what it costs: "This spends one of your 2 skips today, and your streak resets." A one-click skip would quietly undo the only claim this product makes.
The spec originally called for detecting video calls from "active mic/camera state, not hostname sniffing." That turned out to be impossible โ no Manifest V3 API exposes a page holding a microphone stream. So it's a manual toggle, which is honest and has no false positives. Its obvious weakness is that you flip it, forget, and the product goes quiet forever. That's answered with a one-hour expiry rather than left standing.
What it deliberately doesn't do
It never touches your work. The VS Code extension shows a modal and nothing else โ no blocked saves, no intercepted edits, no closed editors. An editor holds unsaved thought, and a reminder that risks losing it is worse than no reminder.
The browser overlay lives in a shadow root and never touches the page's DOM, so form inputs, scroll position and unsaved text all survive a lock.
Honest limitations
- This is friction, not attestation. You can stand on a balcony. The goal was never unbeatable verification โ it was making cheating cost more than complying. One click became: pick up your phone, scan, walk to a door, perform a random gesture, pass four checks.
- Gemma is a 6.6 GB open-weight model, not an oracle. It will sometimes reject a legitimate photo. That's why failures always explain themselves and retakes are unlimited.
- Test photos were stock images, not ones I took. Said plainly because the distinction matters.
- The points system is trivially cheatable. It's local, there's no server, you could edit it in devtools. It exists so skipping costs something visible โ not as a deterrent.
- Peer review is cut. The original design had strangers approving each other's photos. It needs a backend, identity, photo storage, and other users who exist. It also contradicts the only argument that makes this worth building. A private streak does the same job.
Try it
Prerequisites: Ollama, Node 20+, a phone on the same Wi-Fi.
1. Server (required โ this is the brain)
git clone https://github.com/DnyaneshU/BreakTension.git
cd BreakTension
ollama pull gemma4:e4b # 6.6 GB, CPU is fine
npm run server # zero dependencies to install
You'll see:
BreakTension server
โโ extension : http://localhost:8788 (loopback only)
โโ laptop : https://localhost:8787
โโ phone : https://192.168.1.9:8787
Leave it running. Both clients talk to it.
2. Chrome extension (optional)
- Open
chrome://extensions - Toggle Developer mode on (top right)
- Click Load unpacked
- Select the
extension/folder
Pin the icon from the puzzle-piece menu. The popup shows your desk time and should say Verifier: running.
To test it immediately: open any normal website (not
chrome://โ extensions can't inject there), then click the icon โ Take a break now.
3. VS Code extension (optional)
code --extensionDevelopmentPath=/full/path/to/BreakTension/vscode
A second window opens with BreakTension loaded. Confirm with Ctrl+Shift+P โ BreakTension โ three commands should appear:
| Command | What it does |
|---|---|
Show status |
Desk time left + which clients are live |
Take a break now |
Trigger a break immediately |
Pause or resume |
Stop counting editor time |
F5 from inside
vscode/also works, but the extension host starts paused waiting for the debug adapter and can time out before activating. The command above skips that handshake.
Both at once
Run BreakTension: Show status in VS Code while the Chrome extension is running. It reports both clients sharing one budget โ switching apps doesn't dodge the break.
Useful commands
node scripts/verify-image.js photo.jpg # judge any photo from the CLI
node scripts/verify-image.js --list-tasks # see the task pool
BT_MODEL=llava:13b npm run server # swap the model
BT_PORT=9000 npm run server # change the port
The CLI is the quickest way to see the model reason:
$ node scripts/verify-image.js window.jpg
FAIL 28.8s (failed at: glass)
outdoor=y glass=n screen=n task=n
The image clearly shows a window frame and glass panes.
Go all the way outside.
What I'd tell you if you build on this
Test the attack, not the happy path. My prompt looked excellent and failed 5 times out of 5 at the one job that mattered. I only found out because I built the cheat and pointed it at my own system.
And the thing nobody tells you about AI productivity: the hours it gives back don't come with instructions. If you don't decide what they're for, someone else will โ and it'll be more work.
I'm going outside.
Top comments (0)