This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Pixel Leak checks an outdoor photo for clues to where you live before you post it. You drop in a photo, wait about ten seconds, and see what a stranger could learn from it. Then you blur those parts and download a clean copy.
Most people already know the GPS tag is a problem, and plenty of tools strip it. But metadata isn't the only leak. The street sign behind your kid, the number on your gate and the plate on the car in your driveway are all readable by anyone who zooms in, and stripping EXIF does nothing about them.
Pixel Leak checks both:
- The metadata: GPS coordinates, capture time, device.
- The pixels: street and locality names, house numbers, PIN/ZIP/postcodes, phone numbers, number plates, people and screens.
Who it's for: anyone who posts the outdoor part of their life. That includes runners whose route starts at the front door, parents posting the kids in the garden, people who share their morning walk, and anyone who has held back a nice photo because it showed a little too much.
How it fits Touch Grass: the theme asks for builds where the screen is the shortest part of the experience. Pixel Leak is a ten-second stop between the walk and the post. It doesn't keep you on a screen. It gets you back outside faster, with one less thing to worry about when you share.
Demo
Live: https://tarunvashishth.github.io/pixel-leak/
To try it without using your own photo, click Try a sample photo. The first visit downloads the model (about 350 MB), and your browser caches it after that. It runs fastest in Chrome or Edge with WebGPU, and falls back to CPU (slower) elsewhere.
Code
tarunvashishth
/
pixel-leak
Check a photo for location leaks before you post it — EXIF and the pixels. Open-source vision model (Florence-2) running entirely in your browser; nothing is uploaded.
Pixel Leak
Check a photo for location leaks before you post it. It covers the metadata and the pixels.
Stripping EXIF removes the GPS tag. It does nothing about the street sign, house number or number plate in the frame. Pixel Leak runs Florence-2 (MIT-licensed, ~230M params) entirely in your browser to read the photo the way a stranger would. It then lets you blur what it finds and download a clean copy.
Live: https://tarunvashishth.github.io/pixel-leak/
Nothing is uploaded. The only network request is the one-time model download from Hugging Face (~350 MB, cached by the browser). After that you can go offline and it still works.
What it checks
| Source | Finds |
|---|---|
EXIF (via exifr) |
GPS coordinates, capture time, device |
Florence-2 <OCR_WITH_REGION>
|
street / locality names, house numbers, PIN/ZIP/postcodes, phone numbers, emails, licence plates, other visible text |
Florence-2 <OD>
|
people, licence plates, screens, signs, vehicles |
Florence-2 <MORE_DETAILED_CAPTION>
|
a plain-language |
How I Built It
Everything runs in the browser tab. There's no backend at all, and the site is a static page on GitHub Pages.
-
Metadata.
exifrreads GPS, capture time and device from the file. - Pixels. Florence-2-base (about 230M parameters, MIT license) runs in a Web Worker through Transformers.js. It does three passes over the photo:
| Florence-2 task | What I use it for |
| --- | --- |
| <OCR_WITH_REGION> | every piece of text, with a box around it |
| <OD> | people, vehicles, screens, signs |
| <MORE_DETAILED_CAPTION> | a plain-English "what a stranger sees" |
On WebGPU the vision encoder runs in fp16 and the text encoder and decoder in 4-bit. Without WebGPU it falls back to 8-bit on CPU.
-
Rules decide what counts as a leak. A small, unit-tested rules file turns the model's output into findings. It looks for street suffixes (
Road,Marg,Sector,Nagar,Ave…), Indian PIN codes, US ZIPs, UK postcodes, Indian plate formats, phone numbers, emails and house numbers. Each finding gets a severity and a numbered box on the photo. -
Clean copy. The photo is redrawn at full resolution on a
<canvas>with your chosen areas blurred, then re-encoded. Only pixels survive. I checked the output for EXIF, XMP and GPS markers, and none were there.
A scan takes 8–11 seconds on my Mac with WebGPU.
Things I learned building it
Don't let the model be the judge. On my test photo, Florence-2's caption described a house number as "2218". The OCR pass on the same photo read it correctly as "221B". If I had built the leak detection on top of the caption, it would have been confidently wrong. So the model only reads, and plain rules decide what's sensitive. Every flag can be traced to a specific piece of text and a specific regex, and the rules run as tests in CI before every deploy.
Model boxes are approximate. In my tests Florence-2's boxes landed near the text, but were sometimes a little offset or too wide. So the blur covers a padded area around each box rather than the exact box, and I apply the canvas blur twice so large sign lettering doesn't stay readable through a light blur.
Re-encoding is a fair trade here. My other project, remove-exif.com, strips metadata byte by byte so the image data is never touched. Here that wasn't possible: once you blur pixels you have to re-encode them anyway. Doing it through a canvas also guarantees that no metadata block survives, because the canvas never had any.
"Works offline" is easy to say, so I tested it. I loaded the model, cut the browser's network, and ran a scan. It finished in 8.3 seconds and found every leak, with zero network requests during the scan.
What it can't do (yet)
- It flags clues. It can't promise anonymity. Skylines, landmarks, shadows and reflections can still identify a place, and it doesn't reason about those.
- The object detector is the weakest part. On my sample photo it caught every piece of text but didn't flag the (cartoon) car itself. Always look at the photo yourself before posting.
- Small or angled text is hit and miss for OCR.
- It doesn't open HEIC (the iPhone default). For now, export the photo as JPEG first.
Why Does Open Innovation Matter?
Because the honest version of this tool can't be built on a closed API.
A cloud vision API would mean uploading the photo you're worried about, the one with your street name in it, to a third party so it can tell you whether the photo reveals your street. That defeats the point.
An open-weight model changes that:
- You don't have to take my word for it. The model downloads once, and after that nothing leaves the tab. Open DevTools, watch the Network panel and run a scan. Or turn off Wi-Fi.
- It can stay free. There's no API key, no per-image cost and no account. The whole thing is a static site, so hosting costs nothing and there's nothing to shut down when credits run out.
- It's easy to fork for your own country. The rules currently know Indian, US and UK address formats. Adding Brazilian CEPs or German plate formats is a regex in one file. And if a better small vision model comes out next month, swapping it is a change to one worker file.
My Agent Session
I built Pixel Leak pairing with Claude Code. I picked the idea from a shortlist built around my existing repos, and the agent did most of the implementation and helped draft this post. Every claim in it was checked against the running app before it went in: the scan times, the offline test, and the metadata check on the clean copy.

Top comments (0)