This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
A few weeks ago, I was hanging out with my close friend and his partner at their apartment. They were in the middle of this hilarious, chaotic debate while trying to cook dinner together flour on the counter, genuine stomach-hurting laughter, totally absorbed in each other's company.
My friend suddenly stopped, smiled, and said, "Wait, stay right there, let me grab a picture!"
He fumbled for his phone, unlocked it, held up the screen, and counted down. Just like that, the magic vanished. Their shoulders squared up, their smiles froze into polite poses, and the spontaneous energy in the room was gone.
I’ve watched them run into this exact wall dozens of times: the second someone remembers to take a photo, the authenticity of the moment dies.
So for this challenge, I built Canchalant.
It’s an ambient, hands-free web app that turns any spare phone, tablet, or laptop into a discreet, fly-on-the-wall photographer. You simply prop your device up on a shelf, mug, or kitchen counter, hit "Start Watching," and walk away.
Canchalant silently watches the room in real time. It uses computer vision heuristics and Google’s open-weight PaliGemma vision model to detect when people are actually engaged in genuine, unposed interaction,bursting into laughter, sharing quiet eye contact, or gesturing mid-conversation. If someone stares directly into the lens with a stiff smile or moves too fast into a blurry smear, it ignores them. When an authentic moment happens, it silently captures it, writes a story caption, and indexes it into a searchable gallery where they can search their memories using everyday human thoughts.
Demo
- Live App: https://canchalant.vercel.app
- Interactive Backend API: https://canchalant.onrender.com/docs
- Uhmm please give it about 50 seconds for the backend to warm up since I am using a free tier of Render
Code
🎯 Canchalant
Candid + Nonchalant — An intelligent, private ambient photo assistant and semantic memory gallery.
Built as a gift for a friend and his partner. Canchalant uses an active camera stream (laptop webcam or phone browser) that continuously samples frames, runs real-time heuristic filters, classifies authentic unposed candid interactions via Google Gemma Vision (PaliGemma), uploads captures to Cloudinary (canchalant_snaps), and indexes dense CLIP vector embeddings in MongoDB Atlas Vector Search for natural language retrieval.
Say "sitting together laughing over coffee" and find exactly that moment.
🏆 Target Hacktoberfest Categories
- Best Use of Gemma ($200) — PaliGemma 3B open-weight vision model for candid classification, heartfelt captioning, and mood tagging.
-
Best Use of MongoDB Atlas ($100) — Atlas Vector Search (
$vectorSearch) over 512-dim multimodal CLIP embeddings.
✨ Features
-
🎥 Ambient Camera HUD — Real-time HTML5
getUserMediavideo canvas with animated target indicators, Laplacian blur metrics, and…
How I Built It
I wanted zero barrier to entry—my friend isn't going to compile native code, sideload APKs, or run command-line scripts. It had to be a slick, responsive web experience that works on any modern phone browser.
Here’s the journey under the hood:
1. The Living Room Observer (Vite + React & Tailwind CSS)
The frontend uses standard browser media streams (getUserMedia) with a custom camera canvas. It features a clean HUD with sensitivity toggles, capture intervals, and a dynamic flash animation whenever a real candid moment is detected. The gallery tab gives them instant access to high-res captures and mood tags.
2. Guarding the Compute (OpenCV Heuristics)
To prevent overwhelming the system with thousands of useless frames, I built a fast client-side/edge filter using OpenCV Laplacian variance. If someone is dashing across the frame or the lighting is pitch black, the frame gets dropped in under 5 milliseconds before any model even looks at it.
3. The Unposed Eye (Google PaliGemma via Hugging Face)
When a clear candidate frame passes, it hits Google's open-weight PaliGemma (google/paligemma-3b-pt-224). Instead of standard object detection, I prompted PaliGemma to act like an intimate documentary photographer:
- It categorizes the frame (
SPECIFIC_CANDID,GENERAL_CANDID,POSED, orJUNK). - It grades the candid confidence score.
- It writes a heartfelt, human caption of what makes the snapshot authentic along with emotional mood tags (like
["warmth", "belly-laugh", "cozy"]).
4. Semantic Memory Lane (Cloudinary & MongoDB Atlas)
Once a genuine candid is approved:
- The raw image is preserved on Cloudinary so it persists permanently without disappearing when servers cycle.
- The photo is passed through a dense CLIP vectorizer to create an embedding space.
- The metadata, tags, and embeddings are saved directly into MongoDB Atlas, where I configured Atlas Vector Search (
$vectorSearch).
Instead of scrolling through months of camera rolls, my friend can just type: "laughing while cooking" or "quiet afternoon reading" and Atlas Vector Search immediately surfaces the exact candid memories based on mood and meaning.
5. Production Cloud Deployment
The FastAPI backend service runs 24/7 on Render, communicating seamlessly with the cloud database and image CDN, while the frontend is deployed with full mobile camera SSL permissions. However it needs some warm up time [>50sec] for the backend to work since I am working on free tier
Why Does Open Innovation Matter?
Let’s be honest: setting up a camera in someone's home, kitchen, or living room involves deeply personal moments.
The idea of sending non-stop live video frames of my friend and his partner to proprietary, black-box cloud APIs—where data retention rules are opaque and images might be harvested to train private commercial models—felt completely wrong.
Open-weight AI changed the game here:
- Zero Tollbooths on Intimacy: Google’s open-weight PaliGemma allows powerful multimodal vision reasoning without bleeding money on per-frame API bills.
- From Laptop Rig to the Cloud: Because PaliGemma is open, I was able to build and benchmark the entire reasoning loop on my local workstation using quantized GGUF weights before wiring up the deployment pipeline. The architecture remains completely portable—my friend could easily run this on a home mini-PC or edge device down the line with full sovereignty over his memories.
- Tools Built for People, Not Platforms: Open innovation allows developers to build hyper-specific, thoughtful tools for the people in our lives without asking for permission from big tech walled gardens.
Top comments (0)