DEV Community

Cover image for Keepsake: Everyone's Trip Photos, Sorted into Moments Without Anyone Lifting a Finger
Vatsal Lakhmani
Vatsal Lakhmani

Posted on

Keepsake: Everyone's Trip Photos, Sorted into Moments Without Anyone Lifting a Finger

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Every group trip ends the same way.

Everyone comes home with hundreds of photos. Then, for weeks, every conversation about the trip turns into this:

"Bro, send me the photos from the beach."
"Which beach?"
"The one with the sunset. Ask Rajat, his were better."
"Rajat's are on his old phone."

Nobody is lazy. The photos are just scattered across four phones, three chat apps and a shared folder that two people uploaded to and nobody organized. Finding one photo takes longer than the moment it captured, and in the end most of them are never seen again.

I built Keepsake for my friends so the next trip doesn't end that way.

Keepsake turns everyone's photos into one organized album, and nobody has to organize anything:

  1. Someone creates a trip and sends one invite link.
  2. Everyone signs in and drops in their photos.
  3. One button groups the chaos into moments, like Baga Beach, Dinner or Fort Aguada, each with photos from everyone who was there.
  4. You search in plain words, like "sunset at the beach".
  5. Keepsake creates a short narrated recap of the trip from your own photos.

The whole product is one promise: Upload everything. Organize nothing.

What it changes

Keepsake doesn't make something impossible possible. People could always find their photos eventually. It removes the friction that makes people give up, and that friction has real effects:

  • Nobody has to be the organizer. On every trip, one person ends up doing the unpaid work of collecting and sorting everyone's photos, and usually it never gets done. Here, everyone's only job is to upload.
  • Everyone's version of a moment ends up in one place. At dinner, you took two photos and three friends took twenty. Keepsake puts them together, so you see the evening the way the whole table saw it.
  • Photos get found and shared while the memory is still fresh, not weeks later or never.
  • A trip becomes a story, not a folder. The recap turns a gallery into something you can play back, or send to the friend who couldn't come.

Demo

Try it live: https://keepsake-gray.vercel.app

Sign up, create a trip, add a few dozen photos and press Organize photos. The sign-in form shows a "Development mode" badge because the project uses Clerk's test keys; it works normally. If the backend has been idle, the first request takes about a minute while it wakes up.

Code

GitHub logo watzal / KeepSake

Keepsake is a shared trip album. Everyone on the trip adds their photos, and Keepsake groups them into the places and moments you actually remember, names each one, lets you search by describing a scene, and reads a short recap of the trip aloud.

Keepsake

Upload everything. Organize nothing.

Keepsake is a shared trip album. Everyone on the trip adds their photos, and Keepsake groups them into the places and moments you actually remember, names each one, lets you search by describing a scene, and reads a short recap of the trip aloud.

Live demo: https://keepsake-gray.vercel.app

What it does

  • Shared trips. Create a trip and send friends an invite link. Everyone uploads into the same album.
  • Automatic moments. Photos are grouped by what they show and when they were taken, so "the beach", "dinner" and "the fort" become separate sections without anyone sorting anything.
  • Named for you. Each moment gets a short name and a one-line description written by Gemma.
  • Search by description. Type "sunset at the beach" and get the matching photos, no tags needed.
  • Narrated recap. A short story of the trip, written from its moments and read aloud.
  • Handles real uploads.…

The repository has two parts: backend (FastAPI, the photo pipeline, tests) and frontend (Next.js). The README covers local setup and which key switches on which feature. Every outside service is optional except the database and sign-in: without a Gemma key, moments are called "Moment 1", "Moment 2", and without an ElevenLabs key the recap is text only.

How I Built It

The problem with the obvious approach

The obvious way to organize photos with AI is to send every photo to a vision model and ask what's in it. For a thousand photos, that's a thousand slow, costly calls, and the answers are inconsistent ("beach", "seashore" and "coast" mean the same thing to a human).

So I flipped it: don't ask the AI about every photo. Group the photos first, then ask the AI about each group.

Photos from everyone
   → CLIP embeddings (one 512-number vector per photo)
   → HDBSCAN clustering on look + shooting time
   → Gemma names each cluster from 4 representative photos
   → MongoDB Atlas stores moments and vectors
   → Atlas Vector Search powers "find that photo"
   → Gemma writes a recap, ElevenLabs reads it aloud
Enter fullscreen mode Exit fullscreen mode

Stack: Next.js + Tailwind (frontend, on Vercel), FastAPI (backend, on a Hugging Face Docker Space), MongoDB Atlas, Clerk (sign-in), CLIP, HDBSCAN, Gemma 4, ElevenLabs, Cloudflare R2 (photo storage) and Sentry.

Step by step

1. Seeing the photos (CLIP). Each photo becomes a vector, a list of 512 numbers describing what it looks like. Similar photos get similar vectors. The model runs inside my own backend, not behind someone else's API.

2. Finding the moments (HDBSCAN). I cluster the vectors. I picked HDBSCAN because I don't know how many moments a trip has, and it copes with photos that don't fit anywhere. I also add each photo's shooting time as an extra feature, because looks alone can't tell two similar dinners on different nights apart, and time can.

3. Naming them (Gemma). For each cluster, I send the four photos closest to its centre to Gemma and ask for a short name and a one-line description. So a trip costs a dozen or two Gemma calls instead of one per photo. Categories aren't hardcoded either: a beach trip produces beaches and nightlife, and a trip to Tokyo would produce Shinkansen and temples.

4. Finding the one photo (Atlas Vector Search). CLIP puts photos and text in the same space, so I embed whatever you type and ask MongoDB Atlas for the nearest photo vectors, filtered to your trip. If the index is still building, the backend falls back to a local search and tells the UI which one ran.

5. Sharing (Clerk + FastAPI). Trips have members and invite links. Every route checks membership on the server, so someone who isn't on the trip just gets "not found". The backend verifies Clerk's signed session tokens itself.

6. The recap (Gemma + ElevenLabs). Gemma writes a narration of about 100 words from the moments, and ElevenLabs reads it aloud. If either service is down, the app still works: it falls back to a plain script, or to the text without audio.

Seeing inside it (Sentry)

I traced the pipeline as an agent run in Sentry. One trace shows each embedding batch, the clustering step as a tool call, every Gemma call with its token counts, and the database writes. The recap is a second run that adds the ElevenLabs call. When something felt slow, I didn't have to guess which stage to blame, because I could open the trace and see.

Known limits

I'd rather tell you than have you find out:

  • Photo files are protected by unguessable URLs, not by login.
  • It finds scenes and objects, not specific people's faces.
  • Organizing runs as a background task on one server, so a restart interrupts a run. A real deployment would use a job queue.
  • Invite links can't be revoked yet.
  • On the free Gemma tier, a trip with many moments waits on the per-minute request limit, so naming can take a minute or two.

Why Does Open Innovation Matter?

These are photos of real people at their most relaxed, which is the data you'd least like to hand over carelessly. Open models changed what I could build:

  • The photo understanding stays with me. CLIP and the clustering run on my own server. Only a few small thumbnails per moment go out for naming, not the whole library.
  • I could choose where Gemma runs. I used the hosted Gemma API for this submission, but the code has a one-setting switch to local Ollama, so a fully private version is possible. With a closed vision API, that choice wouldn't exist.
  • Cost doesn't grow with every photo. Because the pipeline is built from parts I can run and inspect, I could design the cost out, grouping first and asking the model once per moment instead of once per image.
  • I could debug and tune it. When two moments merged, I changed one number (how strongly time counts) and ran it again. A black box would have left me guessing.

Prize Categories

  • Gemma: names the moments and writes the recap script
  • MongoDB Atlas: application data and Atlas Vector Search for semantic photo search
  • ElevenLabs: the narrated trip recap
  • Sentry: agent tracing across the pipeline

Top comments (0)