DEV Community

Benny
Benny

Posted on

CaptionBot: Say What You’re Selling. Get a WhatsApp Caption. Hacktoberfest Weekend Challenge — Build for a Friend

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

The Problem

A lot of small clothing sellers in Nigeria sell through WhatsApp Status.

The workflow is simple: take a picture, write a caption, post it, repeat.

The pictures are easy.

Writing a fresh caption for every item is the annoying part.

I wanted to make that part almost disappear.

What I Built

CaptionBot lets a seller describe what they’re selling by typing or speaking naturally.

For example:

“I get two Ankara gowns, size 12 and 14, 18k each, Lagos delivery.”

CaptionBot turns that into a clean, ready-to-post WhatsApp caption.

Then the seller can Copy it or Send to WhatsApp.

The idea is deliberately small. I am not trying to build another giant AI writing assistant. I wanted to solve one repetitive task for a specific type of person.

I built it for sellers like the ones already in my own contacts.

I have not put it in a seller’s hands yet, so I am not claiming user validation that I don’t have. I tested the workflow myself and recorded what actually broke.

Demo

https://youtu.be/qfmQImkSp24?is=twic1RjefQTXimIy

Code

CaptionBot GitHub Repository⁠

How It Works

The entire application is intentionally lightweight.

  • Python runs the application.
  • Ollama runs Gemma 3 1B locally.
  • Chrome voice typing lets the seller speak instead of type.
  • A small preprocessing step fixes common transcription mistakes.
  • Gemma generates the caption.
  • Code checks the result before showing it to the user.

I chose Gemma 3 1B because this project does not need a huge model.

The model download is about 815 MB. In a market where mobile data is not something I want a seller constantly spending money on, I wanted the actual AI generation to happen locally on an ordinary laptop.

That led to a simple design:

Gemma handles the language.
My code handles the rules.

The Part That Didn’t Work

My first version had a bug.

I gave CaptionBot a gown with no price.

It generated a caption containing 40k.

I never gave it 40k.

The reason was embarrassing but useful: one of my prompt examples contained a 40k lace gown. The small model copied the example into a completely different product.

That became one of the most important lessons from building CaptionBot:

A small model can produce a very convincing answer and still invent a detail that was never provided.

There were other problems too.

Chrome’s voice typing heard:

  • “lace” as “list”
  • “12k” as “12 key”

It also struggled with Nigerian Pidgin in some cases. Typed Pidgin worked much better.

So I fixed the system instead of hiding the failure.

I:

  1. Removed the misleading examples.
  2. Added a stronger instruction not to invent product details.
  3. Added a code-level price check.
  4. Made the system reject a generated price if that price was never provided.
  5. Added retries, up to three attempts, when the generated caption fails validation.

The repository contains the fixed version.

The demo video still shows the original failure because I wanted to show what actually happened during development.

Why I Used Gemma

Caption generation is a small task.

The seller is not asking the model to research the internet, write an essay, or solve a complicated problem. They are giving it a handful of product details and asking it to turn them into natural language.

Gemma 3 1B was enough for that.

More importantly, running it locally gave me something I would not get from simply calling a hosted API:

I could see exactly where the model ended and my application began.

When I discovered that the model could invent a price, I could put the guardrail directly in my own code.

The model generates.

The application verifies.

Why Open Innovation Matters

For a small seller, paying an API every time they need a caption adds unnecessary cost to a very small task.

CaptionBot can run the model locally, so there is no per-caption AI bill.

But the bigger point for me is control.

Because the model and application pipeline are accessible, I can experiment with the prompt, change the examples, add validation, and test different approaches myself.

The 40k bug is a good example.

I didn’t need to wait for a model provider to fix it.

I found the failure, understood why it happened, changed the prompt, and added a code-level check so the same mistake would be harder to repeat.

That is what open innovation means to me here:

not just using an open model, but being able to build around it, break it, understand it, and improve it.

There is one limitation worth being clear about: the AI inference is local, but Chrome’s voice-typing feature uses Google’s servers. So the entire input-to-output pipeline is not fully local.

What I Would Test Next

The next step is not adding more features.

It is putting CaptionBot in front of actual sellers.

I would want to learn:

  • Whether sellers prefer speaking or typing.
  • How well it handles Nigerian Pidgin and local product names.
  • What happens with incomplete product information.
  • Whether sellers actually use the WhatsApp send button.
  • Which types of clothing descriptions need the most correction.
  • What other repetitive parts of their selling workflow could be automated locally.

That would turn CaptionBot from something I built for sellers into something shaped by actual sellers.

Prize Category

Best Use of Gemma

CaptionBot uses Gemma 3 1B as a local caption-generation model, while application-level validation handles important factual constraints such as prices and sizes.

The result is a small, local-first tool built around one practical problem:

A seller says what they’re selling. CaptionBot makes it ready to post.

Top comments (1)

Collapse
 
respect17 profile image
Kudzai Murimi •

This is a good pick, so many small sellers lose sales because typing out a good post is a hassle. Speak it and get a caption back is exactly the right amount of friction removed.