The Problem
A lot of small clothing sellers in Nigeria sell through WhatsApp Status.
The workflow is simple: take a picture, write a caption, post it, repeat.
The pictures are easy.
Writing a fresh caption for every item is the annoying part.
I wanted to make that part almost disappear.
What I Built
CaptionBot lets a seller describe what they’re selling by typing or speaking naturally.
For example:
“I get two Ankara gowns, size 12 and 14, 18k each, Lagos delivery.”
CaptionBot turns that into a clean, ready-to-post WhatsApp caption.
Then the seller can Copy it or Send to WhatsApp.
The idea is deliberately small. I am not trying to build another giant AI writing assistant. I wanted to solve one repetitive task for a specific type of person.
I built it for sellers like the ones already in my own contacts.
I have not put it in a seller’s hands yet, so I am not claiming user validation that I don’t have. I tested the workflow myself and recorded what actually broke.
Demo
https://youtu.be/qfmQImkSp24?is=twic1RjefQTXimIy
Code
CaptionBot GitHub Repository
How It Works
The entire application is intentionally lightweight.
- Python runs the application.
- Ollama runs Gemma 3 1B locally.
- Chrome voice typing lets the seller speak instead of type.
- A small preprocessing step fixes common transcription mistakes.
- Gemma generates the caption.
- Code checks the result before showing it to the user.
I chose Gemma 3 1B because this project does not need a huge model.
The model download is about 815 MB. In a market where mobile data is not something I want a seller constantly spending money on, I wanted the actual AI generation to happen locally on an ordinary laptop.
That led to a simple design:
Gemma handles the language.
My code handles the rules.
The Part That Didn’t Work
My first version had a bug.
I gave CaptionBot a gown with no price.
It generated a caption containing 40k.
I never gave it 40k.
The reason was embarrassing but useful: one of my prompt examples contained a 40k lace gown. The small model copied the example into a completely different product.
That became one of the most important lessons from building CaptionBot:
A small model can produce a very convincing answer and still invent a detail that was never provided.
There were other problems too.
Chrome’s voice typing heard:
- “lace” as “list”
- “12k” as “12 key”
It also struggled with Nigerian Pidgin in some cases. Typed Pidgin worked much better.
So I fixed the system instead of hiding the failure.
I:
- Removed the misleading examples.
- Added a stronger instruction not to invent product details.
- Added a code-level price check.
- Made the system reject a generated price if that price was never provided.
- Added retries, up to three attempts, when the generated caption fails validation.
The repository contains the fixed version.
The demo video still shows the original failure because I wanted to show what actually happened during development.
Why I Used Gemma
Caption generation is a small task.
The seller is not asking the model to research the internet, write an essay, or solve a complicated problem. They are giving it a handful of product details and asking it to turn them into natural language.
Gemma 3 1B was enough for that.
More importantly, running it locally gave me something I would not get from simply calling a hosted API:
I could see exactly where the model ended and my application began.
When I discovered that the model could invent a price, I could put the guardrail directly in my own code.
The model generates.
The application verifies.
Why Open Innovation Matters
For a small seller, paying an API every time they need a caption adds unnecessary cost to a very small task.
CaptionBot can run the model locally, so there is no per-caption AI bill.
But the bigger point for me is control.
Because the model and application pipeline are accessible, I can experiment with the prompt, change the examples, add validation, and test different approaches myself.
The 40k bug is a good example.
I didn’t need to wait for a model provider to fix it.
I found the failure, understood why it happened, changed the prompt, and added a code-level check so the same mistake would be harder to repeat.
That is what open innovation means to me here:
not just using an open model, but being able to build around it, break it, understand it, and improve it.
There is one limitation worth being clear about: the AI inference is local, but Chrome’s voice-typing feature uses Google’s servers. So the entire input-to-output pipeline is not fully local.
What I Would Test Next
The next step is not adding more features.
It is putting CaptionBot in front of actual sellers.
I would want to learn:
- Whether sellers prefer speaking or typing.
- How well it handles Nigerian Pidgin and local product names.
- What happens with incomplete product information.
- Whether sellers actually use the WhatsApp send button.
- Which types of clothing descriptions need the most correction.
- What other repetitive parts of their selling workflow could be automated locally.
That would turn CaptionBot from something I built for sellers into something shaped by actual sellers.
Prize Category
Best Use of Gemma
CaptionBot uses Gemma 3 1B as a local caption-generation model, while application-level validation handles important factual constraints such as prices and sizes.
The result is a small, local-first tool built around one practical problem:
A seller says what they’re selling. CaptionBot makes it ready to post.
Top comments (1)
This is a good pick, so many small sellers lose sales because typing out a good post is a hassle. Speak it and get a caption back is exactly the right amount of friction removed.