Originally published on the NuWay Biz Solutions blog.
✦ Cover image: Made with Kling 3.0 via OpenArt (paid, standard 720p mode) + SeedVR2 + FLUX.1 schnell (local) — we're transparent about AI. See the exact prompt on the original post.
Halfway through opening a soda can, the monkey pulled a second pull tab out of thin air.
Google's Veo 3.1 had been asked, in those exact words, for "the tab stays attached." It delivered a tab that stayed attached — plus a spare, in case of emergencies. That shot cost 34 cents.
Over three days in October I made the same soda ad nine ways: no AI at all, free models on my local desktop, two paid models, and a mix.
The short version of what I learned has nothing to do with which model to buy. A small business can make a usable AI video ad for pocket change, and the model matters less than the source material you feed it and how carefully you check what comes out. And the bloopers had me laughing out loud.

▶ The can-opening test. Every version's can-opening (the 3D one skips it), in the order they were made. (Sound on.) Watch it on the original article.
| # | Version | Cost to make | What broke |
|---|---|---|---|
| 1 | Pure code, no AI model | $0 | A robotic "ahhh" |
| 2 | 3D cartoon, free models | $0 in model fees | Every take that drank, leaked |
| 3 | Photoreal, free v2 | $0 in model fees | Three different monkeys |
| 4 | Veo 3.1 (Lite mode*) + our sound | $2.38 ($0.34 per 4-second shot) | A second pull tab |
| 5 | Veo 3.1 (Lite mode*), its own sound | Same $2.38 | The splash sounded wrong |
| 6 | Photoreal, free v3 | $0 in model fees | Nuzzled a can that was still closed |
| 7 | Photoreal, free v4 | $0 in model fees | Hands like rubber gloves |
| 8 | Kling 3.0, its own sound | $3.47 (~$0.50 per 5-second shot) | Near-silent on 3 of 7 shots; a stray musical tone |
| 9 ★ | Hybrid: #7 plus one Kling shot | ~$0.50 (one 5-second Kling shot) | A visible seam where the paid shot sits |
*Our tool sent no mode, so OpenArt's default applied. Today that default is Lite, and OpenArt keeps no per-clip record of the mode, so Lite is our best inference.
Why make a monkey open a can?
A short, 5-second clip portraying a monkey in sunglasses drinking a can of pop… a marketing-type commercial… enjoy the can with a lip-smacking 'ahhh'.
— My brief, 4 October 2026 (trimmed)
The soda, BANANA BOOM, doesn't exist, which is a shame, because I'd try a banana soda. The brief sounds simple, but its five seconds hold four things that turned out to be exactly where things went wrong: hands, a tiny mechanism, liquid going where it should, and a crack that has to land on the exact frame.
The AI doing the work was Claude Code, Anthropic's coding agent: it wrote the scripts, ran every model, cut the edits and mixed the sound, and separate Claude agents reviewed each result. I wrote the brief, watched every cut and decided what shipped. When Claude got something wrong, I'll say so by name.
Level 0: no AI at all
Monkey #1 is drawn in code with Remotion, and every sound was synthesized from scratch. Twelve minutes to build, 38 seconds to render, $0. The "ahhh" was robotic, because math is a poor substitute for a satisfied primate.

▶ #1: no AI at all. Drawn in code with Remotion, every sound synthesized. $0. 5 seconds. (Sound on.) Watch it on the original article.
For a logo sting or a mascot that must look identical every time, code still wins: the brand name never misspells itself.
Level 1: a 3D cartoon, free, on my desktop
The first real AI version ran entirely on my own desktop, with two free models doing the picture work, both licensed for commercial use:
| Part | What I used | Its job | Cost |
|---|---|---|---|
| The computer | My desktop PC (NVIDIA RTX 3080 graphics card, 10 GB; Ryzen 9 5900X; 64 GB RAM) | Ran everything locally | Already owned |
| Picture model | FLUX.1 schnell | Drew each starting picture | $0 |
| Video model | Wan 2.2 (Alibaba's open model) | Animated each picture, about 10 minutes per 5-second clip at 720p | $0 |
Then the monkey drank the soda, and the soda went down his chest.
So the next take's prompt said, in so many words, "no spilling and no dripping."
It still spilled.
Claude then asked for four short clips to bridge the gaps, each one told "completely dry: no liquid, no dripping, no spilling." All four dripped. A product shot with no monkey in it at all poured soda over a can that was still closed. The only dry take was the one where he didn't drink.
Two drinks, four bridges, six leaks. I've never seen an instruction ignored with such commitment.
We shipped it anyway, by cutting around the leaks. A lot of AI video editing is exactly that: rescuing the usable frames.

▶ #2: 3D cartoon, free, on my desktop. FLUX.1 schnell + Wan 2.2, voice by Chatterbox, all on my desktop. $0 in model fees. 9.8 seconds. (Sound on.) Watch it on the original article.
The blooper reel collects nine things that went wrong across the whole project, one clip each.

▶ Bloopers: what went wrong. The spills, the drips, the conjured can, the human hand, three monkeys, the second pull tab, the closed can, the smoke hand and Kling's musical tone. (Sound on.) Watch it on the original article.
Level 2: photoreal, still free
Next the brief went photoreal and gained a story, with the monkey finding the can first, so everything from here runs longer, from 18 to 27 seconds.
3, the first free photoreal cut, cast three different monkeys in one ad: one with a black cap and a peach face, one with coarse brown fur, one with a slate-blue face and a white ruff. The hand opening the can melted into a fingerless mitten.
On an empty jungle shot, Claude's prompt still mentioned the can, so Wan grew a giant, misspelled one in the forest, like a monument to the prompt. And in another take with no monkey, a human hand reached in and took the can. Somebody else filming on set wanted a soda too...
The fix never touched the video model. Claude rewrote the starting pictures to describe the same monkey, hands included, in every shot, and three species became one.
For #7, Qwen-Image-Edit, a free image-editing model, built every picture of the monkey from one reference image. Wan drafted each shot at 480p, and SeedVR2, a free upscaler, scaled it up to 1080p.
Same monkey in every shot, a finger that lifts the tab like a lever, a wide "AHHH." His hands look like rubber gloves, but model fees were still $0.
Lesson one: the starting pictures decide more than the video model does.

▶ #7: photoreal, free v4. Locked identity + Wan 2.2 at 480p + 1080p upscale, on my desktop. $0 in model fees. 23.5 seconds. Watch it on the original article.
Level 3: what does $6 of paid AI video buy you?
Both paid models came through OpenArt, a site that resells credits for many models; its Plus plan is $34 a month for 12,000 credits. Claude gave both models the identical starting pictures and prompts, though not the same settings: Veo ran at 1080p in four-second shots, Kling in its standard 720p mode in five-second shots.
Veo 3.1 (from Google) cost $0.34 a shot, $2.38 for seven. It grew the spare tab and acted the "ahhh" without voicing it: mouth open, a lick of the lips, silence.
One caveat: Veo most likely ran in Lite, the lowest of OpenArt's three Veo 3.1 settings (see the note under the first table), so this probably wasn't Veo at its best.
Kling 3.0 (from Kuaishou) cost about $0.50 a shot, $3.47 for seven. A finger hooks the ring, the tab rises, the spray comes out of the top, and the crack lands right on it.
Best sequence and sound of a can opening by far.
— My note on Kling's can-opening
Then the monkey's palm turns magenta, as the can's color bleeds into his hand. Apparently the paint was still wet.
Two flaws they shared were ours.
The monkey changes coats between shots (jet black in one shot, a small brown juvenile in another, a cream mane in a third), and the can-opening shot has a grey hand with flat, human-looking fingernails. Both were already in the starting pictures, and both models faithfully animated what they were handed. Lesson one again, from the paid side.

▶ All nine at once. All nine versions, lined up so every monkey's "ahhh" lands on the same frame. (Sound on.) Watch it on the original article.
Can the AI do the sound too?
Both paid models make sound along with the picture. Neither voiced the "ahhh," the one sound the brief asked for by name.
Kling's own audio was near-silent on three of seven shots. And in the shot that introduces the can, where Claude's prompt asked for stream water, a "glint" chime and "no music," Kling added a sustained musical tone. Interesting addition...
I called Veo's own splash "laughably bad," and part of the joke was on me: Claude's audio processing had squashed it.
The sound that won was the one we built: every effect placed on a measured frame. Our own "ahhh" came from Chatterbox, a free voice model, and a reviewer measured its pitch: an adult man's voice coming out of a small monkey.
So, as a tenth step, we swapped in real recordings for the jungle, the birds and the "ahhh," from Sonniss's free sound library. It was more believable, though still too low for a small monkey, and now I could hear the last fakes: two synthesized metal "tinks" when the monkey's hand grabs the can. A palm on an aluminum can doesn't ring like a bell.
Level 4: the winner was a hybrid
The version I'd stand behind, despite the remaining flaws, is #9: the free #7, with one shot swapped for Kling's can-opening, crack and fizz included. The paid part is one five-second shot, about 50 cents. My note when I watched it: "the best one so far."
It also has the flaw I'm most embarrassed by.
Look at the image at the top of this page: the monkey opening the can has a cream torso. A second later, drinking, he has black arms. Same monkey, two wardrobes.
The seam came from our side
That Kling shot was made from the comparison's older starting pictures, before the monkey's identity was locked. Claude reused it in the locked cut without re-making it. A re-take costs about 50 cents a go. The rule: any shot you splice in has to start from the same pictures as the rest of the cut.
Lesson two: pay only for the hard shot. Free models were good enough for the scenery, the can on its rock and the drink. The paid model was clearly better at a hand working a tiny mechanism, so that's the shot worth paying for.

▶ #9 ★: the hybrid (featured). #7 plus Kling 3.0's can-opening shot and its own crack and fizz. ~$0.50 for the one Kling shot. 23.4 seconds. Watch it on the original article.
What it really cost, in money and time
| Item | Cost |
|---|---|
| Veo 3.1 (Lite mode*), seven 4-second 1080p shots | 840 credits = $2.38 |
| Kling 3.0, seven 5-second 720p shots | 1,225 credits = $3.47 |
| OpenArt Plus plan | $34 a month (12,000 credits) |
| Free models and the Sonniss sound library | $0 in fees (electricity not metered) |
| Claude, at published prices | About $120 |
| Claude, what I actually paid | $0 extra on my $100-a-month plan |
| Time | About ten hours of active work over three days |
*As above.
So all fourteen paid shots came to (drumroll please, though you've already seen it twice)...
$5.85.
Measured from the session logs, the Claude work behind the nine versions came to about 1,200 model calls: the main session plus 24 review and research agents. At Anthropic's published API rates, that would have cost about $120.
I didn't pay that. It all ran on the $100-a-month Claude Max plan I already had, so the Claude side added $0 in cash.
That $120 also isn't the price of one ad. It covers all nine versions and a lot of trial and error, and along the way Claude wrote the scripts and we worked out a process the next video can reuse. So the next one should need less Claude work, though I haven't measured how much less.
Doing all of this by hand would take longer, a lot longer, but I can't tell you how much longer: I'm not a professional video editor.
Lesson three: check it like it's trying to fool you
From the 3D cartoon on, every file went to a reviewer before I signed off, and the reviewers caught things before they shipped: Veo's second tab, which we cut down to three frames in the final cut; audio running 52 milliseconds late; a "recorded" jungle sound that turned out to be synthesized; and Claude telling a reviewer a gap in the sound was covered when it wasn't.
We've written about this for text: making got cheap and checking didn't. With video it's worse, because a defect can hide in three frames. Step through the frames at full resolution before anyone else sees them.
Where AI video stands in October 2026
The free video model I used, Wan 2.2, released in July 2025, is fifteen months old. Alibaba's Wan team hasn't released a newer general-purpose video model for download since; its newer ones are rent-only. Every free-route gain here came from what we fed Wan 2.2 and the tools around it.
RigidBench, a University of Cambridge preprint from August 2026, tested eight video models, including Veo 3.1 and Kling 3.0, though at higher settings than ours (see the note under the first table). No model led on all ten measurements, and the usual "does it look right" score ranked them almost backwards from how accurately they moved objects. Looking right and moving right are different skills, and my extra pull tab is a small, silly example.
What I didn't test: OpenAI's Sora, Runway, and the AI video features in Canva and CapCut. I stuck to what a small business can run on its own desktop or rent one shot at a time, so I can't say how they'd do. Sounds like a test for another day.
If this were your ad, what would I use?
| If you need… | Use | Rough cost (our numbers) |
|---|---|---|
| A logo sting, a price card, a mascot that must match every time | Code or a motion template (Remotion is free for individuals and companies of up to three employees; larger ones pay) | $0 in model fees |
| Scenery, b-roll, a product looking good | Free open models on a desktop with a decent graphics card | $0 in fees; ~10 min per 5-second clip at 720p |
| The hero moment: hands, a small mechanism, liquid | Pay for that one shot, and budget a few attempts | ~$0.34–$0.50 per attempt through a reseller |
| Sound | Real library recordings placed on the frame | $0 with a free library |
| Your real label, a real person, a product claim | Film it | Most of our AI versions restyled or smeared our made-up label |
**No desktop and no coding agent? **Use a browser reseller like OpenArt and pay per shot.
8 was all paid ($3.47), though its starting pictures came from my desktop. On OpenArt you'd make them with its image models or start from your own photos, and commercial use needs its Plus plan or above.
**A computer, but no graphics card that can handle it? **Rent one. RunPod lists an RTX 4090, a step up from mine, from $0.34 an hour on its Community Cloud, and fal.ai will run the same Wan 2.2 model for you at $0.08 per second of 720p video, about 40 cents for a five-second clip. Those are list prices from October 2026. I haven't tried either, so I can't tell you what a finished ad costs that way.
Whichever route: make one picture of your character that you like, build every other starting picture from it, keep each prompt to what's actually in the frame, and check the result frame by frame. If nobody on your team has time for that, the video fees aren't the real price.
If video is your craft
This is about a business making its own marketing. For creative studios our position hasn't moved: automate the business, never the art. And purely AI-generated footage may not be something you can own; the slop piece covers why.
If you'd rather not spend three days learning which shots to pay for, start a no-pressure conversation. We'll tell you which route fits what you're selling, including when the honest answer is code, or a camera.
Practical AI. Clear process. Real business value.
— Brian, NuWay Biz Solutions
_P.S. _Claude has since made a free fixed cut. The tinks are gone (I picked silence over real hand sounds), the "ahhh" is pitched up to suit a small monkey (I picked it blind from ten takes), and a patch covers the end of the label smear. The patch left a faint seam of its own. Of course it did. Next time: before and after, side by side.
Top comments (0)