I Made the Same AI Video Ad 9 Ways
One soda ad, one monkey, nine routes, from pure code to free models on my desktop to two paid models. The video fees came to $5.85. The starting pictures mattered more than the model, the checking caught what I missed, and the bloopers were worth the price of admission.
Kling 3.0 via OpenArt (paid, standard 720p mode) + SeedVR2 + FLUX.1 schnell (local)Same prompt, 4 AI models — swipe to compare. Showing 1 of 4.
Made with Kling 3.0 via OpenArt (paid, standard 720p mode) + SeedVR2 + FLUX.1 schnell (local) (Kling 3.0 (Kuaishou) image-to-video, upscaled to 1080p by SeedVR2 3B (ByteDance Seed, Apache-2.0); keyframe by FLUX.1 [schnell] (Black Forest Labs, Apache-2.0))view prompthide prompt
VIDEO, Kling 3.0 (as sent; OpenArt keeps no prompts, seed not exposed): "The monkey's thumb lifts the pull tab and the can cracks open: a quick puff of fine fizzy mist rises from the opening in the top of the can, glittering in the sunlight; the tab stays attached. Sound: A crisp, loud can crack and a fizzing hiss, close and satisfying. Audio: natural, realistic sound design with no music and no voiceover." || KEYFRAME, FLUX.1 [schnell] (exact; seed 101, 4 steps, CFG 1.0, 1920x1088): "Close-up of a real white-faced capuchin monkey: a neat black cap of fur on top of its head, a creamy white face, throat and shoulders, pinkish-tan bare facial skin, dark brown eyes, and a glossy black-brown body and arms holding an ice-cold hot-pink aluminum soda can with bold yellow letters "BANANA BOOM", beaded with condensation droplets in both of its small black-skinned monkey hands with slender dark fingers and dark fingernails, the can upright in front of its chest, one thumb hooked under the can's standard stay-on aluminum pull tab, the monkey's face looking down at the can, backlit by sunlight, jungle bokeh behind. Photorealistic wildlife commercial cinematography, shot on ARRI Alexa 35 with anamorphic lenses, natural warm sunlight, shallow depth of field, rich saturated tropical greens, crisp detail, film grain, high-end soda advertisement." || UPSCALE: SeedVR2 3B takes no prompt (seed 42), frames 24-100 of the Kling clip.
Frame at 14.33 s of version 9 (the featured hybrid). The cream torso here vs the black-armed monkey in the next shot is the seam the article explains: this shot came from the Veo/Kling comparison keyframes, made before the identity lock, and was reused without a re-take.
Halfway through opening a soda can, the monkey pulled a second pull tab out of thin air.
Google's Veo 3.1 had been asked, in those exact words, for "the tab stays attached." It delivered a tab that stayed attached — plus a spare, in case of emergencies. That shot cost 34 cents.
Over three days in October I made the same soda ad nine ways: no AI at all, free models on my local desktop, two paid models, and a mix.
The short version of what I learned has nothing to do with which model to buy. A small business can make a usable AI video ad for pocket change. The model matters less than the source material you feed it and how carefully you check what comes out. And the failures in the blooper reel had me laughing out loud.
# | Version | Cost to make | What broke |
|---|---|---|---|
1 | Pure code, no AI model | $0 | A robotic "ahhh" |
2 | 3D cartoon, free models | $0 in model fees | Every take that drank, leaked |
3 | Photoreal, free v2 | $0 in model fees | Three different monkeys |
4 | Veo 3.1 (Lite mode*) + our sound | $2.38 ($0.34 per 4-second shot) | A second pull tab |
5 | Veo 3.1 (Lite mode*), its own sound | Same $2.38 | The splash sounded wrong |
6 | Photoreal, free v3 | $0 in model fees | Nuzzled a can that was still closed |
7 | Photoreal, free v4 | $0 in model fees | Hands like rubber gloves |
8 | Kling 3.0, its own sound | $3.47 (~$0.50 per 5-second shot) | Near-silent on 3 of 7 shots; a stray musical tone |
9 ★ | Hybrid: #7 plus one Kling shot | ~$0.50 (one 5-second Kling shot) | A visible seam where the paid shot sits |
*Our tool sent no mode, so OpenArt's default applied. Today that default is Lite, and OpenArt keeps no per-clip record of the mode, so Lite is our best inference.
Why make a monkey open a can?
A short, 5-second clip portraying a monkey in sunglasses drinking a can of pop… a marketing-type commercial… enjoy the can with a lip-smacking 'ahhh'.
The soda is called BANANA BOOM and doesn't exist in the real world, despite the intrigue a banana-flavored soda brings to my taste buds. The brief sounds trivial, but it packs four things into five seconds that turned out to be exactly where things broke: hands, a tiny mechanism, liquid going where it should, and a crack that has to land on the exact frame.
The AI doing the work was Claude Code, Anthropic's coding agent: it wrote the scripts, ran every model, cut the edits and mixed the sound, and separate Claude agents reviewed each result. I wrote the brief, watched every cut and decided what shipped. When Claude got something wrong, I'll say so by name.
Level 0: no AI at all
Monkey #1 is drawn in code with Remotion, and every sound was synthesized from scratch. Twelve minutes to build, 38 seconds to render, $0. The "ahhh" was robotic, because math is a poor substitute for a satisfied primate.
For a logo sting or a mascot that must look identical every time, code still wins: the brand name never misspells itself.
Level 1: a 3D cartoon, free, on my desktop
The first real AI version ran entirely on my own desktop, with two free models doing the picture work, both licensed for commercial use:
Part | What I used | Its job | Cost |
|---|---|---|---|
The computer | My desktop PC (NVIDIA RTX 3080 graphics card, 10 GB; Ryzen 9 5900X; 64 GB RAM) | Ran everything locally | Already owned |
Picture model | FLUX.1 schnell | Drew each starting picture | $0 |
Video model | Wan 2.2 (Alibaba's open model) | Animated each picture, about 10 minutes per 5-second clip at 720p | $0 |
Then the monkey drank the soda, and the soda went down his chest.
So the next take's prompt said, in so many words, "no spilling and no dripping."
It still spilled.
Claude then asked for four short clips to bridge the gaps, each one told "completely dry: no liquid, no dripping, no spilling." All four dripped. A product shot with no monkey in it at all poured soda over a can that was still closed. The only dry take was the one where he didn't drink.
Two drinks, four bridges, six leaks. I've never seen an instruction ignored with such commitment.
We shipped it anyway, by cutting around the leaks. A lot of AI video editing is exactly that: rescuing the usable frames.
Level 2: photoreal, still free
Then the brief went realistic, with the monkey finding the can first, which is why everything from here runs 18 to 27 seconds.
#3, the first free photoreal cut, cast three different monkeys in one ad: one with a black cap and a peach face, one with coarse brown fur, one with a slate-blue face and a white ruff. The hand opening the can melted into a fingerless mitten.
On an empty jungle shot, Claude's prompt still mentioned the can, so Wan grew a giant, misspelled one in the forest, like a monument to the prompt. And in another take with no monkey, a human hand reached in and took the can. Somebody else filming on set wanted a soda too...
The fix had nothing to do with the video model. Claude rewrote the starting pictures to describe the same monkey, hands included, in every shot, and three species became one.
Then for #7, Qwen-Image-Edit built every picture with the monkey in it from one reference image, so he stayed the same monkey from shot to shot. Wan drafted each shot at 480p and SeedVR2 upscaled it to 1080p.
Same monkey in every shot, a finger that lifts the tab like a lever, a wide "AHHH." His hands look like rubber gloves, but model fees were still $0.
Lesson one: the starting pictures decide more than the video model does.
Level 3: what does $6 of paid AI video buy you?
Both paid models came through OpenArt, a site that resells credits for many models; its Plus plan is $34 a month for 12,000 credits. Claude gave both models the identical starting pictures and prompts, though not the same settings: Veo ran at 1080p in four-second shots, Kling in its standard 720p mode in five-second shots.
Veo 3.1 (from Google) cost $0.34 a shot, $2.38 for seven. It grew the spare tab and acted the "ahhh" without voicing it: mouth open, a lick of the lips, silence.
One caveat: Veo most likely ran in Lite mode (see the note under the first table), the lowest of OpenArt's three Veo 3.1 settings, so it probably didn't run at its best.
Kling 3.0 (from Kuaishou) cost about $0.50 a shot, $3.47 for seven. A finger hooks the ring, the tab rises, the spray comes out of the top, and the crack lands right on it.
Best sequence and sound of a can opening by far.
Then the monkey's palm turns magenta, as the can's color bleeds into his hand. Apparently the paint was still wet.
Two flaws they shared were ours.
The monkey changes coats between shots (jet black in one shot, a small brown juvenile in another, a cream mane in a third), and the can-opening shot has a grey hand with flat, human-looking fingernails. Both were already in the starting pictures, and both models faithfully animated what they were handed. Lesson one again, from the paid side.
Can the AI do the sound too?
Both paid models make sound along with the picture. Neither voiced the "ahhh," the one sound the brief asked for by name.
Kling's own audio was near-silent on three of seven shots, and in the shot that introduces the can, where Claude's prompt asked for stream water, a "glint" chime and "no music," Kling added a sustained musical tone. Interesting addition...
Veo's own splash I called "laughably bad," and part of the joke was on me: Claude's audio processing had squashed it.
The sound that won was the one we built: every effect placed on a measured frame. Our own "ahhh" came from an open voice model, Chatterbox, and a reviewer measured its pitch: an adult man's voice coming out of a small monkey.
So, as a tenth step, we swapped in real recordings from the free Sonniss sound library for the jungle, the birds and the "ahhh." More believable, still too low for a small monkey, and now I could hear the last fakes: two synthesized metal "tinks" when the monkey's hand grabs the can. A palm on an aluminum can doesn't ring like a bell.
Level 4: the winner was a hybrid
The version I'd stand behind, despite the remaining flaws, is #9: the free #7, with one shot swapped for Kling's can-opening and its own crack and fizz. The paid part is one five-second shot, about 50 cents. My note when I watched it: "the best one so far."
It also has the flaw I'm most embarrassed by.
Look at the image at the top of this page: the monkey opening the can has a cream torso. A second later, drinking, he has black arms. Same monkey, two wardrobes.
Lesson two: pay only for the hard shot. Free models were good enough for the scenery, the can on its rock and the drink. The paid model was clearly better at a hand working a tiny mechanism, so that's the shot worth paying for.
What it really cost, in money and time
Item | Cost |
|---|---|
Veo 3.1 (Lite mode*), seven 4-second 1080p shots | 840 credits = $2.38 |
Kling 3.0, seven 5-second 720p shots | 1,225 credits = $3.47 |
OpenArt Plus plan | $34 a month (12,000 credits) |
Free models and the Sonniss sound library | $0 in fees (electricity not metered) |
Claude, at published prices | About $120 |
Claude, what I actually paid | $0 extra on my $100-a-month plan |
Time | About ten hours of active work over three days |
*As above.
So all fourteen paid shots came to (drumroll please, though you've already seen it twice)
$5.85.
Measured from the session logs, the Claude work behind the nine versions came to about 1,200 model calls: the main session plus 24 review and research agents. At Anthropic's published API rates, that would have cost about $120.
I didn't pay that. It all ran on the $100-a-month Claude Max plan I already had, so the Claude side added $0 in cash.
That $120 also isn't the price of one ad. It covers all nine versions and a lot of trial and error, and along the way Claude wrote the scripts and we worked out a process the next video can reuse. So the next one should need less Claude work, though I haven't measured how much less.
An AI agent wrote the scripts and drove every step. By hand it would take longer, and I can't tell you how much because I am not a professional video editor.
Lesson three: check it like it's trying to fool you
From the 3D cartoon on, every file went to a reviewer before I signed off, and they caught things BEFORE they were shipped: Veo's second tab, which we cut down to just three frames in the final cut; audio running 52 milliseconds late; a "recorded" jungle sound that turned out to be synthesized; and Claude telling a reviewer a gap in the sound was covered when it wasn't.
We've written about this with text: making got cheap and checking didn't. With video it's worse, because a defect can hide in three frames. Step through the frames at full resolution before anyone else sees them.
Where AI video stands in October 2026
The free video model I used, Wan 2.2, released in July 2025, is fifteen months old, and Alibaba's Wan team hasn't put a newer general-purpose video model up for download since; its newer ones you can only rent. Every free-route gain here came from what we fed it and the tools around it, not from a newer model.
RigidBench, a University of Cambridge preprint from August 2026, tested eight models, including Veo 3.1 (its full and Fast tiers; ours most likely ran in Lite) and Kling 3.0 (in its pro mode, not our standard one). No model led on all ten measurements, and the usual "does it look right" score ranked them almost backwards from how accurately they moved objects. Looking right and moving right are different skills. My extra pull tab is a small, silly example.
What I didn't test: OpenAI's Sora, Runway, and the AI video features in Canva and CapCut. The scope was what a small business can run on its own desktop or rent one shot at a time, so I have nothing to add about those tools and how they'd do. Sounds like a test for another day.
If this were your ad, what would I use?
If you need… | Use | Rough cost (our numbers) |
|---|---|---|
A logo sting, a price card, a mascot that must match every time | Code or a motion template (Remotion is free for individuals and companies of up to three employees; larger ones pay) | $0 in model fees |
Scenery, b-roll, a product looking good | Free open models on a desktop with a decent graphics card | $0 in fees; ~10 min per 5-second clip at 720p |
The hero moment: hands, a small mechanism, liquid | Pay for that one shot, and budget a few attempts | ~$0.34–$0.50 per attempt through a reseller |
Sound | Real library recordings placed on the frame | $0 with a free library |
Your real label, a real person, a product claim | Film it | Most of our AI versions restyled or smeared our made-up label |
No desktop and no coding agent? Use a browser reseller like OpenArt and pay per shot.
#8 was all paid ($3.47), though its starting pictures came from my desktop; on the site you'd make them with its image models or use your own photos. Commercial use of what OpenArt makes needs its Plus plan or above.
Whichever route: make one picture of your character you like, build every other starting picture from it, keep each prompt to what's actually in the frame, and check the result frame by frame. If nobody on your team has time for that, the video fees aren't the real price.
If you'd rather not spend three days learning which shots to pay for, start a no-pressure conversation. We'll tell you which route fits what you're selling, including when the honest answer is code, or a camera.
Cheers, from the boring side of the business,

P.S. Claude has since made a free fixed cut. The tinks are gone (I picked silence over real hand sounds), the "ahhh" is pitched up to suit a small monkey (I picked it blind from ten takes), and a patch covers the end of the label smear. The patch left a faint seam of its own. Of course it did. Next time: before and after, side by side.
Want help applying this to your business? Start a no-pressure conversation →
Frequently asked questions
- Can a small business make a video ad with AI in 2026?
- Yes. In our October 2026 test, free open models on a desktop PC made most of a 23-second soda ad, and one paid 5-second shot from Kling 3.0 (about 50 cents on a $34-a-month plan) handled the hardest moment. Every version had visible flaws, so check it frame by frame before a customer sees it.
- How much does an AI video ad cost to make?
- Our fourteen paid shots cost $5.85 through OpenArt's $34-a-month plan: $0.34 per 4-second Veo 3.1 shot, about $0.50 per 5-second Kling 3.0 shot. The AI that directed and checked the work would have cost about $120 at Anthropic's published prices but ran on a $100-a-month subscription I already had. It took about ten hours of active work.
- Are free AI video models good enough for marketing?
- For most shots, yes. Wan 2.2 on a desktop with an RTX 3080, with Qwen-Image-Edit keeping one consistent character and SeedVR2 upscaling to 1080p, was good enough for scenery, product shots and the drink. It lost on the hardest shot, a hand opening a can.
- Which is better for ads, Veo 3.1 or Kling 3.0?
- We can't call it cleanly: Veo ran at 1080p in what was most likely its Lite mode, Kling at 720p in standard mode. On identical starting pictures, Veo grew a second pull tab and Kling gave the best can-opening of the project, then turned the monkey's palm magenta. Neither voiced the "ahhh".
- Why do AI video models struggle with hands and small objects?
- Our test shows that they do, not why. Researchers find the same gap between looking right and moving right: a University of Cambridge preprint from August 2026 found the usual "does it look right" score ranked models almost backwards from how accurately they moved objects. Neither it nor Google DeepMind's earlier Physics-IQ tested hands.
Know someone who needs this?



