I've been experimenting with video generation workflows lately, specifically with how different AI models behave when they're given the same tools, environment, and instructions.
So I decided to run a small experiment.
I created a reusable video-generation skill — a set of detailed instructions describing how the task should be approached — and gave the same skill to four different models through OpenCode.
The results were interesting.
The models produced broadly comparable outputs, but their approaches to the task were noticeably different.
The Setup
The video-generation skill was created with the help of Opus 5.5.
It contained the workflow, instructions, verification steps, and expectations for generating the video.
I then gave the same skill to:
- GLM 5.3 Flash
- DeepSeek v4.1 Flash
- MiMo v2.6 Flash
- LongCat 2.5 Preview
Each model had access to the same general environment and tools.
The goal wasn't to create a formal benchmark or determine which model is "best."
I wanted to see something simpler:
How differently do models behave when they're given the same well-defined workflow?
The Results
Here are the execution times and token usage from each run:
| Model | Time | Tokens |
|---|---|---|
| GLM 5.3 Flash | 28 min | 200.1K |
| DeepSeek v4.1 Flash | 29 min | 265.6K |
| MiMo v2.6 Flash | 47 min | 266.3K |
| LongCat 2.5 Preview | 1 hr 5 min | 276.8K |
A few things immediately stood out.
GLM completed the task in 28 minutes, while DeepSeek was just a minute behind at 29 minutes.
GLM also used considerably fewer tokens — about 200K, compared with roughly 266–277K for the other three.
MiMo took 47 minutes, despite using almost the same number of tokens as DeepSeek.
LongCat was the slowest at 1 hour 5 minutes, and also used the most tokens.
That made sense once I looked at how each model approached the task.
1. GLM 5.3 Flash — Fast and Focused 🏎️
GLM understood the assignment and got straight to work with what it had been given.
Its instruction following was particularly strong.
There was very little deviation from the predefined workflow, and it didn't spend much time exploring unnecessary paths.
It simply followed the instructions and executed.
The result was a relatively efficient run:
28 minutes · 200.1K tokens
If I had to describe its behavior in one word, it would be:
Focused.
It didn't try to reinvent the workflow.
It just executed it.
2. DeepSeek v4.1 Flash — The Most Exploratory 🎨
DeepSeek understood the assignment well, but its approach was quite different.
Instead of immediately executing the provided workflow, it spent more time understanding the environment.
It went digging.
It launched Playwright MCP and other tools to inspect how things actually looked in the environment rather than relying only on the information already provided.
That extra exploration made it the most interesting model from a creativity perspective.
It wasn't necessarily looking for the shortest path.
It was trying to build a better understanding of the context before producing the result.
The run took:
29 minutes · 265.6K tokens
So despite using roughly 65K more tokens than GLM, it completed the task almost exactly as quickly.
3. MiMo v2.6 Flash — The Thinker 🧘♂️
MiMo spent considerably more time understanding the task and the environment before getting started.
It was more deliberate.
It explored the context, worked through the instructions, and then executed the workflow.
The interesting part was that the final output was still broadly comparable to the others.
The run took:
47 minutes · 266.3K tokens
And this highlighted something important for me:
The quality wasn't necessarily coming from the model figuring everything out from scratch.
The skill itself was doing a lot of the heavy lifting.
Because the instructions were already clear and structured, MiMo had a roadmap to follow.
It could spend more effort understanding the context without having to invent the entire workflow.
4. LongCat 2.5 Preview — The Most Iterative 😼
LongCat took the longest of the four.
1 hour 5 minutes · 276.8K tokens
Like MiMo, it spent considerable time understanding the context and environment.
But there was one thing that stood out.
It performed two verification rounds.
After generating the initial result, it checked its work, identified areas that could be improved, and then performed another iteration.
That made its workflow particularly interesting:
Understand → Generate → Verify → Improve → Verify
The final quality was broadly on par with the other models, but its process was more iterative.
It wasn't just generating the result and stopping.
It was actively trying to improve it.
The Interesting Part: They All Verified Their Work
One of the most interesting observations from this experiment was that every model performed at least one verification round at the end.
These weren't simply:
Prompt → Generate → Done
workflows.
The models actually checked their output.
LongCat went one step further and performed another iteration specifically to improve the result.
So despite the differences in their approaches, they all converged on a similar pattern:
Understand → Execute → Verify
Some models explored more.
Some were faster.
Some spent more tokens.
Some iterated more.
But verification was present across all four runs.
Was This Really "One-Shot"?
Technically, each video was generated in a single run.
But there's an important caveat.
The models weren't starting from a blank page.
They were given a pre-built skill containing clear instructions and a defined workflow.
So calling this a pure "one-shot" experiment would be misleading.
The actual process was closer to:
One prompt + a predefined workflow + autonomous execution
And that's arguably more interesting.
The experiment wasn't really testing:
Which model can figure out video generation from nothing?
It was testing:
How differently do models behave when they're given the same structured workflow?
The Skill May Have Mattered More Than the Model
This was probably my biggest takeaway.
The final outputs were surprisingly close in quality despite the models taking very different approaches.
The predefined skill gave every model a strong starting point.
Instead of having to figure out:
- Which tools to use
- What steps to follow
- How to inspect the environment
- How to validate the output
- What constitutes a good result
the models already had a roadmap.
That changes the role of the model.
Instead of asking the model to invent the entire workflow, you're asking it to execute and adapt a workflow.
And that seems to reduce the differences between models for this particular task.
Model Behavior at a Glance
| Model | Time | Tokens | What stood out |
|---|---|---|---|
| GLM 5.3 Flash 🏎️ | 28 min | 200.1K | Fast execution, strong instruction following |
| DeepSeek v4.1 Flash 🎨 | 29 min | 265.6K | Environment exploration and creativity |
| MiMo v2.6 Flash 🧘♂️ | 47 min | 266.3K | Deliberate reasoning and context gathering |
| LongCat 2.5 Preview 😼 | 65 min | 276.8K | Verification and additional iteration |
These numbers come from a single task/run per model, so I wouldn't treat them as a general benchmark of model speed, efficiency, or capability.
They're simply measurements of how these models behaved in this particular workflow.
What I Took Away From This
We've spent a lot of time comparing models based on raw capabilities.
Which model is smarter?
Which model follows instructions better?
Which model writes better code?
But agentic workflows introduce another variable:
How good is the workflow you're giving the model?
A good skill can provide:
- Context
- Tool usage instructions
- A repeatable workflow
- Validation steps
- Quality criteria
- Recovery strategies
And suddenly the model doesn't need to figure everything out itself.
It needs to execute the process well.
That makes the combination of:
Model + Skill + Tools + Verification
potentially much more important than looking at the model alone.
Final Thoughts
This was a small experiment, not a benchmark.
Four models.
One task.
One predefined skill.
But the differences in how they approached the same problem were fascinating.
GLM was fast and focused.
DeepSeek explored and experimented.
MiMo was deliberate.
LongCat iterated.
And despite those differences, they all managed to produce broadly comparable results.
For me, that's the interesting part.
Maybe the future of agentic AI isn't just about finding a smarter model.
Maybe it's also about getting better at telling the model how to work.
TL;DR
🏎️ GLM 5.3 Flash — 28 min / 200.1K tokens
🎨 DeepSeek v4.1 Flash — 29 min / 265.6K tokens
🧘♂️ MiMo v2.6 Flash — 47 min / 266.3K tokens
😼 LongCat 2.5 Preview — 65 min / 276.8K tokens
All four performed at least one verification round.
The bigger takeaway wasn't which model was "best."
It was how much a well-designed skill can shape the behavior of very different models.
Top comments (0)