DEV Community

Cover image for I Gave 4 AI Models the Same Agent Skill. Here's What Happened
Prakhar Yadav
Prakhar Yadav

Posted on Originally published at prakhar.hashnode.dev

I Gave 4 AI Models the Same Agent Skill. Here's What Happened

I've been experimenting with video generation workflows lately, specifically with how different AI models behave when they're given the same tools, environment, and instructions.

So I decided to run a small experiment.

I created a reusable video-generation skill — a set of detailed instructions describing how the task should be approached — and gave the same skill to four different models through OpenCode.

The results were interesting.

The models produced broadly comparable outputs, but their approaches to the task were noticeably different.


The Setup

The video-generation skill was created with the help of Opus 5.5.

It contained the workflow, instructions, verification steps, and expectations for generating the video.

I then gave the same skill to:

  • GLM 5.3 Flash
  • DeepSeek v4.1 Flash
  • MiMo v2.6 Flash
  • LongCat 2.5 Preview

Each model had access to the same general environment and tools.

The goal wasn't to create a formal benchmark or determine which model is "best."

I wanted to see something simpler:

How differently do models behave when they're given the same well-defined workflow?


The Results

Here are the execution times and token usage from each run:

Model Time Tokens
GLM 5.3 Flash 28 min 200.1K
DeepSeek v4.1 Flash 29 min 265.6K
MiMo v2.6 Flash 47 min 266.3K
LongCat 2.5 Preview 1 hr 5 min 276.8K

A few things immediately stood out.

GLM completed the task in 28 minutes, while DeepSeek was just a minute behind at 29 minutes.

GLM also used considerably fewer tokens — about 200K, compared with roughly 266–277K for the other three.

MiMo took 47 minutes, despite using almost the same number of tokens as DeepSeek.

LongCat was the slowest at 1 hour 5 minutes, and also used the most tokens.

That made sense once I looked at how each model approached the task.


1. GLM 5.3 Flash — Fast and Focused 🏎️

GLM understood the assignment and got straight to work with what it had been given.

Its instruction following was particularly strong.

There was very little deviation from the predefined workflow, and it didn't spend much time exploring unnecessary paths.

It simply followed the instructions and executed.

The result was a relatively efficient run:

28 minutes · 200.1K tokens

If I had to describe its behavior in one word, it would be:

Focused.

It didn't try to reinvent the workflow.

It just executed it.


2. DeepSeek v4.1 Flash — The Most Exploratory 🎨

DeepSeek understood the assignment well, but its approach was quite different.

Instead of immediately executing the provided workflow, it spent more time understanding the environment.

It went digging.

It launched Playwright MCP and other tools to inspect how things actually looked in the environment rather than relying only on the information already provided.

That extra exploration made it the most interesting model from a creativity perspective.

It wasn't necessarily looking for the shortest path.

It was trying to build a better understanding of the context before producing the result.

The run took:

29 minutes · 265.6K tokens

So despite using roughly 65K more tokens than GLM, it completed the task almost exactly as quickly.


3. MiMo v2.6 Flash — The Thinker 🧘‍♂️

MiMo spent considerably more time understanding the task and the environment before getting started.

It was more deliberate.

It explored the context, worked through the instructions, and then executed the workflow.

The interesting part was that the final output was still broadly comparable to the others.

The run took:

47 minutes · 266.3K tokens

And this highlighted something important for me:

The quality wasn't necessarily coming from the model figuring everything out from scratch.

The skill itself was doing a lot of the heavy lifting.

Because the instructions were already clear and structured, MiMo had a roadmap to follow.

It could spend more effort understanding the context without having to invent the entire workflow.


4. LongCat 2.5 Preview — The Most Iterative 😼

LongCat took the longest of the four.

1 hour 5 minutes · 276.8K tokens

Like MiMo, it spent considerable time understanding the context and environment.

But there was one thing that stood out.

It performed two verification rounds.

After generating the initial result, it checked its work, identified areas that could be improved, and then performed another iteration.

That made its workflow particularly interesting:

Understand → Generate → Verify → Improve → Verify

The final quality was broadly on par with the other models, but its process was more iterative.

It wasn't just generating the result and stopping.

It was actively trying to improve it.


The Interesting Part: They All Verified Their Work

One of the most interesting observations from this experiment was that every model performed at least one verification round at the end.

These weren't simply:

Prompt → Generate → Done

workflows.

The models actually checked their output.

LongCat went one step further and performed another iteration specifically to improve the result.

So despite the differences in their approaches, they all converged on a similar pattern:

Understand → Execute → Verify

Some models explored more.

Some were faster.

Some spent more tokens.

Some iterated more.

But verification was present across all four runs.


Was This Really "One-Shot"?

Technically, each video was generated in a single run.

But there's an important caveat.

The models weren't starting from a blank page.

They were given a pre-built skill containing clear instructions and a defined workflow.

So calling this a pure "one-shot" experiment would be misleading.

The actual process was closer to:

One prompt + a predefined workflow + autonomous execution

And that's arguably more interesting.

The experiment wasn't really testing:

Which model can figure out video generation from nothing?

It was testing:

How differently do models behave when they're given the same structured workflow?


The Skill May Have Mattered More Than the Model

This was probably my biggest takeaway.

The final outputs were surprisingly close in quality despite the models taking very different approaches.

The predefined skill gave every model a strong starting point.

Instead of having to figure out:

  • Which tools to use
  • What steps to follow
  • How to inspect the environment
  • How to validate the output
  • What constitutes a good result

the models already had a roadmap.

That changes the role of the model.

Instead of asking the model to invent the entire workflow, you're asking it to execute and adapt a workflow.

And that seems to reduce the differences between models for this particular task.


Model Behavior at a Glance

Model Time Tokens What stood out
GLM 5.3 Flash 🏎️ 28 min 200.1K Fast execution, strong instruction following
DeepSeek v4.1 Flash 🎨 29 min 265.6K Environment exploration and creativity
MiMo v2.6 Flash 🧘‍♂️ 47 min 266.3K Deliberate reasoning and context gathering
LongCat 2.5 Preview 😼 65 min 276.8K Verification and additional iteration

These numbers come from a single task/run per model, so I wouldn't treat them as a general benchmark of model speed, efficiency, or capability.

They're simply measurements of how these models behaved in this particular workflow.


What I Took Away From This

We've spent a lot of time comparing models based on raw capabilities.

Which model is smarter?

Which model follows instructions better?

Which model writes better code?

But agentic workflows introduce another variable:

How good is the workflow you're giving the model?

A good skill can provide:

  • Context
  • Tool usage instructions
  • A repeatable workflow
  • Validation steps
  • Quality criteria
  • Recovery strategies

And suddenly the model doesn't need to figure everything out itself.

It needs to execute the process well.

That makes the combination of:

Model + Skill + Tools + Verification

potentially much more important than looking at the model alone.


Final Thoughts

This was a small experiment, not a benchmark.

Four models.

One task.

One predefined skill.

But the differences in how they approached the same problem were fascinating.

GLM was fast and focused.

DeepSeek explored and experimented.

MiMo was deliberate.

LongCat iterated.

And despite those differences, they all managed to produce broadly comparable results.

For me, that's the interesting part.

Maybe the future of agentic AI isn't just about finding a smarter model.

Maybe it's also about getting better at telling the model how to work.


TL;DR

🏎️ GLM 5.3 Flash — 28 min / 200.1K tokens
🎨 DeepSeek v4.1 Flash — 29 min / 265.6K tokens
🧘‍♂️ MiMo v2.6 Flash — 47 min / 266.3K tokens
😼 LongCat 2.5 Preview — 65 min / 276.8K tokens

All four performed at least one verification round.

The bigger takeaway wasn't which model was "best."

It was how much a well-designed skill can shape the behavior of very different models.

Top comments (0)