Two papers, one familiar question
I keep having this strange experience lately: I spend time experimenting with something in a very small, practical way, and then a paper appears exploring a remarkably similar question at a much larger scale.
Two recent papers made that happen again.
The first is “Evaluating Skills, Not Just Agents”, which introduces ACES and Skill Lift: execute the same task with and without a skill and measure the difference.
That connects directly with what I was exploring in skill-eval.
But my main takeaway from that experiment was actually about the limits of A/B testing.
An A/B comparison gives evidence for that particular model, task, context and harness. It does not prove that a skill is universally good. Change the model and the skill may become less useful, unnecessary, or even harmful.
Knowledge does not have to become instructions
Then I read Google Research's “WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution”.
And another piece clicked into place.
WikiSkill separates raw experience, accumulated knowledge and the actual executable skill. Experience can keep accumulating outside the skill, while proposed skill changes are validated before being promoted.
Even more interestingly, the paper shows cases where low-level workarounds learned for one model transfer poorly to a stronger model. Instructions that helped the weaker model could constrain the stronger one or cause redundant tool calls.
That is very close to another experiment I did inside skill-eval.
Evolving a skill by making it smaller
I tried to “evolve” my implement-execplan skill not by adding more instructions, but by removing them.
The original baseline had 228 lines. The adaptive version had 165.
Instead of requiring a large fixed plan structure and continuous updates, the newer version relies more on the model's own capabilities: fewer mandatory sections, conditional information only when useful, fewer plan updates and less routine bookkeeping.
The skill is here:
And the case study is here:
skill-eval: implement-execplan case study
In the first valid comparison, both versions passed the same quality gates, while the smaller version used about 18% fewer candidate input + output tokens.
That is only one observation, not proof that the new version is universally better. The replication was incomplete, and I deliberately documented that limitation.
But it changed how I think about “skill evolution”.
Evolution does not necessarily mean accumulating more rules.
Sometimes the model improved and part of the harness became obsolete.
Sometimes something that used to require an instruction can become a deterministic test, hook or static check and disappear from the skill entirely.
This is also a continuation of something I noticed while gradually giving TDD to an agent. In When I Entrusted TDD to an AI Agent, some architectural guidance became much more reliable once it stopped being only an instruction and became executable feedback in the repository.
And sometimes a failure is worth remembering, but not yet worth turning into an instruction. This is where the WikiSkill separation between experience, persistent knowledge and executable skills makes a lot of sense to me.
I prefer to discover the workflow first
This also changed how I create skills.
I prefer to experience the problem first, work with the model manually, adjust the workflow until I can make it work reliably, and only then consolidate what I learned into a skill.
My latest implement-approved-plan skill came from exactly that process:
implement-approved-plan/SKILL.md
This one is interesting because the skill itself coordinates a multi-chat workflow inside ChatGPT Desktop.
A dedicated Coordinator keeps the approved plan and controls a set of specialized chats. Each role is created once and then reused throughout the implementation: an Implementer works through TDD, a Committer owns commits, a Validator runs the complete quality gates, a Habit Curator handles mechanically detected design findings, and an independent Structural Reviewer looks for design problems after the implementation is already green.
The specialization is not only about giving each chat a different responsibility. The skill also assigns different models to different roles. The Coordinator, Implementer, Habit Curator, and Structural Reviewer run on gpt-5.6-sol at xhigh; the Committer runs on gpt-5.6-terra at xhigh; and the Validator runs on gpt-5.6-luna at xhigh.
That routing makes the workflow more economical. Sol is reserved for open-ended implementation, orchestration, and design judgment; Terra handles commit work, which is narrower but still requires judgment about repository conventions; and Luna runs the mostly deterministic validation gates. The point is not merely to create specialists, but to avoid using the flagship model where a more economical model is sufficient.
implement-approved-plan skill to implement habit-hooks issue #160 in the amazing habit-hooks project.Conceptually, it looks roughly like this:
Approved Plan
↓
Coordinator
↓
Implementer
↓
Coordinator
↓
Committer
↓
Coordinator
↓
Validator
↓
Coordinator
↓
Habit Curator
↓
Coordinator
↓
Structural Reviewer
↓
Final Validation
↓
PR + CI
↓
Ready for Merge
The interesting part for me is that these chats do not freely coordinate with each other. The Coordinator is the only one talking to the specialists, and access to the same working tree is serialized through a persistent workflow ledger.
That ledger works almost like a small state machine. It records the current phase, specialist chat IDs, commits, validation gates, Habit state, pull request and CI evidence. It rejects invalid transitions, concurrent ownership of the checkout, premature pull requests and cleanup before GitHub actually confirms that the PR was merged.
The shape is more elaborate than the three nested loops I described in I Had Already Built Three Agentic Loops Without Naming Them, but the underlying idea feels similar: autonomy becomes safer when feedback and exit conditions live in the workflow instead of depending on the model remembering everything.
So the skill is not trying to make one giant agent smarter. It is encoding a workflow I had already been performing manually, while moving as much coordination and validation as possible into deterministic machinery around the model.
PR that introduced it:
And instead of stopping when the SKILL.md looked good, I exercised the workflow against a public fixture and a real pull request:
implement-approved-plan-fixture
implement-approved-plan-fixture#2
My current mental model
So my current mental model is becoming something like this:
Experience the problem.
Make the workflow work manually.
Turn deterministic knowledge into deterministic mechanisms.
Keep uncertain experience outside the skill until it earns its place.
Add instructions only for what the model still needs.
And when the model changes, question those instructions again.
No magic skill. No assumption that more instructions mean a better agent.
Just a workflow that keeps adapting as both the model and my understanding of the problem change.

Top comments (19)
Moving structural constraints out of the prompt and into a state machine or git hook is usually where agent setups stop regressing. When instructions pile up in a SKILL.md, newer frontier models end up over-indexing on edge-case scaffolding meant for older ones, often adding defensive tool calls where a direct patch would work.The workflow ledger you describe for serializing working tree access solves the exact race condition that usually breaks multi-agent setups. Once verification and checkout locking live in the harness, you can prune the prompt down to the actual task delta.
Thanks! Same experience here. I am digging into the AI Agent building process to understanding more and more.
This is a useful way to think about skill maintenance. Instructions are not automatically more durable because there are more of them. As model behavior improves, keeping only the constraints, acceptance checks, and procedures the model cannot safely infer can make the skill both clearer and more robust.
What I have in mind now, actually, is figuring out how I can do things without needing a skill.
So I try to find deterministic ways to handle them. When I see that I cannot make something deterministic, or I still have not found a solid and concrete approach that is worth implementing, I start by experimenting with skills.
At that stage, I use the skill for the parts that still depend on behavior and judgment. Then, as I use it and patterns start to emerge, I can identify opportunities to turn those patterns into deterministic mechanisms. As that happens, the skill gradually becomes smaller.
Sometimes, the skill itself also contains deterministic scripts. So some parts of the skill still rely on nondeterministic behavior, while other parts are handled deterministically by scripts inside the skill.
That makes sense. I like that approach a lot. Use the skill while the pattern is still messy, then make it smaller as the repeatable parts become clear. Eventually the skill only needs to handle the parts that really still need judgment.
The distinction you draw between accumulating instructions and moving knowledge into deterministic mechanisms gets right to the heart of why agent setups degrade over time.
When developers treat a
SKILL.mdas an append-only log of past post-mortems, the skill becomes brittle scaffolding. Frontier models then waste tokens reasoning through defensive edge-case guardrails designed for weaknesses in earlier models—often triggering speculative tool hops instead of direct patches.In our production harness, we found three concrete patterns essential when migrating rules out of prompts and into runtime mechanics:
Precondition Tokens over Imperative Guidelines: Instead of instructing the agent to "always verify prerequisites before mutating", the tool layer enforces a two-phase dry-run/intent attestation. The mutation tool rejects any execution that lacks a freshly generated content-addressed intent token (
intent_sha256). This eliminates an entire class of prompt instructions and makes accidental double-writes physically impossible in code.Decoupling Cognitive Intent from State Receipts: When coordinating specialist roles (Coordinator vs. Implementer vs. Validator), relying on the model to "remember where it left off" creates silent drift. Serializing working tree access through an external ledger—where state transitions (
planned -> submitted -> verifying -> merged) are enforced by the harness—frees the model from routine operational bookkeeping. The prompt only ever receives the current delta and verified readback.Treating Reversion as First-Class Evolution: Just as you observed with
implement-execplanshrinking from 228 to 165 lines, skill maintenance needs a deprecation cycle. If a compiler check, AST fuzzer, or git hook can catch a regression deterministically, the corresponding prompt paragraph should be deleted immediately.The coordinator-specialist chat routing you mapped out is a great illustration of this: real reliability doesn't come from a mega-prompt that tries to anticipate every failure mode, but from narrow execution scopes bounded by rigid external state machines.
I am loving the coordenador workflow. When I have second thoughts about what is implementing, I talk to the coodinator and it will handle the new instructions immediately and orchestrated everything for me.
I have to much to learn.
We read the current
implement-execplan/SKILL.mdbefore writing this. It is 165 lines today, so the artifact matches the post, and reading it changes what the 18 percent figure probably means.The removal is concentrated in one place. Plan structure went adaptive: five mandatory core sections, everything else conditional, plus an explicit instruction to prefer the smallest plan that stays safe and handoff-ready. But the validation and handoff section at the end is still unconditional and dense, a re-read of the request against the final workspace, a check that documentation names the exact legacy and current identifiers, a summary with five required bullets. So the skill did not get uniformly smaller. It got smaller where it dictates the artifact, and stayed the same where it dictates the audit.
Which suggests the token saving may not be coming from where the framing implies. Sixty-three fewer lines of SKILL.md is a fixed cost paid once per run, and on a 165-line file that is a few hundred tokens. A plan with fewer mandatory sections and a lower update cadence is a variable cost paid on every milestone. If most of the 18 percent is the second thing, then the lesson is not "smaller skills use fewer tokens", it is "skills that mandate smaller artifacts use fewer tokens", and those two evolve in different directions. The second one gets better by being more prescriptive about what to leave out, which means writing more instruction, not less.
Do your case study numbers let you separate those? If the plan artifact's own size dropped roughly in line with the total, that settles it, and it would make the WikiSkill transfer result land harder: an instruction that shrinks output is worth keeping across a model upgrade, while an instruction that scaffolds a weaker model's reasoning is the one that goes stale.
That mid-flight steering is where a dedicated coordinator really shines.
The architectural reason it works so well is boundary isolation: when you change requirements mid-task, the coordinator can halt child workers and ensure the working tree or intermediate state is reconciled before fanning out new tasks. Without a coordinator layer, mid-flight pivots usually create split-brain states where specialists continue executing on stale assumptions.
Decoupling high-level intent orchestration from low-level execution leaves room to explore and iterate safely without leaving messy partial state behind.
That was a really good and relevant comment. As I explained, I ended up building this workflow quite intuitively. I was basically copying and pasting between one chat and another, and before I knew it, the workflow had taken shape and eventually became a skill.
But the split brain issue you mentioned is a really important point. The agents can end up operating somewhat independently, almost as if they are no longer properly coordinated.
I really liked your comment. Thank you!
Very interesting. I have been seeing the same pattern while developing CCA: newer models sometimes skip or reinterpret instructions that older models followed more literally. That makes a long SKILL.md increasingly fragile as an enforcement mechanism.
The distinction between experience, knowledge, and executable skill is useful. My current conclusion is that skills should contain only guidance the model still needs. Invariants should move into hooks, deterministic checks, and state transitions.
I also strongly agree that skill evolution can mean deletion. When a rule becomes executable, removing it from the prompt reduces token cost and avoids carrying workarounds designed for weaker models into stronger ones. This feels like the right direction for making agent workflows resilient to model changes.
Evolving a skill by deleting it smaller is the direction that feels wrong and keeps being right. Instructions are hypotheses about the model, and they expire when the model changes - the workaround that rescued a weak model becomes a constraint or a redundant tool call on a stronger one. The A/B caveat is the load-bearing line for me: lift measured on one model, task, and harness is evidence for that combination, not a property of the skill. "Change the model and the skill may become harmful" should be taped above every prompt library.
I totally agree! The repository skills in github are a great source of knowledge, but I won't use it blindly 😃
the A/B limit you're naming is real — "this skill helped on this task with this model" doesn't tell you whether the skill is doing the work or the model already knew how to do it. ACES tries to isolate that, but only if you're willing to run the same task without the skill many times on the same model, which gets expensive fast.
the removal direction is interesting because it suggests a different fitness signal: if the model handles a class of task without the instruction, the instruction was just scaffolding. you're measuring independence, not just performance.
what triggered the first removal? was it a test that just kept passing?
I changed the model from gpt 5.5 to gpt 5.6 sol xhigh, and i realized that the detailed execution plan created a lot of steps and large files. So I decided to test how sol could follow the instructions with fewer harnesses. That was the starting point.
The idea of removing obsolete instructions is something I need to apply to my own eval sets. Stale rubrics are as dangerous as stale prompts and twice as hard to notice.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.