Two papers, one familiar question
I keep having this strange experience lately: I spend time experimenting with something in a very smal...
For further actions, you may consider blocking this person and/or reporting abuse
Moving structural constraints out of the prompt and into a state machine or git hook is usually where agent setups stop regressing. When instructions pile up in a SKILL.md, newer frontier models end up over-indexing on edge-case scaffolding meant for older ones, often adding defensive tool calls where a direct patch would work.The workflow ledger you describe for serializing working tree access solves the exact race condition that usually breaks multi-agent setups. Once verification and checkout locking live in the harness, you can prune the prompt down to the actual task delta.
Thanks! Same experience here. I am digging into the AI Agent building process to understanding more and more.
This is a useful way to think about skill maintenance. Instructions are not automatically more durable because there are more of them. As model behavior improves, keeping only the constraints, acceptance checks, and procedures the model cannot safely infer can make the skill both clearer and more robust.
What I have in mind now, actually, is figuring out how I can do things without needing a skill.
So I try to find deterministic ways to handle them. When I see that I cannot make something deterministic, or I still have not found a solid and concrete approach that is worth implementing, I start by experimenting with skills.
At that stage, I use the skill for the parts that still depend on behavior and judgment. Then, as I use it and patterns start to emerge, I can identify opportunities to turn those patterns into deterministic mechanisms. As that happens, the skill gradually becomes smaller.
Sometimes, the skill itself also contains deterministic scripts. So some parts of the skill still rely on nondeterministic behavior, while other parts are handled deterministically by scripts inside the skill.
That makes sense. I like that approach a lot. Use the skill while the pattern is still messy, then make it smaller as the repeatable parts become clear. Eventually the skill only needs to handle the parts that really still need judgment.
The distinction you draw between accumulating instructions and moving knowledge into deterministic mechanisms gets right to the heart of why agent setups degrade over time.
When developers treat a
SKILL.mdas an append-only log of past post-mortems, the skill becomes brittle scaffolding. Frontier models then waste tokens reasoning through defensive edge-case guardrails designed for weaknesses in earlier models—often triggering speculative tool hops instead of direct patches.In our production harness, we found three concrete patterns essential when migrating rules out of prompts and into runtime mechanics:
Precondition Tokens over Imperative Guidelines: Instead of instructing the agent to "always verify prerequisites before mutating", the tool layer enforces a two-phase dry-run/intent attestation. The mutation tool rejects any execution that lacks a freshly generated content-addressed intent token (
intent_sha256). This eliminates an entire class of prompt instructions and makes accidental double-writes physically impossible in code.Decoupling Cognitive Intent from State Receipts: When coordinating specialist roles (Coordinator vs. Implementer vs. Validator), relying on the model to "remember where it left off" creates silent drift. Serializing working tree access through an external ledger—where state transitions (
planned -> submitted -> verifying -> merged) are enforced by the harness—frees the model from routine operational bookkeeping. The prompt only ever receives the current delta and verified readback.Treating Reversion as First-Class Evolution: Just as you observed with
implement-execplanshrinking from 228 to 165 lines, skill maintenance needs a deprecation cycle. If a compiler check, AST fuzzer, or git hook can catch a regression deterministically, the corresponding prompt paragraph should be deleted immediately.The coordinator-specialist chat routing you mapped out is a great illustration of this: real reliability doesn't come from a mega-prompt that tries to anticipate every failure mode, but from narrow execution scopes bounded by rigid external state machines.
I am loving the coordenador workflow. When I have second thoughts about what is implementing, I talk to the coodinator and it will handle the new instructions immediately and orchestrated everything for me.
I have to much to learn.
We read the current
implement-execplan/SKILL.mdbefore writing this. It is 165 lines today, so the artifact matches the post, and reading it changes what the 18 percent figure probably means.The removal is concentrated in one place. Plan structure went adaptive: five mandatory core sections, everything else conditional, plus an explicit instruction to prefer the smallest plan that stays safe and handoff-ready. But the validation and handoff section at the end is still unconditional and dense, a re-read of the request against the final workspace, a check that documentation names the exact legacy and current identifiers, a summary with five required bullets. So the skill did not get uniformly smaller. It got smaller where it dictates the artifact, and stayed the same where it dictates the audit.
Which suggests the token saving may not be coming from where the framing implies. Sixty-three fewer lines of SKILL.md is a fixed cost paid once per run, and on a 165-line file that is a few hundred tokens. A plan with fewer mandatory sections and a lower update cadence is a variable cost paid on every milestone. If most of the 18 percent is the second thing, then the lesson is not "smaller skills use fewer tokens", it is "skills that mandate smaller artifacts use fewer tokens", and those two evolve in different directions. The second one gets better by being more prescriptive about what to leave out, which means writing more instruction, not less.
Do your case study numbers let you separate those? If the plan artifact's own size dropped roughly in line with the total, that settles it, and it would make the WikiSkill transfer result land harder: an instruction that shrinks output is worth keeping across a model upgrade, while an instruction that scaffolds a weaker model's reasoning is the one that goes stale.
That mid-flight steering is where a dedicated coordinator really shines.
The architectural reason it works so well is boundary isolation: when you change requirements mid-task, the coordinator can halt child workers and ensure the working tree or intermediate state is reconciled before fanning out new tasks. Without a coordinator layer, mid-flight pivots usually create split-brain states where specialists continue executing on stale assumptions.
Decoupling high-level intent orchestration from low-level execution leaves room to explore and iterate safely without leaving messy partial state behind.
That was a really good and relevant comment. As I explained, I ended up building this workflow quite intuitively. I was basically copying and pasting between one chat and another, and before I knew it, the workflow had taken shape and eventually became a skill.
But the split brain issue you mentioned is a really important point. The agents can end up operating somewhat independently, almost as if they are no longer properly coordinated.
I really liked your comment. Thank you!
Very interesting. I have been seeing the same pattern while developing CCA: newer models sometimes skip or reinterpret instructions that older models followed more literally. That makes a long SKILL.md increasingly fragile as an enforcement mechanism.
The distinction between experience, knowledge, and executable skill is useful. My current conclusion is that skills should contain only guidance the model still needs. Invariants should move into hooks, deterministic checks, and state transitions.
I also strongly agree that skill evolution can mean deletion. When a rule becomes executable, removing it from the prompt reduces token cost and avoids carrying workarounds designed for weaker models into stronger ones. This feels like the right direction for making agent workflows resilient to model changes.
Evolving a skill by deleting it smaller is the direction that feels wrong and keeps being right. Instructions are hypotheses about the model, and they expire when the model changes - the workaround that rescued a weak model becomes a constraint or a redundant tool call on a stronger one. The A/B caveat is the load-bearing line for me: lift measured on one model, task, and harness is evidence for that combination, not a property of the skill. "Change the model and the skill may become harmful" should be taped above every prompt library.
I totally agree! The repository skills in github are a great source of knowledge, but I won't use it blindly 😃
the A/B limit you're naming is real — "this skill helped on this task with this model" doesn't tell you whether the skill is doing the work or the model already knew how to do it. ACES tries to isolate that, but only if you're willing to run the same task without the skill many times on the same model, which gets expensive fast.
the removal direction is interesting because it suggests a different fitness signal: if the model handles a class of task without the instruction, the instruction was just scaffolding. you're measuring independence, not just performance.
what triggered the first removal? was it a test that just kept passing?
I changed the model from gpt 5.5 to gpt 5.6 sol xhigh, and i realized that the detailed execution plan created a lot of steps and large files. So I decided to test how sol could follow the instructions with fewer harnesses. That was the starting point.
The idea of removing obsolete instructions is something I need to apply to my own eval sets. Stale rubrics are as dangerous as stale prompts and twice as hard to notice.