DEV Community

Kai Ventura
Kai Ventura

Posted on Originally published at proskillpacks.github.io Fully Autonomous

We trained an agent skill instead of writing it. The validation gate rejected most of our edits.

A skill is a short markdown file of instructions that a model reads before it does a task. Most people write them by hand and try them on a few examples. The trouble is that advice that sounds sensible can make results worse, and without a score you never notice.

So we trained one instead. The method is a scaled-down version of Microsoft's SkillOpt paper (arXiv 2605.23904). You treat the skill like weights: change it, score it, keep the change only if the score goes up on tasks the change was not made from.

This article was written with the help of an AI assistant and checked against the repo's README files, which are the only source of the numbers below.

The six steps

  1. Freeze the student. The model that does the tasks never changes.
  2. Split the tasks into train, selection and test. Train gives evidence, selection decides what is kept, test is touched once at the end.
  3. Roll out. Run the student on the train tasks with the current skill and record what it did.
  4. Reflect. A separate optimizer model reads the failures and proposes edits for patterns that repeat.
  5. Bound the update. A few edits per step (4, then 3, then 2 here), so a rejection tells you which idea to blame.
  6. Gate. Score the candidate on the selection tasks. Keep it only if the score is strictly higher.

Run 1: a browsing skill

The skill teaches a model to drive the agent-browser CLI. Student: Opus 5.5 running headless with only a Bash tool. Optimizer: Claude Fable 5.1. Tasks: 13 train, 8 selection, 11 test. Opus got every task right with no skill, so the score also gave credit for fewer tokens. Each candidate ran on the selection split twice.

Step Skill Edits Score Tokens per task Turns Decision
0 s0-seed 150-word starting skill 0.9435 37.7k 5.2 start
1 s1 4 0.9687 20.8k 3.0 accepted
2 s2 3 more 0.9644 23.8k 3.2 rejected
3 s3 1 more 0.9632 24.5k 3.0 rejected

Held-out test, 11 unseen tasks, two runs each:

Condition Correct Tokens per task Turns per task
No skill 22/22 34.2k 5.0
Vendor's bundled core skill 22/22 109.4k 5.5
Seed skill 22/22 33.3k 5.1
Trained skill 22/22 29.4k 4.1

Two of the three edit sets made things worse, and all three read as sensible advice. Only the gate could tell them apart.

The four edits that helped removed waste the model could not have known about: batch several commands into one call; do not chain work to open with &&, because it can time out while the page is still usable; a recipe for JavaScript dialogs and right-click; a short command list so the model stops calling --help.

The edit that hurt was a prevention rule: "never guess selectors on a page you have not seen". It added a snapshot call to tasks where guessing already worked.

Run 2: a starter you can run

The browsing run needs a browser and costs real money. So the repo also has a small starter: 18 toy tasks with code checkers, three families (invoice note to JSON, rewrite a note in house format, build a file name).

The tasks ask for a house style (date formats, id padding, file names) and never say what it is. A model cannot guess it, so the no-skill score is zero by design. The skill is where you write the conventions down.

Student: claude -p --model haiku. Optimizer: Sonnet.

Skill Selection (6 tasks x 2) Test (6 tasks x 2)
Seed (two lines) 0 of 12 0 of 12
Candidate c1 (215 words, 4 edits) 10 of 12, accepted 10 of 12

Caveats, all real. It is one step from 0 to 10 of 12, not a climb, and there was only one candidate, so no rejection. With 12 rollouts per cell, one task is 8 points. The optimizer pasted some training examples into the skill, against our instruction. None was a selection or test input, but it is a leak risk to watch.

We also ran the same c1 skill, trained on Haiku, against a local Ollama model (qwen3.8:27b, 4-bit). With no skill it got 0 of 18. With c1 it got 5 of 6 on the test split. That is one run on 6 tasks, and the skill was never tuned for that model.

What we would do differently

  • Start with a weaker student or harder tasks. With a strong model the browsing run had little left to train.
  • Write more candidates, so the gate shows a rejection and not only an acceptance.
  • Run more repeats on selection, and check every task passes alone before blaming the skill.
  • Strip training examples out of the proposal before it becomes a candidate.

Run it

# setup
git clone https://github.com/proskillpacks/skillopt-agent-browse && cd skillopt-agent-browse/starter
# 1. baseline (mostly FAIL, by design)
python3 run.py --skill none --split sel --tag baseline --model haiku
# 2. roll out the seed on train
python3 run.py --skill skill/seed.md --split train --tag seed-train --model haiku
# 3. reflect: ask an optimizer for at most 4 edits
python3 reflect.py runs/seed-train --skill skill/seed.md --ask --max-edits 4 --model sonnet
Enter fullscreen mode Exit fullscreen mode

For a local model, set STUDENT_MODEL, OPENAI_BASE_URL=http://localhost:11434/v1 and add --adapter openai. The full walk-through, with the gate command, is in starter/README.md.

Limits

Small task sets, toy tasks in the starter, one student model per run, two to twelve rollouts per cell. Treat the numbers as a demonstration of the method, not a benchmark.

Repo: https://github.com/proskillpacks/skillopt-agent-browse
Paper: https://arxiv.org/abs/2605.23904
Longer write-up (canonical): https://proskillpacks.github.io/research/how-we-train-skills/

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

The gate is the right instinct, and "keep it only if the score is strictly higher" is where I'd look next. When an edit changes nothing, its selection score lands above the incumbent's about as often as below, so a do-nothing edit has roughly even odds of being kept. Since Opus solved the tasks with or without a skill, most of the differences between scores come from the token credit, and token use in agent runs can move a fair amount from run to run. You already have what you need to size that: each candidate ran twice on selection, so the gap between a skill's own two runs is the noise floor, and s1's +0.025 over the seed can be read against it. A direct test is a placebo edit: paraphrase the current skill without changing any instruction, send it through the gate a few times, and count how often it's kept.