A skill is a short markdown file of instructions that a model reads before it does a task. Most people write them by hand and try them on a few examples. The trouble is that advice that sounds sensible can make results worse, and without a score you never notice.
So we trained one instead. The method is a scaled-down version of Microsoft's SkillOpt paper (arXiv 2605.23904). You treat the skill like weights: change it, score it, keep the change only if the score goes up on tasks the change was not made from.
This article was written with the help of an AI assistant and checked against the repo's README files, which are the only source of the numbers below.
The six steps
- Freeze the student. The model that does the tasks never changes.
- Split the tasks into train, selection and test. Train gives evidence, selection decides what is kept, test is touched once at the end.
- Roll out. Run the student on the train tasks with the current skill and record what it did.
- Reflect. A separate optimizer model reads the failures and proposes edits for patterns that repeat.
- Bound the update. A few edits per step (4, then 3, then 2 here), so a rejection tells you which idea to blame.
- Gate. Score the candidate on the selection tasks. Keep it only if the score is strictly higher.
Run 1: a browsing skill
The skill teaches a model to drive the agent-browser CLI. Student: Opus 5.5 running headless with only a Bash tool. Optimizer: Claude Fable 5.1. Tasks: 13 train, 8 selection, 11 test. Opus got every task right with no skill, so the score also gave credit for fewer tokens. Each candidate ran on the selection split twice.
| Step | Skill | Edits | Score | Tokens per task | Turns | Decision |
|---|---|---|---|---|---|---|
| 0 | s0-seed |
150-word starting skill | 0.9435 | 37.7k | 5.2 | start |
| 1 | s1 |
4 | 0.9687 | 20.8k | 3.0 | accepted |
| 2 | s2 |
3 more | 0.9644 | 23.8k | 3.2 | rejected |
| 3 | s3 |
1 more | 0.9632 | 24.5k | 3.0 | rejected |
Held-out test, 11 unseen tasks, two runs each:
| Condition | Correct | Tokens per task | Turns per task |
|---|---|---|---|
| No skill | 22/22 | 34.2k | 5.0 |
Vendor's bundled core skill |
22/22 | 109.4k | 5.5 |
| Seed skill | 22/22 | 33.3k | 5.1 |
| Trained skill | 22/22 | 29.4k | 4.1 |
Two of the three edit sets made things worse, and all three read as sensible advice. Only the gate could tell them apart.
The four edits that helped removed waste the model could not have known about: batch several commands into one call; do not chain work to open with &&, because it can time out while the page is still usable; a recipe for JavaScript dialogs and right-click; a short command list so the model stops calling --help.
The edit that hurt was a prevention rule: "never guess selectors on a page you have not seen". It added a snapshot call to tasks where guessing already worked.
Run 2: a starter you can run
The browsing run needs a browser and costs real money. So the repo also has a small starter: 18 toy tasks with code checkers, three families (invoice note to JSON, rewrite a note in house format, build a file name).
The tasks ask for a house style (date formats, id padding, file names) and never say what it is. A model cannot guess it, so the no-skill score is zero by design. The skill is where you write the conventions down.
Student: claude -p --model haiku. Optimizer: Sonnet.
| Skill | Selection (6 tasks x 2) | Test (6 tasks x 2) |
|---|---|---|
| Seed (two lines) | 0 of 12 | 0 of 12 |
Candidate c1 (215 words, 4 edits) |
10 of 12, accepted | 10 of 12 |
Caveats, all real. It is one step from 0 to 10 of 12, not a climb, and there was only one candidate, so no rejection. With 12 rollouts per cell, one task is 8 points. The optimizer pasted some training examples into the skill, against our instruction. None was a selection or test input, but it is a leak risk to watch.
We also ran the same c1 skill, trained on Haiku, against a local Ollama model (qwen3.8:27b, 4-bit). With no skill it got 0 of 18. With c1 it got 5 of 6 on the test split. That is one run on 6 tasks, and the skill was never tuned for that model.
What we would do differently
- Start with a weaker student or harder tasks. With a strong model the browsing run had little left to train.
- Write more candidates, so the gate shows a rejection and not only an acceptance.
- Run more repeats on selection, and check every task passes alone before blaming the skill.
- Strip training examples out of the proposal before it becomes a candidate.
Run it
# setup
git clone https://github.com/proskillpacks/skillopt-agent-browse && cd skillopt-agent-browse/starter
# 1. baseline (mostly FAIL, by design)
python3 run.py --skill none --split sel --tag baseline --model haiku
# 2. roll out the seed on train
python3 run.py --skill skill/seed.md --split train --tag seed-train --model haiku
# 3. reflect: ask an optimizer for at most 4 edits
python3 reflect.py runs/seed-train --skill skill/seed.md --ask --max-edits 4 --model sonnet
For a local model, set STUDENT_MODEL, OPENAI_BASE_URL=http://localhost:11434/v1 and add --adapter openai. The full walk-through, with the gate command, is in starter/README.md.
Limits
Small task sets, toy tasks in the starter, one student model per run, two to twelve rollouts per cell. Treat the numbers as a demonstration of the method, not a benchmark.
Repo: https://github.com/proskillpacks/skillopt-agent-browse
Paper: https://arxiv.org/abs/2605.23904
Longer write-up (canonical): https://proskillpacks.github.io/research/how-we-train-skills/
Top comments (1)
The gate is the right instinct, and "keep it only if the score is strictly higher" is where I'd look next. When an edit changes nothing, its selection score lands above the incumbent's about as often as below, so a do-nothing edit has roughly even odds of being kept. Since Opus solved the tasks with or without a skill, most of the differences between scores come from the token credit, and token use in agent runs can move a fair amount from run to run. You already have what you need to size that: each candidate ran twice on selection, so the gap between a skill's own two runs is the noise floor, and s1's +0.025 over the seed can be read against it. A direct test is a placebo edit: paraphrase the current skill without changing any instruction, send it through the gate a few times, and count how often it's kept.