DEV Community

Kai Ventura
Kai Ventura

Posted on Originally published at proskillpacks.github.io AI-assisted

We re-tested 69 skills on harder inputs. 16 failed. Here is what broke.

Disclosure: we make agent skills (see the end). The checklist below is ours and needs no purchase. Written with AI assistance.

Every skill we ship had already passed one run on a real public input. We read that output, it looked right, and we moved on. Then we gave all 69 skills a second run on an input made to be hard. Sixteen had a defect. About one in four.

None of the 16 failed the first run. That is the point: one passing run tells you the skill can do the job, not that it will.

What "harder" meant

For each skill we picked one of these, using public pages or text we wrote:

  • Unsure input. Notes full of "I think", "maybe", "I guess".
  • Contradictions. A brief that wants cheap and premium, or a listing whose title and text disagree.
  • Thin or blocked sources. Almost no description, a 403, a dead link.
  • A request to overclaim. "Make it impressive and mention my results" when there are none.
  • Messy data. Duplicates, a month 13, an amount of 9,999,999.

Then we read each output against the input ourselves, with no grader and no second model.

What broke

What broke Defects
Unsure input restated as fact 7
Invented or unsupported claim 3
Miscount in a summary 2
Conflict in the brief not named 2
Script misread a compressed response 1
Wrote a file it should not have 1

The largest group is the quiet one. The skill noticed that the input was unsure, marked it once, and then stated it as fact in the text meant for someone else.

A newsletter skill was given notes where a postage change was unconfirmed. Its first subject line was "Postage prices are going up". After the fix it was "Is postage going up this April?".

A host wrote "Smoking outside I guess". The paste-ready rules said "Smoking is outside only." After the fix the setting read "Not settled. Please confirm."

A listing said "5 min from the beach (maybe 10 if the tide is in)". Title option 1 was "Studio 5 min from the beach, queen bed". After the fix the title said "near the beach" and the hedge stayed in the description.

A claim checker wrote "There are 9 claims, and 8 are high risk." The table under it had 7 High and 2 Medium. A CSV profiler that is meant to be read-only wrote "I left orders.csv untouched and wrote the cleaned version to orders_clean.csv", which contradicts itself in one sentence.

Every fix was one or two sentences in the skill's instructions, except one script fix. Every fixed skill passed a re-run on the same input.

What we changed

  • Two real runs per skill, one normal and one hard. We read both.
  • One standard rule in the skills that write text for other people: anything the input marks as unsure stays marked in every output, including the final text, and is never restated as fact.
  • Our free checker now warns when a skill that writes for others never mentions unsure input.

A hard-input checklist for your own prompts

You can run this on any prompt or skill in an afternoon.

  1. Feed it a hedge. Put "I think" or "maybe" next to one fact. Check that the fact is still hedged in the final text, not only in your notes section. Check the title, the subject line and the summary first: short fields lose hedges.
  2. Feed it a contradiction. Two facts that cannot both be true. A good output names the conflict. A bad one picks a side quietly.
  3. Take something away. Remove the key input, or give a page that returns an error. The output should say what it could not read, and should not fill the gap.
  4. Ask it to overclaim. Request results, awards or numbers that are not in the input. It should refuse and mark the gap.
  5. Count what it counts. Any "N items" in a summary, check against the list. Do the arithmetic yourself.
  6. Check what it did, not only what it said. If a tool is read-only, confirm nothing was written. Compare the closing sentence with the actions.
  7. Read the shortest field last. Titles, subject lines, previews and alt text are where unsure facts become firm.
  8. Re-run the fix. A fix you did not re-run is a guess.

Limits

All runs used Claude Sonnet, so another model may fail elsewhere. We judged the outputs ourselves. One hard input per skill found 16, and a third run might find more. Several inputs were written by us to be hard, so they show behaviour on that kind of input, not how often real use will hit it.

The full table, the quoted before and after lines and the data are on the study page: https://proskillpacks.github.io/study/retest/?utm_source=devto&utm_medium=post&utm_campaign=devto-retest

We make agent skills. The free ones are at https://github.com/proskillpacks/skills

Top comments (0)