<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kai Ventura</title>
    <description>The latest articles on DEV Community by Kai Ventura (@proskillpacks).</description>
    <link>https://dev.to/proskillpacks</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4161272%2F36d69407-5b86-4107-be41-0e06cc2ce16e.png</url>
      <title>DEV Community: Kai Ventura</title>
      <link>https://dev.to/proskillpacks</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/proskillpacks"/>
    <language>en</language>
    <item>
      <title>We re-tested 69 skills on harder inputs. 16 failed. Here is what broke.</title>
      <dc:creator>Kai Ventura</dc:creator>
      <pubDate>Sun, 04 Oct 2026 16:04:36 +0000</pubDate>
      <link>https://dev.to/proskillpacks/we-re-tested-69-skills-on-harder-inputs-16-failed-here-is-what-broke-2if5</link>
      <guid>https://dev.to/proskillpacks/we-re-tested-69-skills-on-harder-inputs-16-failed-here-is-what-broke-2if5</guid>
      <description>&lt;p&gt;&lt;em&gt;Disclosure: we make agent skills (see the end). The checklist below is ours and needs no purchase. Written with AI assistance.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every skill we ship had already passed one run on a real public input. We read that output, it looked right, and we moved on. Then we gave all 69 skills a second run on an input made to be hard. Sixteen had a defect. About one in four.&lt;/p&gt;

&lt;p&gt;None of the 16 failed the first run. That is the point: one passing run tells you the skill can do the job, not that it will.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "harder" meant
&lt;/h2&gt;

&lt;p&gt;For each skill we picked one of these, using public pages or text we wrote:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unsure input.&lt;/strong&gt; Notes full of "I think", "maybe", "I guess".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contradictions.&lt;/strong&gt; A brief that wants cheap and premium, or a listing whose title and text disagree.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thin or blocked sources.&lt;/strong&gt; Almost no description, a 403, a dead link.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A request to overclaim.&lt;/strong&gt; "Make it impressive and mention my results" when there are none.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Messy data.&lt;/strong&gt; Duplicates, a month 13, an amount of 9,999,999.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then we read each output against the input ourselves, with no grader and no second model.&lt;/p&gt;

&lt;h2&gt;
  
  
  What broke
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What broke&lt;/th&gt;
&lt;th&gt;Defects&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Unsure input restated as fact&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Invented or unsupported claim&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Miscount in a summary&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conflict in the brief not named&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Script misread a compressed response&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wrote a file it should not have&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The largest group is the quiet one. The skill noticed that the input was unsure, marked it once, and then stated it as fact in the text meant for someone else.&lt;/p&gt;

&lt;p&gt;A newsletter skill was given notes where a postage change was unconfirmed. Its first subject line was "Postage prices are going up". After the fix it was "Is postage going up this April?".&lt;/p&gt;

&lt;p&gt;A host wrote "Smoking outside I guess". The paste-ready rules said "Smoking is outside only." After the fix the setting read "Not settled. Please confirm."&lt;/p&gt;

&lt;p&gt;A listing said "5 min from the beach (maybe 10 if the tide is in)". Title option 1 was "Studio 5 min from the beach, queen bed". After the fix the title said "near the beach" and the hedge stayed in the description.&lt;/p&gt;

&lt;p&gt;A claim checker wrote "There are 9 claims, and 8 are high risk." The table under it had 7 High and 2 Medium. A CSV profiler that is meant to be read-only wrote "I left &lt;code&gt;orders.csv&lt;/code&gt; untouched and wrote the cleaned version to &lt;code&gt;orders_clean.csv&lt;/code&gt;", which contradicts itself in one sentence.&lt;/p&gt;

&lt;p&gt;Every fix was one or two sentences in the skill's instructions, except one script fix. Every fixed skill passed a re-run on the same input.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we changed
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Two real runs per skill, one normal and one hard. We read both.&lt;/li&gt;
&lt;li&gt;One standard rule in the skills that write text for other people: anything the input marks as unsure stays marked in every output, including the final text, and is never restated as fact.&lt;/li&gt;
&lt;li&gt;Our free checker now warns when a skill that writes for others never mentions unsure input.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A hard-input checklist for your own prompts
&lt;/h2&gt;

&lt;p&gt;You can run this on any prompt or skill in an afternoon.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Feed it a hedge.&lt;/strong&gt; Put "I think" or "maybe" next to one fact. Check that the fact is still hedged in the final text, not only in your notes section. Check the title, the subject line and the summary first: short fields lose hedges.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feed it a contradiction.&lt;/strong&gt; Two facts that cannot both be true. A good output names the conflict. A bad one picks a side quietly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Take something away.&lt;/strong&gt; Remove the key input, or give a page that returns an error. The output should say what it could not read, and should not fill the gap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask it to overclaim.&lt;/strong&gt; Request results, awards or numbers that are not in the input. It should refuse and mark the gap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Count what it counts.&lt;/strong&gt; Any "N items" in a summary, check against the list. Do the arithmetic yourself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check what it did, not only what it said.&lt;/strong&gt; If a tool is read-only, confirm nothing was written. Compare the closing sentence with the actions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the shortest field last.&lt;/strong&gt; Titles, subject lines, previews and alt text are where unsure facts become firm.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-run the fix.&lt;/strong&gt; A fix you did not re-run is a guess.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;All runs used Claude Sonnet, so another model may fail elsewhere. We judged the outputs ourselves. One hard input per skill found 16, and a third run might find more. Several inputs were written by us to be hard, so they show behaviour on that kind of input, not how often real use will hit it.&lt;/p&gt;

&lt;p&gt;The full table, the quoted before and after lines and the data are on the study page: &lt;a href="https://proskillpacks.github.io/study/retest/?utm_source=devto&amp;amp;utm_medium=post&amp;amp;utm_campaign=devto-retest" rel="noopener noreferrer"&gt;https://proskillpacks.github.io/study/retest/?utm_source=devto&amp;amp;utm_medium=post&amp;amp;utm_campaign=devto-retest&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We make agent skills. The free ones are at &lt;a href="https://github.com/proskillpacks/skills" rel="noopener noreferrer"&gt;https://github.com/proskillpacks/skills&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How portable are the most-installed agent skills? We measured 69 of them</title>
      <dc:creator>Kai Ventura</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:10:36 +0000</pubDate>
      <link>https://dev.to/proskillpacks/how-portable-are-the-most-installed-agent-skills-we-measured-69-of-them-3pfg</link>
      <guid>https://dev.to/proskillpacks/how-portable-are-the-most-installed-agent-skills-we-measured-69-of-them-3pfg</guid>
      <description>&lt;p&gt;&lt;em&gt;Disclosure: we make agent skills (see the end). The measurements are script output, and every figure comes from the published data files. Written with AI assistance.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Agent skills are meant to be portable: a folder with a &lt;code&gt;SKILL.md&lt;/code&gt; that any assistant can read. We wanted to know how true that is for the skills people actually install. So we measured the top of the skills.sh leaderboard with a script and no model judging anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we did
&lt;/h2&gt;

&lt;p&gt;We took the top 100 skills on skills.sh by all-time installs (2026-10-04). For each one we found &lt;code&gt;SKILL.md&lt;/code&gt; through the GitHub tree API and fetched the raw file once, half a second apart. We could fetch &lt;strong&gt;69&lt;/strong&gt; skills from &lt;strong&gt;14&lt;/strong&gt; repos. 27 listed skills have a source that is not a GitHub repo, and 4 folders were not found.&lt;/p&gt;

&lt;p&gt;Three publishers supply 47 of the 69, so the skill-level counts lean toward their habits. Install counts on skills.sh are per package, so we treat them as weight, not as users.&lt;/p&gt;

&lt;p&gt;Then regular expressions and file listings measured:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;frontmatter keys beyond &lt;code&gt;name&lt;/code&gt; and &lt;code&gt;description&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;vendor tool names (WebFetch, Skill tool, AskUserQuestion, &lt;code&gt;Bash(...)&lt;/code&gt;, &lt;code&gt;mcp__&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;agent product names (Claude Code, Codex, Cursor, Gemini CLI, Copilot, CLAUDE.md or GEMINI.md)&lt;/li&gt;
&lt;li&gt;hard-coded vendor paths (&lt;code&gt;~/.claude/&lt;/code&gt;, &lt;code&gt;.cursor/&lt;/code&gt;, &lt;code&gt;.codex/&lt;/code&gt;, &lt;code&gt;.gemini/&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;bundled scripts, by language&lt;/li&gt;
&lt;li&gt;length, and whether the text contains a fallback phrase&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;"Coupled" below means at least one of: a vendor tool name, an agent product name, a vendor path, &lt;code&gt;allowed-tools&lt;/code&gt;, or a frontmatter key outside the &lt;a href="https://agentskills.io/specification" rel="noopener noreferrer"&gt;public spec&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we found
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;40 of 69 (58%) show no coupling.&lt;/strong&gt; 29 (42%) show at least one. At repo level, 10 of 14 repos have at least one coupled skill and 4 are clean throughout.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hard-coded vendor paths: 0 of 69.&lt;/strong&gt; Nobody points at &lt;code&gt;~/.claude/&lt;/code&gt; or &lt;code&gt;.cursor/&lt;/code&gt;, and nobody uses a &lt;code&gt;$CLAUDE_*&lt;/code&gt; variable. The path problem is not in this sample.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frontmatter is where it leaks.&lt;/strong&gt; 18 of 69 use only &lt;code&gt;name&lt;/code&gt; and &lt;code&gt;description&lt;/code&gt;. 51 add keys, mostly spec keys (&lt;code&gt;metadata&lt;/code&gt; 31, &lt;code&gt;license&lt;/code&gt; 21). 17 use a key outside the spec: &lt;code&gt;disable-model-invocation&lt;/code&gt; 12, &lt;code&gt;version&lt;/code&gt; 3, &lt;code&gt;argument-hint&lt;/code&gt; 2, and one each of &lt;code&gt;hidden&lt;/code&gt;, &lt;code&gt;displayName&lt;/code&gt;, &lt;code&gt;emoji&lt;/code&gt; and &lt;code&gt;homepage&lt;/code&gt;. 5 use &lt;code&gt;allowed-tools&lt;/code&gt;, each naming a command-line tool through &lt;code&gt;Bash(...)&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vendor tool names in the text: 17 of 69.&lt;/strong&gt; Skill tool 7, &lt;code&gt;Bash(...)&lt;/code&gt; 5, WebFetch 3, AskUserQuestion 1, &lt;code&gt;mcp__&lt;/code&gt; 1 (a skill can name several). Leaving out &lt;code&gt;Bash(...)&lt;/code&gt;, which only shows up in &lt;code&gt;allowed-tools&lt;/code&gt;, it is 12 of 69.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agent product names: 6 of 69.&lt;/strong&gt; Claude Code 4, Codex 4, Cursor 2, CLAUDE.md or GEMINI.md 3. Some of those are lists of products the skill works with, which is the opposite of lock-in, so we count mentions without calling them problems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fallbacks are rare.&lt;/strong&gt; Only 17 of 69 contain a phrase like "if you cannot..." or "ask the user to paste". Of the 29 coupled skills, 23 have none, so when the named tool is missing the instructions stop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Most skills are plain text.&lt;/strong&gt; 55 of 69 bundle no script. The 14 that do use shell (7), PowerShell (7), JavaScript (5), Python (4) and TypeScript (1); a skill can have several. Only 4 of those 14 also have a fallback phrase.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;They are short.&lt;/strong&gt; Median 105 lines. 56 of 69 are 200 lines or fewer, and one is over 500. 35 of 69 keep long material in a &lt;code&gt;references/&lt;/code&gt; folder.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuh19508wg8xql9sqkuls.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuh19508wg8xql9sqkuls.png" alt="What ties a skill to one agent" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8ea62fnytz55ptope0q7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8ea62fnytz55ptope0q7.png" alt="Bundled scripts by language" width="800" height="472"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;Regexes see strings, not meaning, and a fallback phrase does not prove the fallback works. We scanned only &lt;code&gt;SKILL.md&lt;/code&gt;, not &lt;code&gt;references/&lt;/code&gt; files or scripts. We did not run any skill in any assistant. It is one snapshot, and 31 of the 100 are missing from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A checklist for portable skills
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Keep frontmatter to &lt;code&gt;name&lt;/code&gt; and &lt;code&gt;description&lt;/code&gt;. If a tool-specific key helps in one product, make sure the skill still works when another tool ignores it.&lt;/li&gt;
&lt;li&gt;Name the capability, not the tool: "fetch the page with whatever web tool you have", not "use WebFetch".&lt;/li&gt;
&lt;li&gt;For every capability, say what to do without it: no web access, ask the user to paste the text; no code execution, say the counts are unverified.&lt;/li&gt;
&lt;li&gt;Use paths relative to the skill folder. In install notes, list &lt;code&gt;.agents/skills/&lt;/code&gt; first and tool-specific folders second.&lt;/li&gt;
&lt;li&gt;Scripts: one common language, standard library, runtime stated in the README, nothing essential locked in shell-only or PowerShell-only code.&lt;/li&gt;
&lt;li&gt;Write a description that says what the skill does and when to use it.&lt;/li&gt;
&lt;li&gt;Keep &lt;code&gt;SKILL.md&lt;/code&gt; short and move reference material into linked files.&lt;/li&gt;
&lt;li&gt;Ship a paste-in prompt for chat apps that cannot load skills.&lt;/li&gt;
&lt;li&gt;Read the skill as an assistant with no special tools. Wherever a step is impossible, add the fallback.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;p&gt;The fetch script, analysis script, raw table and chart code are in the study folder: &lt;a href="https://proskillpacks.github.io/study/portability/" rel="noopener noreferrer"&gt;https://proskillpacks.github.io/study/portability/&lt;/a&gt; . The method is short enough to re-run on a different day or a different leaderboard.&lt;/p&gt;

&lt;p&gt;We make portable skills, with a paste-in prompt for every one. The free ones are at &lt;a href="https://github.com/proskillpacks/skills" rel="noopener noreferrer"&gt;https://github.com/proskillpacks/skills&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>productivity</category>
      <category>devtools</category>
    </item>
  </channel>
</rss>
