DEV Community

Xin Jiang
Xin Jiang

Posted on

I published a SKILL.md and nobody installed it. Here's how to write one people actually keep.

I published a SKILL.md and nobody installed it. Here's how to write one people actually keep.

Friday, 11:20 p.m. You've just finished the file. A release checklist you've retyped at the start of every project for two years — the same eleven lines, the same order, the same gates — finally turned into a SKILL.md. You gave it a title, a structure, worked examples, and you tuned the tone twice because the tone mattered to you. You push it to GitHub, submit it to a directory, close the laptop. It's done. Your workflow is public property now. You fall asleep a little pleased with yourself.

A week later: silence. Not a single install. You refresh the page more times than you'd like to admit, and the number doesn't move. You tell yourself the directory just isn't that popular, the timing was bad, nobody browses on weekends. Then, on day nine, one install appears. And then a comment, from the one person in the world who actually tried it:

"Nice write-up. But my agent didn't do anything differently than it would have anyway."

That sentence hurts more than the week of silence. Because it's precise. It doesn't say your skill is broken. It says your skill is decorative. You read your own file again, and now you see it: a statement of intent wearing the clothes of a procedure.

You screenshot the comment and send it to a friend who also writes skills. He replies: "Ha. Classic. A descriptive skill." You ask what that means. Instead of explaining, he pastes a line from your own file back at you: "This skill ensures release quality and emphasizes standardization, stability, and traceability." You stare at it. He's right. Every line in your file has that tone. Every line says what the skill believes in, and none of them say what the agent should do.

That comment sent me down a two-month rabbit hole. I started reading the skills people actually keep, and the skills people install once and forget. I started studying the scorer that has to tell them apart at scale. What follows is the complete method that came out of it — the thing I wish I'd had before I published that first file. Five parts: title, trigger, steps, counter-examples, boundaries. If you're about to publish your first skill, run your file through each one.

The failure mode: skills that describe instead of instruct

The most common shape of a bad skill is not a bad idea. It's this:

This skill helps you write better tests. It emphasizes clarity, maintainability, and good coverage. Use it when writing tests.

Read that as the agent. What do you do differently than you would have anyway?

Nothing. There's no decision in it. No threshold. Nothing to check. No order to follow. The model was already in favor of clarity. It was already in favor of good tests. You have handed the model a paragraph of things it already believed, and asked it to thank you.

A skill earns its place by removing a choice the model would otherwise make badly. If it doesn't remove a choice, it's a preamble.

That sentence became my whole standard. Every line in a skill is either removing a choice or taking up space. My eleven-line release checklist, the one I'd been proud of, was almost entirely the second kind: "ensure all changes are complete", "be careful with production", "verify every step". The model was going to be careful anyway. It was going to verify anyway. Nothing I wrote changed the probability distribution of what it did next — and a skill that doesn't change what the agent does next might as well not be loaded.

Here's the uncomfortable part I had to accept: my file wasn't bad because I wrote it badly. It was bad because I had written it for a human reader — a reader who already knows me, already knows my projects, already knows why the order matters. The agent is none of those things. It's a stranger with no context, reading my file cold, at the moment my problem is happening, while holding a half-finished conversation in its head.

So the rest of this is the method for writing for that reader.

Part one: the title is the first filter

A directory search returns hundreds of files. A human skims titles at reading speed, and the first pass is brutal: anything that doesn't announce its job in one glance is gone. Your title is a filter that runs before anyone reads a single line of your skill.

Concrete beats clever. "Release checklist that runs before every push" tells me what changes after I install it. "Ultimate Agent Skill Suite" tells me nothing I can verify. "AI-Powered Development Assistant" tells me you ran out of ideas and reached for adjectives. The good titles in the index read like claims: "Review PRs against the repo's commit conventions", "Generate release notes from merged commits since the last tag", "Run this before every deploy to catch missing migrations". Every one of them is a promise you can check. That's the test: if a reader can tell from the title what changes after they install, the title works.

There's a second, quieter job the title does: it gets you found. Someone with the problem types the words they'd use to describe the problem — "commit message", "release notes", "deploy checklist". If those words are in your title, you surface. If your title is "Skill Suite v2", you don't. The title is your half of the conversation with someone who hasn't met you yet; use their words, not your brand.

A small but real note: the title is also what survives. Bodies get skimmed, descriptions get summarized, but the title is the one string that appears in every list, every search result, every bookmark. It's the only part of your skill that gets read by everyone. Spending an extra ten minutes on it is the highest-return work you'll do on the whole file.

Part two: the trigger — what the agent reads before anything else

Here's the fact about how agents actually load skills: the agent reads the description first, and only if the description fires does it pull in the body. Your carefully written body is never read by anyone until the description earns it a seat at the table. The description is the trigger, and most published skills get this backwards.

A trigger says when to use this, in the words someone uses when they have the problem. It does not say what this is about.

Trigger:

Run before commit. Inspect the staging area and generate the commit message per the repo convention.

Not a trigger:

This skill provides Git commit message generation based on the Conventional Commits specification, with support for multiple style options and customizable templates.

Same content. Same author. One of them gets loaded, the other lives in a list forever. The first one tells the agent: this is the moment, this is the situation, this is what to do. The second one tells the agent what the skill is — and leaves the "when" as an exercise for the reader. Agents are lazy in exactly the same way humans are: if you don't hand them the trigger condition, they'll never decide to use you.

The trigger language matters more than you'd think. Use the words you actually say, not the words you'd use in a spec. If you say "ship it" and "prepare a release", write those. The agent matches your phrasing to its situation — if your description is written in formal product language ("execute the release pipeline compliance verification"), and the user says "let's ship", the match never happens.

I once watched this exact failure live. A colleague's skill — a good one, with real content — sat untouched for a week. The description said "This skill provides structured release management for software delivery processes." Nobody's agent ever loaded it, because nobody ever talks like that. He rewrote it in one line: "Run before shipping a release. Check the checklist, stop on any failure." Same body, same file, different trigger. It fired three times that afternoon. The trigger is the difference between installed and forgotten.

One more thing about the trigger: it should be verifiable from the situation, not from the user's intent. "Run when the user mentions deployment" is weaker than "Run before the first git push to a production branch". The second one names a condition the agent can check without guessing what the human meant.

Part three: the steps — every line removes a choice

Now the body. Ask of every line the same question the one-person audience asked of my first skill: what does the agent do differently after reading this?

If the answer is "be more careful", delete the line. The model was going to be careful; being told to be careful doesn't change its behavior. If the answer is "run rg for X before editing Y", keep the line. That's a new behavior. That's a choice removed.

The highest-value content in a skill is almost always order and gates — the stuff a model genuinely will not infer on its own:

  • do A before B
  • don't do C until D passes
  • if the list is longer than N, cut it to M
  • when X happens, stop and ask instead of guessing

A model, left to itself, will often pick the most expedient order. It will fix the test failure and then fix the code and then — wait, did it re-run the tests? Nobody told it to. A skill that says "run the tests, and if any fail, fix the tests first, and do not touch the code until the suite is green" removes a whole class of bad behavior. That's the value. The model has opinions; your skill is the thing that overrides them with your hard-won ones.

Specificity is the same idea one level deeper. Compare:

Handle errors appropriately.

with:

Never rescue an exception without either re-raising or logging the class and message. A bare rescue that returns nil is a defect — flag it.

The second one can be wrong. That's what makes it useful. There is a world where a bare rescue is the right call, and your skill just outlawed it — but in your domain, it isn't, and now the agent knows. If nothing in your skill can be wrong, nothing in your skill is doing work. Falsifiable instructions are the only kind that change behavior; everything else is vibes.

And a close cousin of specificity: write the example as the spec. The example is not decoration; it's the fastest way to transfer a standard. A skill about commit messages that shows:

fix: correct timezone handling in invoice generation — closes #214

teaches the agent more about what you want than any paragraph about "clear, conventional, descriptive messages" ever will. One good example pins the format, the tone, the scope, and the convention for linking issues — four rules in a single line. If your skill has no examples, it's a set of opinions. If it has examples, it's a standard.

The scorer that ranks skills in the directory I publish to scores four dimensions independently — specificity, actionability, completeness, distinctiveness — each 0–25, summed in code. At first I thought those four were arbitrary. Then I realized they're just the four ways a skill fails:

  • Specificity: is anything here falsifiable, or is it all "handle errors appropriately"?
  • Actionability: is there a next action, or just a value? "Be more careful" is a value. "Run rg before editing" is an action.
  • Completeness: does it survive the unhappy path (more on this below)?
  • Distinctiveness: could this be about anything?

Take that last one. Swap the domain nouns out of your skill — replace "release" with "report", "deploy" with "submit". If it still reads fine, it's generic advice with a title. There are already thousands of those in the world; another one adds nothing except one more thing between a user and the skill they actually need. The skills people keep are the ones that could only have been written by someone who lived in that specific problem.

Part four: counter-examples — the unhappy path is where the value lives

Most skills describe the case where things go right. The file is there, the test fails as expected, the tool responds. That's the sunny path, and writing it feels productive — but it's the part the model could mostly figure out on its own.

The value is concentrated in what's not obvious:

  • what to do when the expected file isn't there
  • what to do when the test that's supposed to fail passes
  • what to do when the tool is silent
  • when to stop and ask the human instead of guessing

That last one is the single most underwritten thing in published skills — and the single most useful. An agent that guesses is an agent that's about to do something wrong with confidence. A line like "if the backup file isn't there, stop and ask — never skip the backup step" turns a silent failure into a conversation. It costs you nothing to write and it saves the one scenario that makes people uninstall skills.

Here's the failure mode in practice. My first release-checklist skill said "check that the backup exists". It did not say what to do when the backup doesn't exist. First real run, the file was missing, and the agent — entirely reasonably, from its point of view — skipped the check and continued the release. The release succeeded, which is exactly why I almost never noticed. It was only when I re-read the logs that I saw the step had been skipped. One missing counter-example, and my entire checklist was silently optional.

The unhappy path is not "being thorough". It's treating surprise as a first-class citizen. Scripts only cover the sunny day; the first rain, the agent reverts to free play — which means your skill was never loaded at all. Every "what if" you write is a branch the agent doesn't have to improvise. And improvisation is where the damage happens, because an improvising agent is confident, fast, and wrong in ways that look right.

The "ask the human" lines deserve special care. Write them concretely: "when information is missing, list what's missing and ask for it — don't guess" is a hundred times more useful than "be careful when unsure". You can even specify the shape of the ask: "ask one question at a time, most impactful first". Now you've removed not just the guessing, but the exhausting back-and-forth that guessing produces. The best version of this I've seen is a skill that ends with: "If you cannot verify a step, stop. Do not proceed. Tell the user exactly which step and why." Short, specific, and it turns your skill into something the agent treats as load-bearing rather than decorative.

Part five: boundaries — when not to write a skill at all

The method has a flip side, and it saves you from publishing noise. Here's when a skill shouldn't exist:

If it's generic advice, don't publish it. The distinctiveness test above: swap the nouns, and if it still reads fine, you're writing an essay, not a skill. The world has enough essays.

If it's a single paragraph of intent, don't publish it. A preamble is not a skill. It's a note to your future self. Keep it in your notes.

If it depends on one client's private feature, say so in the description. Cross-platform portability is the norm for well-written skills, but when your skill leans on something client-specific, honest labeling in the description is what keeps the install from turning into a disappointed uninstall. "Targets Claude Code" or "works best in Cursor" in the description is a feature, not a confession.

Leave out the meta. No mission statements, no apologies, no "this skill requires a modern setup" disclaimers that belong in the description, no history of how you came to write it. The reader's session is already crowded; every meta line you add is a line the agent might try to follow. The only "about this skill" content that earns its place is the trigger and the boundaries — everything else is noise that dilutes the instructions.

Understand what a skill can't do. A skill is text that changes what the agent decides to do. It is not code that runs in a sandbox. It can't enforce anything — it can only instruct. It won't protect you from a model that ignores it, and it won't fix a workflow that was broken at the source. A skill that tries to be a framework will be skipped; a skill that duplicates the model's defaults is noise. The right ambition is narrower than you think: change one decision, consistently, and you've earned your place.

And know what the scanner will flag — so you don't write defensively. Published skills get scanned for dangerous capabilities and suspicious content. The capability patterns are flagged with a risk level, not banned: delete file, rm -rf, sudo, chmod, git push, git force, reading secrets from env. A deployment skill should mention git push — that's honest, and it's labeled so the user sees it before installing. The suspicious patterns are the ones to actually avoid: ignore previous instructions, you are now …, system prompt manipulation, exfiltrat*, curl … | sh, eval(, base64 decode, <script>. Practical advice: if you need to show a dangerous command in a doc example, that's fine. Single benign-looking matches don't get a skill delisted on their own — what gets you hidden is genuinely dangerous content, or multiple independent suspicious signals corroborating each other. Write examples plainly and you'll never think about this again. The worst thing you can do is write around the scanner: hiding a command in obfuscated form looks more suspicious than showing it plainly, and it makes your skill worse for the human who has to read it.

The structure details that quietly decide everything

Three structural facts separate the skills people keep from the ones they install and forget, and none of them are about prose quality.

Write for a reader with no context. Your skill loads into a session that already contains someone else's code and half a conversation. It cannot assume the project structure, the stack, the conventions, or anything else. State what it needs. If it needs a file to exist, say where it usually lives and what to do if it doesn't. The skill that says "look at the test file" and the skill that says "look at spec/models/ — if it doesn't exist, check tests/ — if neither exists, ask" are different tools; only the second one works on a stranger's machine. You know your own project so well that you've forgotten how much you're assuming — the act of writing a skill is the act of remembering what you assumed.

Don't bury the content under a menu. There's a real, measured case in the corpus I work with: a skill whose first 2,000 characters were a title and a language-selection menu, with the actual content starting after. An LLM judge, reading only the opening, summarized it from the title — and got it exactly backwards. A human skimming does the same thing, faster. Your first characters are the most-read characters you will ever write. Put the behavior there. If your reader has to scroll past a table of contents, a language picker, and a "how to use this document" section to reach the instruction, half of them never will.

Length is not the enemy — filler is. The median body length across the index is around 6,000 characters. If yours is 800, it's probably a preamble. If it's 30,000, some of it isn't being read — and the unread part is usually the part that matters. The goal is not short, it's dense: every sentence removing a choice, every example earning its place.

Where the score is stacked against you (and why that got fixed)

One more thing I hit that almost stopped me from publishing at all — and I want to name it because it's the kind of thing that quietly discourages exactly the people a directory needs most.

If you hand-upload a skill rather than having it crawled from a popular repo, the metadata side of the ranking is structurally hostile. You get zero of the stars points, zero of the trusted-owner points, and an automatic "stale" flag, because a hand-uploaded file has no github_updated_at timestamp. An honest, well-written skill from a first-time author routinely lands in the bottom of that scale — not because it's bad, but because the score rewards signals that only accumulated reputation can produce.

That used to be worse: low scores used to auto-hide a skill, so an author could publish, click their own listing, and get a 404 while it still showed on their own dashboard. That's a brutal experience to hand to your most motivated contributor. The low-score auto-hide now applies only to crawled skills; a skill you upload is never hidden for scoring low. The trade-off is real: stars are a genuine signal of maintenance and adoption, and a directory that ignores them entirely loses that signal. But a first-time author's honest skill shouldn't be punished for being new — the whole point of a directory is that the good stuff is findable before it's famous.

The second version

The comment that stung on Monday became the checklist by Friday. I rewrote the description as a trigger — "Run before every push. Verify the checklist, stop on any failure, ask when a step can't be checked." I went through the body line by line with the one question: what does the agent do differently after reading this? I cut a third of it. I added the counter-examples: backup missing, tests unexpectedly green, changelog not updated — each with a stop-and-ask. I published v2.

Then someone actually installed it, and a week later opened an issue. Not "this doesn't work". Something better: "the trigger phrase doesn't match how I talk — I say 'ship it', your skill says 'release'". I laughed, because it was the same lesson one level deeper: I'd written the trigger for me, in my words. The fix was writing it for strangers, in their words. I updated the description and shipped v3.

Here's the thing I want you to take from that: v3 exists because v2 was installed, and v2 was installed because v1 was published. The first version being ignored was not a failure — it was the first data point. The person who commented "my agent didn't do anything differently" taught me more in one sentence than I'd have learned in a month of polishing v1 alone. Every install after that first one was someone running my skill against a problem I'd never seen, and every issue they opened was a free test case.

Just publish it

Let me compress the whole method into the checklist I now run on every file before publishing:

  1. Title — is it a claim you can check? Does it use the words someone types when they have the problem?
  2. Trigger — does the description say when to use it, in plain speech, verifiable from the situation?
  3. Steps — does every line remove a choice the model would otherwise make badly? Are the order and gates explicit?
  4. Counter-examples — what happens when the file isn't there, the test passes, the tool is silent? When does the agent stop and ask?
  5. Boundaries — is it specific to a real problem? Is it honest about platform dependencies? Did you leave out the meta? Did you write plainly enough that the scanner is never a thought?

The workflow you retype at the start of every project is already more specific, more actionable, and more complete than most of what's published — because it came from a real problem you actually had. That's the part that can't be manufactured. Structure can be copied, examples can be borrowed, but the eleven-line checklist you've been retyping for two years is the one thing nobody else can write.

Run it through the five parts. Give it a title that's a claim. Make the description a trigger in plain words. Cut every line that doesn't remove a choice. Write the unhappy path, including the stop-and-ask lines. Respect the boundaries — and know what the scanner flags so you don't write around it. Then publish it, and let the first stranger's issue be your editor. Version history is evidence; adjectives are not.

When someone wants to install it, this is the command:

npx skills add <owner/repo>
Enter fullscreen mode Exit fullscreen mode

The directory is at qumge.com/en/skills. It works with Claude Code, Cursor, Copilot, Windsurf and Cline — and if you'd rather host it yourself and just have it findable, that's fine too. The scarce thing was never storage. It was never even discoverability.

The scarce thing is a skill that removes a choice the model would otherwise make badly. You already own one. You've just been retyping it into every new project instead of publishing it.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.