Five kinds of skills that earn their place in your agent
The Monday morning standup is the moment it hits you. You're going around the room, and when it's your turn, you open your mouth to tell everyone what you got done with the agent last week — and you realize you can't. Not because you did nothing. Because you installed twelve skills last week, and by Friday you couldn't name three of them.
You remember Wednesday night pretty well, actually. That's when the spree happened. You'd seen a thread on Hacker News about someone's agent setup, which linked to an awesome-list, which linked to a collection, which linked to a dozen more collections. One browser window became six, then twelve. You cloned the big repos. You scrolled the categories. You read the comments — "absolute game changer," "this one fixed my whole release flow" — and you installed the ones with the most stars, the ones with the nicest READMEs, the one a commenter said they "literally can't work without." Twelve installs in an hour, each one a one-liner:
npx skills add <owner/repo>
Easy. And by Friday, the honest accounting: your agent behaved... about the same. Marginally better in a couple of spots you couldn't quite pin down. You couldn't have told anyone which of the twelve were earning their keep, because you'd never run any of them on a real task in isolation. They were just there — twelve files loaded into context on the off chance that they'd help, each one a paragraph of good intentions adding a little more noise to every session.
Your teammate across the desk has exactly one skill. He installed it in February. He talks about it the way people talk about a favorite chair. "It's just a checklist," he says when you ask, "tells the agent when to stop and ask me before it does something stupid. It's saved my ass twice." One skill. Named it, trusts it, knows what it does. Meanwhile you have twelve and you can't describe any of them. That contrast sat in your head all weekend, and it's still sitting there on Monday.
By Tuesday you'd stopped being annoyed and started being curious. What does the guy with one skill know that the guy with twelve doesn't? The answer, you suspected, wasn't that he'd found a better skill — it was that he'd found a better way to choose. You spent that week testing the hypothesis: instead of installing by repository, install by job. Instead of counting skills, count the jobs each one does. The five kinds below are what that week produced, and they've held up ever since.
And here's the thing that actually bothered you, standing there in the standup: the problem isn't finding skills. Finding was never the problem — you found twelve in an hour. The problem is that most of them are a title plus a paragraph of good intentions, and there is no reliable way to tell which ones will actually change what your agent does. Word of mouth doesn't transfer — the skill that's life-changing for your teammate's Rails monolith is pure noise in your Go service. Stars don't tell you about the body. And nobody has time to read forty files to find the two that matter.
So you need a different way to shop. Stop shopping by repo. Shop by job.
What "shop by job" means, concretely: think about the jobs you actually give your agent every week — reviewing a PR, doing a release, writing a commit message, deciding whether to delete something. Those are the jobs. Then find skills that do those jobs, not skills that live in impressive-sounding repositories. A skill earns its place by taking one job off your plate and doing it the way you'd do it. That's the filter. Everything else is a hobby.
Five kinds of skills earn their place in your agent. Install one from each category, run each on a real task, and keep only what changes your output — you will end up with a smaller, better kit than most people's, and you'll be able to name every single thing in it.
1. Project-convention enforcers
Every repository has rules the model cannot know. How tests are named. How errors are handled. What a commit message looks like. What "done" means. The model has never seen your team's pull requests; it has no way to infer that your group uses [JIRA-123] prefixes and rejects anything else, or that a bare rescue swallowing an exception is treated as a defect in code review.
You know this from the incident. Last month, you handed the agent a routine refactor and it came back with commit messages that were grammatically perfect and completely useless — "fix bug" for a change that touched three modules, "improve code" for the one that rolled back an entire feature. Your lead rejected the PR with a single comment: "what happened to our commit format?" Nothing was wrong with the code. Everything was wrong with the part nobody had told the model about. The convention existed — it just lived in your head and in six years of commit history.
A good convention skill states those rules as gates, not vibes. Consider the difference:
- Good: "Never rescue an exception without re-raising or logging the class and message. A bare rescue returning nil is a defect — flag it."
- Skip: "Handle errors appropriately."
The second one was true before the skill existed. The model already knows to "handle errors appropriately"; that sentence changes nothing. The first one removes a choice the model would otherwise make badly — it defines what "appropriate" means in your codebase, with a concrete rule and a concrete consequence. That is the whole test: does the skill remove a choice the model would otherwise get wrong? If the model would have done the right thing anyway, the skill is decoration. If the model would have guessed wrong, the skill is load-bearing.
The convention skill you ended up writing after the incident is almost embarrassing in its simplicity: one paragraph, five rules, zero fluff. "Commit messages: result first, no process verbs, max three lines, prefix with the ticket number. A 'fix bug' message is a defect — rewrite it." That's it. That's the entire skill. It took ten minutes to write and it fixed the thing that had been quietly embarrassing you for a month. The lesson stuck: the most valuable skills look too small to be valuable.
2. Multi-step procedures
Anything you do in a fixed order — a release, a PR review, an onboarding pass, a migration. The value lives in the ordering and the gates: do A before B, don't merge until C passes, run D only after E is green. A model will not infer your team's sequence, because it has never seen your team. It has never stood in the room while your lead said "we always bump the version before the changelog, not after." That sequence is invisible to it, and that's exactly the job: encode the sequence.
The release incident is the one that made this category real for you. You let the agent "do a release" once, unprompted, to see what would happen. What happened: it updated the changelog, then bumped the version, then ran the tests. Three steps, all correct, two in the wrong order — which broke the build, which broke the tag, which cost you an hour of untangling a release that should have taken ten minutes. The model didn't do anything insane. It just didn't know your sequence, because your sequence was never written down anywhere the agent could read.
That's the category's whole job: make the order and the gates explicit so the agent doesn't have to guess. And it's where agents drift most when left unprompted — ask a model to "do a release" and it will produce a plausible-sounding but slightly-wrong order, and in a release, slightly wrong means broken. A procedure skill pins the order and the gates. The great thing about this category is that it's self-testing: the skill either produces the sequence you know or it doesn't, and you'll find out on the first real run.
The onboarding pass is the quiet sibling of the release here, and it's worth a mention because it's the one people forget. You have a list of things a new repo needs before anyone can work in it: the env file template, the pre-commit hook, the test command, the "how we name branches" doc. That's a procedure — fixed order, fixed gates — and it's exactly the kind of thing that currently lives in your head and in three different READMEs, and gets half-done every time a new repo is created. A procedure skill makes the half-done thing impossible.
The gate concept deserves its own paragraph, because it's the part people skip. A procedure without gates is just a list with good intentions — the model will cheerfully do all five steps in the wrong situation. The gates are what make it safe: "don't merge until C passes," "if the tests fail, stop and show me the output, don't fix it silently," "if this is a Friday, ask before you even start." Gates are the difference between a skill that guides and a skill that guards. Both are valuable; only one of them has saved anyone's week.
3. Stop-and-ask checklists
The most undervalued category. A skill that tells the agent when to stop and ask a human — before deleting, before force-pushing, before rewriting shared history, before any decision with no safe default — protects you more than a hundred "be careful" instructions ever will.
Think about why. "Be careful" is a sentiment; the model cannot act on it. "Before you delete anything, list the files and ask me" is an instruction the model can follow mechanically. The difference between the two is the difference between hoping and specifying. And the near-miss that sold you on this category is still fresh: you once watched the agent propose a "cleanup" that involved force-pushing over the main branch history. It asked permission, in a sense — it presented the command in its plan, right in the middle of a long list, formatted identically to the harmless steps around it. You caught it only because you were reading line by line. Your teammate's one skill exists precisely so that a moment like that is impossible: the agent stops, presents the destructive action as its own item, and waits.
The best stop-and-ask skills give the agent a small list of tripwires: when you see X, do not proceed; stop and ask. If a skill has nothing else, this alone earns its install — it's the only category that pays you back during the failures, which is when you actually need it.
The reason this category is so undervalued is that it earns nothing on the happy path. On a normal Tuesday, a stop-and-ask skill does nothing at all, and it's easy to conclude it's dead weight — the same way you concluded, around skill number nine on Wednesday night, that checklists are boring. Then the day comes when the agent proposes something irreversible, and the skill is the difference between "hold on, let me check" and a force-push you'll be explaining for a week. Insurance is boring until it isn't.
Your teammate's one skill, the famous favorite chair, turns out to be exactly this category. When you finally asked him to show you the file, it was short — "before deleting anything, before force-pushing, before changing shared config, list what you're about to do and wait for explicit approval." Six lines of markdown, six months of service, two incidents prevented that you know about. "The thing is," he said, "it doesn't do anything 99% of the time. But the 1% is where I live." That's the pitch for the whole category, delivered by a man with one skill and no regrets.
4. Output-format skills
Fixed formats are where models drift most. Commit messages. Changelogs. Release notes. Standup summaries. The model is already in favor of clarity; what it needs is a mold. A good output-format skill gives a template plus the rules for filling it: result first, process verbs not allowed, maximum three bullet points, dates in this format, links in that position.
You have a changelog story too, because of course you do. The agent wrote a genuinely good changelog entry — accurate, well-worded, complete — in entirely the wrong shape: prose paragraphs instead of bullet points, no version grouping, the date format your team had explicitly banned three years ago. Anyone reading it would have said "that's fine," and that's exactly the problem. Drift is silent. The content was right; the shape was off; nobody notices until the release notes get pasted somewhere public and look wrong next to every previous release.
A format skill doesn't just improve the output; it makes the output consistent, and consistency is what makes the output reviewable. You stop checking the shape and start checking the content. That's a real workflow win, and it's available for every fixed format you own: commit messages, changelogs, release notes, standup summaries, incident reports. Write the mold once, and the model fills it without supervision.
The standup summary is the sneaky one, because everyone has the format and nobody has written it down. Yours is: three lines, past tense, one line per item, result first, no "worked on." The model will happily produce "worked on the auth refactor" as a standup line, which is exactly the line that makes your manager's eye twitch. One template, three rules, and the daily ritual stops being a negotiation with the agent about what "summarize" means. Formats are the category where the smallest file wins the most often.
5. Domain and philosophy skills
Two different things, both worth having:
- Domain skills encode how a specific framework, system, or codebase should be worked with: the idioms, the sharp edges, the parts of the docs everyone forgets, the deployment quirks. Concretely: your service has a deployment process that involves a blue-green switch with a fifteen-second drain window, and the official docs describe a completely different process. The person who knows the real process is you, and right now that knowledge lives in your head and in two PR comments from last year. A domain skill puts it where the agent can reach it: "drain for 15 seconds before switching; the health check endpoint is /healthz, not /health; never deploy on Fridays unless the on-call has acknowledged." If you maintain a library or a service, writing the domain skill is the highest-leverage documentation you will produce this quarter — it's your org's memory, extracted.
-
Philosophy skills encode how you want the agent to operate overall — when to be terse, when to question the request, when to push back instead of complying.
obra/superpowersis the canonical example of that genre and is worth reading even if you never install it; it's a masterclass in writing operational rules the model can actually follow, and half of what it demonstrates is that this genre works at all.
And while we're naming the ecosystem's backbone: anthropics/skills is where the first-party examples live — the best place to start when you're unsure what a good SKILL.md looks like. Those two repos are the backbone of the community; everything else, including any directory, sorts around them. If you read nothing else this week, read those.
If you're unsure how to write any of these five kinds yourself, the reading order is simple: start with anthropics/skills to calibrate what a good file looks like, then read obra/superpowers to see how far the philosophy genre can go, then write your smallest workflow — one convention, one gate — and publish it even if it feels too small. It isn't. Small and concrete beats ambitious and vague in this format every single time.
How to tell before you install
Same test every time, thirty seconds. Does the description say when to use it, in the words you would use when you have the problem — or what it's about? "Run before commit. Inspect the staged changes and generate a commit message per repo convention" is a trigger. "This skill provides conventional commit message generation with multiple style options" is a topic. The first one tells the agent when to load it; the second one asks the agent to figure that out, and it won't.
You can run this test on anything, including the twelve skills still sitting in your agent. Go through them right now, one by one, and read only the description. The ones that say "when" are candidates. The ones that say "what" are dead weight — they're not failing to be useful, they're failing to be loaded at all. Most of your twelve, if you're honest, will be "what" skills. That's not a judgment on their authors; it's a description of the format. The description is the front door, and most skills forgot to put a doorknob on it.
You did this on Sunday, and it took twenty minutes: twelve descriptions, read at a pace you'd normally reserve for skimming headlines. Seven of your twelve were "what" skills. You deleted them on the spot, without ceremony, before you'd even gotten to the body-reading stage. That's the beautiful thing about the thirty-second test — it's a filter that costs nothing and runs in bulk. You can apply it to fifty skills in an afternoon and come out the other side with five candidates, without having read a single full body.
Then you skim the bodies of the five survivors, and this is where the test gets personal. One of your five has a gate you don't agree with — "always rebase before merging" — and you know your team's rule is the opposite. That's not a mark against the skill; it's a mark against the fit, and fit is the thing you're actually shopping for. Thirty seconds of body-skimming just saved you from installing a skill that would have fought your team's process every single day. That's the second filter doing its job.
Then skim the body for the three things above: gates (rules that remove a choice), order (sequence and checkpoints), and a stop-and-ask line (when to halt and consult you). If a skill has one of the three, it's doing real work. If it has none of the three, it's a paragraph of good intentions, no matter how nice the README is.
The mechanical scores you see on directory pages — a metadata heuristic over eight dimensions, or an LLM reading the body across four — are useful for sorting and no more. Two minutes of reading beats any score, including ours.
One honest caveat: the categories overlap, and a skill that changes your teammate's workflow can be pure noise in yours. The five kinds are a way to shop, not a promise that everything inside a category is good. A stop-and-ask checklist written for a team with a strict approval process will nag you into submission if your process is "just do it and tell me later." The overlap is real; the sorting is the point. You're not looking for the five perfect skills; you're looking for the five kinds of value, and then for the one skill in each kind that fits your actual work.
One more note on overlap, because it's where people get stuck: the categories are a map, not a taxonomy. A stop-and-ask checklist can also be an output-format skill — "before you reply with a plan, format it as: what you'll change, what you'll touch, what could break" is both a tripwire and a mold. A release procedure is a convention enforcer for the specific convention of "how we ship." Don't argue about which box a skill belongs in; argue about whether it does a job you have. The map is only there so you remember to shop for all five jobs instead of the two obvious ones.
Install, run, keep
npx skills add <owner/repo>
Pick one skill from each of the five kinds. Run each on a real task this week — not a toy, the actual thing you'd do anyway. Keep the ones that changed what your agent produced. Delete the rest. The deleting is the productive part, and it's the part everyone skips: a skill you keep without evidence is a skill you installed for the feeling of installing it.
Your Sunday cleanout is the closest thing you've had to a religious experience this year. You went through the twelve, read the descriptions, ran the three that had real triggers on real tasks, and deleted nine without ceremony. The agent got quieter. Sessions got faster to read. The noise — the paragraphs of good intentions loading into every conversation — was gone, and you felt the difference within a day. What you kept were the two you could describe out loud, plus the one you discovered while cleaning: a stop-and-ask skill you'd forgotten you had, the closest thing to your teammate's favorite chair that you own.
Then came the actual experiment, which is the part the cleanout had only set up. You picked one skill from each of the five kinds — a commit-message convention skill, a release procedure, the stop-and-ask checklist, a changelog format skill, and one domain skill for your stack — and you ran each on a real task, one per day. The convention skill produced a commit message your lead didn't send back. The release procedure ran the steps in your order, including the gate you'd never written down. The changelog skill made the next release notes the first ones in a year nobody had to fix. Two of the five visibly changed your output; the other three didn't, and you deleted those too. Five candidates, two keepers — and both keepers are now the things you can name at standup.
A month later, the kit has stabilized at four: the two keepers, the stop-and-ask checklist you rediscovered, and one domain skill you wrote yourself after the pattern clicked. Adding a skill is no longer a spree; it's a decision. Each one earns its place by surviving a real task, and each one you delete makes the survivors more visible. The agent is quieter, the sessions are sharper, and you finally understand why your teammate talks about his checklist like it's furniture: because it's the kind of thing you keep.
You'll end up with five skills you can name, five you've seen work, and a session that's quieter because the noise is gone. That's a better kit than the twelve you had on Monday. Next standup, when it's your turn, you'll have something to say.
And the thing you'll have to say is the proof that the whole approach works: not "I installed a bunch of skills," but "I have four skills and here's what each one does for me, and here's the one that stopped a release from breaking." One sentence per skill, all of them true, none of them borrowed from a README. That's what a curated kit sounds like. It's quieter than a spree, and it's worth more.
That's the honest end of the arc: twelve skills you couldn't name, four skills you can defend. The number went down. The value went up. And the way you install — by job, with evidence, keeping only what survives contact with real work — is a habit now, which means it compounds. The next skill you add will have to beat the four you already have. That's the standard. It's the right one.
The directory at qumge.com/en/skills aggregates community skills so you can shop by category instead of by repo. That is what curation is for: the sorting, not the hosting.
This week's plan, if you want the same outcome: pick one job from the five kinds that you actually do — not the one you wish you did, the one on your calendar. Find one skill for it, run it on a real task, and keep it only if your output changed. Then pick the next job. Five weeks, five jobs, and a kit you can describe out loud. That's the whole method, and it beats another Wednesday-night spree every time.
Top comments (2)
tr.ee/dev-to
Some comments may only be visible to logged-in visitors. Sign in to view all comments.