How to evaluate whether agents find and use your skill, how it affects their behavior, and whether it helps or harms the use of your product.
Agent skills introduce a new layer between coding agents and developer products. Instead of leaving an agent to piece everything together from documentation or other hints, a skill can give it targeted instructions, scripts, and resources when they become relevant. A skill goes beyond documentation by helping shape how an agent discovers, navigates, and uses your product.
Of course, a SKILL.md doesn’t necessarily mean the agent will find it, use it, or use your product more effectively just because it exists. Early evidence suggests skills can change agent behavior, but not always for the better. That raises a more useful set of questions about how coding agents discover and use skills, starting with whether they find them at all.
Can the agent find the skill?
Before we can ask whether a skill helps, the agent has to know it exists. And that starts with something surprisingly mundane: where the skill lives.
A SKILL.md sitting somewhere in a repository isn't automatically discoverable. The exact location still depends on the agent and runtime. Claude Code, for example, supports project skills under .claude/skills/skill-name/SKILL.md, while OpenAI Codex uses the same basic skill format, discovers repository skills in .agents/skills, and can also load skills bundled in plugins. When testing a skill, the first thing to verify is that it has been installed somewhere the agent is designed to discover it.
SKILL.md Description
Once the skill is available, there's still another step. Both OpenAI and Anthropic use progressive disclosure rather than putting every installed skill’s full instructions into context at once. The agent initially sees lightweight metadata, particularly the skill’s name and description, and uses that information to decide whether the skill is relevant enough to load. That makes the skill description especially important. Both OpenAI and Anthropic recommend using skill metadata to explain what it does and when the agent should use it.
Ask the right questions
Test skill activation at different levels of specificity:
- Name the skill: Tests whether the agent can use the skill when explicitly directed. Example: “Use the Quantiles skill to run an evaluation.”
- Describe the task: Tests whether the agent identifies the relevant skill from the task. Example: “Configure authentication for this application.”
- Describe the goal: Tests whether the agent infers the skill’s relevance from the developer’s goal. Example: “Get this service connected so I can start evaluating agents.”
That’s much closer to how we actually use coding agents. We tell them what we want done, not which skill, MCP server, documentation page, or tool to use along the way. OpenAI recommends testing this range, from explicit and contextual requests to cases where the skill shouldn’t activate at all.
Does the skill make the product easier for agents to use?
Skill discovery is only one part of the evaluation. You also want to know whether the skill changes the agent’s ability to use your product successfully. Compare the same task with and without the skill while holding the agent, repository, documentation, tools, and environment constant. This isolates the skill as the primary experimental variable and makes differences in agent behavior and outcomes easier to interpret.
So how do we know if the skill actually helped? The execution trace gives us a few clues. Looking at these signals can show how the skill changed the agent’s behavior and whether it made the workflow more efficient, including its time to completion, or more reliable.
Measuring the Effect of a Skill
| Measurement | What to look for |
|---|---|
| Task success | Did the agent reach the verified outcome, and was success more consistent with the skill? |
| Tool calls | Did the skill change how much work the agent needed to complete the task? |
| Tool errors | Did the skill help the agent avoid incorrect commands, arguments, APIs, or operations? |
| Repeated tool calls | Did the skill reduce unnecessary retries or loops, or help the agent recover more efficiently? |
| Latency | Did the skill shorten the path to completion or introduce additional overhead? |
| Token usage | Did the skill reduce searching and reasoning, or add more context and processing? |
| Human intervention | Did the agent need less clarification, redirection, or assistance to finish the task? |
| Skill activation | Did the agent load the skill, and at what point in the task did it become useful? |
Where SKILL.md helps and where it hurts
Skills may affect agent performance differently across task types. For installation and authentication, a defined product workflow can reduce ambiguity and help the agent avoid incorrect tool calls or unnecessary steps. Recovering from errors often requires more flexibility because the appropriate next action depends on the error and the current environment. Advanced workflows may benefit more from guidance about product conventions and important decision points than from detailed step-by-step instructions.
The effect can also vary by coding agent. One agent may already navigate your documentation well and gain little from additional guidance, while another benefits significantly from the structure. That can change as models improve too. The goal isn't for the skill to improve everything. It's to find where a little extra guidance makes a meaningful difference and where the agent is already better off figuring things out on its own.
The table below shows examples of where skills may help or hurt and how the right level of guidance can reduce uncertainty while leaving room for the agent to adapt.
How SKILL.md Affects Different Tasks
| More likely to help | More likely to hurt |
|---|---|
| There is a preferred path. Several approaches are possible, but one is recommended or supported. | The path is already obvious. The agent can infer the correct approach from the environment and normal documentation. |
| The task has important sequencing. Steps need to happen in a particular order. | The task is simple or direct. Extra instructions add process without adding useful information. |
| There are non-obvious prerequisites. Authentication, setup, state, dependencies, or permissions must exist first. | The skill repeats easily discoverable information. It mostly duplicates CLI help, API schemas, or straightforward documentation. |
| There are product-specific decisions. The agent needs to know which API, command, configuration, or workflow is appropriate in a particular situation. | The skill overprescribes the solution. It tells the agent exactly how to solve situations where adapting to the environment would be better. |
| There are important constraints. Some approaches appear reasonable but are unsupported, unsafe, expensive, destructive, or likely to produce the wrong state. | The instructions reduce flexibility. The agent keeps following the documented path even when errors or environment state suggest it should change course. |
| Success is not obvious. A successful command or API response does not necessarily mean the task is complete, and the agent needs to know what to verify. | The skill adds unnecessary verification or steps. The agent already has reliable signals that the task succeeded. |
| The product has unusual conventions. Important behavior differs from what an agent might reasonably infer from similar products. | The skill describes conventional behavior. It spends context explaining patterns the agent already understands. |
| A small amount of guidance generalizes. A rule or decision principle can help across many variations of the task. | The skill is tutorial-shaped. It encodes one specific trajectory that does not transfer well to unfamiliar situations. |
| The guidance is current and stable. The instructions describe behavior unlikely to change unexpectedly. | The guidance becomes stale easily. Commands, parameters, defaults, or workflows change frequently and can conflict with the current product. |
| The skill is concise and targeted. It supplies information the agent needs and leaves the rest to the agent. | The skill is large and noisy. Relevant guidance competes with examples, edge cases, explanations, and instructions that do not apply to the current task. |
SKILL.md should evolve with your product
Changes to your product and the services it depends on, including AI models, can change how agents experience it. Keep a stable set of evaluation tasks for your most important workflows and use them as regression tests when the product, documentation, or SKILL.md changes. When a new agent failure appears, turn it into another evaluation task so you can test for it going forward.
Coding agents are constantly changing too. Guidance that helps today may become unnecessary as models improve, while new agent behavior may expose problems you've never seen before. That makes agent evaluation an ongoing part of building your product. Keep testing important workflows as both the product and the agents using it change. Sometimes that will lead you to update SKILL.md. Other times, the better fix will be in your documentation, API, error messages, or the product itself.
Top comments (0)