<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Janz</title>
    <description>The latest articles on DEV Community by Janz (@janzong).</description>
    <link>https://dev.to/janzong</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4124400%2Fd36a6924-4bf7-43e9-9d96-8d59fe598a4e.jpg</url>
      <title>DEV Community: Janz</title>
      <link>https://dev.to/janzong</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/janzong"/>
    <language>en</language>
    <item>
      <title>Your agent run passed. Can you prove it was allowed?</title>
      <dc:creator>Janz</dc:creator>
      <pubDate>Wed, 23 Sep 2026 03:03:42 +0000</pubDate>
      <link>https://dev.to/janzong/your-agent-run-passed-can-you-prove-it-was-allowed-3d5i</link>
      <guid>https://dev.to/janzong/your-agent-run-passed-can-you-prove-it-was-allowed-3d5i</guid>
      <description>&lt;p&gt;&lt;strong&gt;Short version:&lt;/strong&gt; Most agent teams can show a run. Few can show, in one re-runnable command, whether that run was allowed under a policy. I added a policy-as-code audit to &lt;code&gt;agent-lab-trust&lt;/code&gt;: it checks cost and call caps, the declared artifact contract, and forbidden markers, then emits a canonical &lt;code&gt;audit_hash&lt;/code&gt;. On 13 archived runs, the default contract passed 0/13 and a declared GenMentor contract passed 13/13. The data did not change. The contract did.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap
&lt;/h2&gt;

&lt;p&gt;Dashboards answer "what happened". Governance needs "what was allowed, and by which rule". A run can succeed and still violate a budget, write an unexpected artifact, or carry a marker that should never ship.&lt;/p&gt;

&lt;p&gt;The audit is deliberately boring:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;agent-lab-trust audit &amp;lt;run-root&amp;gt; &lt;span class="nt"&gt;--policy&lt;/span&gt; policy.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A policy can declare:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;max_cost_usd&lt;/code&gt; and &lt;code&gt;max_calls&lt;/code&gt; per run;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;required_artifacts&lt;/code&gt; plus &lt;code&gt;required_artifacts_mode: all|any&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;forbidden_markers&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The output is a finding list and a canonical &lt;code&gt;audit_hash&lt;/code&gt;. &lt;code&gt;deletion-proof --output &amp;lt;path&amp;gt;&lt;/code&gt; writes the deletion evidence used in the same policy flow.&lt;/p&gt;

&lt;h2&gt;
  
  
  The demonstration
&lt;/h2&gt;

&lt;p&gt;I ran the audit over 13 archived GenMentor runs.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Trust-layer default contract (&lt;code&gt;output/structured.json&lt;/code&gt; or &lt;code&gt;results/structured.json&lt;/code&gt;): &lt;strong&gt;0/13 passed&lt;/strong&gt;, &lt;code&gt;missing_artifact&lt;/code&gt; ×13, &lt;code&gt;cost_exceeded&lt;/code&gt; ×1.&lt;/li&gt;
&lt;li&gt;Declared GenMentor contract (&lt;code&gt;archive.json&lt;/code&gt; and &lt;code&gt;summary.json&lt;/code&gt;, mode &lt;code&gt;all&lt;/code&gt;, cap &lt;code&gt;5.00&lt;/code&gt;): &lt;strong&gt;13/13 passed&lt;/strong&gt;, no findings.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Audit hashes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;default structured contract: &lt;code&gt;2c122c4ff2c0d5eed11a0fc23b4ff717002864dbd941402e8fe18d6d17c96ce6&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;GenMentor contract: &lt;code&gt;368b75a06c0f5c205436d287881cfddcbabd94b7ad34a860e852fba59dad7ab5&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The one cost finding is real: &lt;code&gt;replay-8of8-20260920-a&lt;/code&gt; recorded &lt;code&gt;3.9529266&lt;/code&gt; under the trust-layer cap &lt;code&gt;0.60&lt;/code&gt;. Under the GenMentor policy cap &lt;code&gt;5.00&lt;/code&gt;, it passes. Caps are part of the contract too.&lt;/p&gt;

&lt;p&gt;The lesson is not "the audit is noisy". It is that an audit without an explicit run-family contract silently embeds one family's format. The policy has to name the contract, the caps, and the markers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;--rm&lt;/span&gt; ghcr.io/janzong/agent-lab-trust:rc2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or &lt;code&gt;bash scripts/reproduce.sh&lt;/code&gt;. Expected: &lt;code&gt;13 passed&lt;/code&gt; under both &lt;code&gt;TZ=UTC&lt;/code&gt; and &lt;code&gt;TZ=Asia/Shanghai&lt;/code&gt;, &lt;code&gt;report_hash&lt;/code&gt; &lt;code&gt;a841b192981fd7e7&lt;/code&gt;, deletion &lt;code&gt;audit_hash&lt;/code&gt; &lt;code&gt;4f0193abbd49a0f9&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;:rc2&lt;/code&gt; tag points at the latest published build; &lt;code&gt;SERIES.md&lt;/code&gt; pins the verification digest &lt;code&gt;sha256:2e178e63fff30e70ac68501cb17e5316097875932d1d7085e669ce4f8d5a105f&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not prove
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;no real governance deployment;&lt;/li&gt;
&lt;li&gt;no external reproduction yet;&lt;/li&gt;
&lt;li&gt;no real decision changed yet;&lt;/li&gt;
&lt;li&gt;the 13 runs are a private synthetic family, not published data.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The series
&lt;/h2&gt;

&lt;p&gt;This is piece 3 of a three-part line: &lt;strong&gt;write it down → test it → govern it&lt;/strong&gt;.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;agent-charters&lt;/code&gt;: what people actually tell agents.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;agent-lab-trust&lt;/code&gt;: whether an agent run is valid and reproducible.&lt;/li&gt;
&lt;li&gt;this audit: whether a run was allowed under a declared policy.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Index: &lt;a href="https://github.com/janzong/agent-lab-trust/blob/main/SERIES.md" rel="noopener noreferrer"&gt;https://github.com/janzong/agent-lab-trust/blob/main/SERIES.md&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Call
&lt;/h2&gt;

&lt;p&gt;Run the audit on a run family you own. Tell me where the policy and the artifacts disagree. I am also looking for 3 independent reproductions of the trust layer: &lt;a href="https://github.com/janzong/agent-lab-trust/issues/1" rel="noopener noreferrer"&gt;https://github.com/janzong/agent-lab-trust/issues/1&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>governance</category>
      <category>opensource</category>
    </item>
    <item>
      <title>We ran 2 vs 4 agents six times. Four agents cost 2.1 and did not improve success</title>
      <dc:creator>Janz</dc:creator>
      <pubDate>Tue, 22 Sep 2026 21:51:04 +0000</pubDate>
      <link>https://dev.to/janzong/we-ran-2-vs-4-agents-six-times-four-agents-cost-21x-and-did-not-improve-success-k98</link>
      <guid>https://dev.to/janzong/we-ran-2-vs-4-agents-six-times-four-agents-cost-21x-and-did-not-improve-success-k98</guid>
      <description>&lt;p&gt;&lt;strong&gt;Short version:&lt;/strong&gt; I ran a preregistered 2-agent versus 4-agent comparison on a deterministic task.&lt;br&gt;
Six live runs completed under hard caps. Both groups succeeded in exactly one of three repeats.&lt;br&gt;
The 4-agent group cost &lt;strong&gt;2.115×&lt;/strong&gt; more per run and per complete success. More agents were operationally&lt;br&gt;
viable; they were not better on this task.&lt;/p&gt;

&lt;h2&gt;
  
  
  The task
&lt;/h2&gt;

&lt;p&gt;The scenario is a deterministic public-repair contribution task:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;eight named participants;&lt;/li&gt;
&lt;li&gt;each participant chooses how many repair units to contribute;&lt;/li&gt;
&lt;li&gt;the task succeeds only if the total reaches a fixed threshold;&lt;/li&gt;
&lt;li&gt;participant coverage is forced, so every selected agent acts;&lt;/li&gt;
&lt;li&gt;no LLM reasoning is used for the choice; the model returns a strict JSON choice.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The point of forcing coverage was to avoid the earlier failure mode where one agent dominated every&lt;br&gt;
turn and the other participants never acted.&lt;/p&gt;

&lt;h2&gt;
  
  
  Protocol
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Groups&lt;/td&gt;
&lt;td&gt;2 agents / 3 steps; 4 agents / 5 steps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repeats&lt;/td&gt;
&lt;td&gt;3 per group&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Caps per run&lt;/td&gt;
&lt;td&gt;150 calls / USD 0.50 / 600s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Choice model&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;deepseek-flash&lt;/code&gt; on a Responses API contract&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Outcome&lt;/td&gt;
&lt;td&gt;machine-decidable &lt;code&gt;outcome.json&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Validity&lt;/td&gt;
&lt;td&gt;provenance, coverage, integer choices, arithmetic, caps, port closure, archive hashes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All six runs passed every validity gate. No parser failure, no budget breach, no port leak.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Agents&lt;/th&gt;
&lt;th&gt;Total / threshold&lt;/th&gt;
&lt;th&gt;Success&lt;/th&gt;
&lt;th&gt;Zero contributors&lt;/th&gt;
&lt;th&gt;Calls&lt;/th&gt;
&lt;th&gt;Total tokens&lt;/th&gt;
&lt;th&gt;Cost USD&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2-agent r1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2 / 3&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;10,060&lt;/td&gt;
&lt;td&gt;0.0269116&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2-agent r2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;3 / 3&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;11,070&lt;/td&gt;
&lt;td&gt;0.0313716&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2-agent r3&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0 / 3&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;10,255&lt;/td&gt;
&lt;td&gt;0.0281116&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4-agent r1&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;3 / 5&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;21,493&lt;/td&gt;
&lt;td&gt;0.0597772&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4-agent r2&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;2 / 5&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;22,679&lt;/td&gt;
&lt;td&gt;0.0641372&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4-agent r3&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;5 / 5&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;21,459&lt;/td&gt;
&lt;td&gt;0.0588042&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Group aggregates:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;2-agent&lt;/th&gt;
&lt;th&gt;4-agent&lt;/th&gt;
&lt;th&gt;Ratio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;complete successes&lt;/td&gt;
&lt;td&gt;1 / 3&lt;/td&gt;
&lt;td&gt;1 / 3&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mean calls&lt;/td&gt;
&lt;td&gt;17.00&lt;/td&gt;
&lt;td&gt;33.00&lt;/td&gt;
&lt;td&gt;1.941&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mean total tokens&lt;/td&gt;
&lt;td&gt;10,461.67&lt;/td&gt;
&lt;td&gt;21,877.00&lt;/td&gt;
&lt;td&gt;2.091&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mean cost USD&lt;/td&gt;
&lt;td&gt;0.028798&lt;/td&gt;
&lt;td&gt;0.060906&lt;/td&gt;
&lt;td&gt;2.115&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cost per complete success USD&lt;/td&gt;
&lt;td&gt;0.0863948&lt;/td&gt;
&lt;td&gt;0.1827186&lt;/td&gt;
&lt;td&gt;2.115&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mean zero contributors&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.333&lt;/td&gt;
&lt;td&gt;1.333&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What I take from it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;More agents are not automatically better. In this task they doubled cost and left more participants
contributing nothing on average.&lt;/li&gt;
&lt;li&gt;The result is a valid negative result, not an apparatus failure. Every run completed and every
artifact passed its gates.&lt;/li&gt;
&lt;li&gt;The 4-agent group did produce one complete success, so the mechanism is not broken. It is simply
not worth 2.115× the cost on this task.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What this does not show
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;It does not show that four agents are worse in general.&lt;/li&gt;
&lt;li&gt;It does not transfer to another task or to a real user.&lt;/li&gt;
&lt;li&gt;It says nothing about agent quality under human review, because no human scored these runs.&lt;/li&gt;
&lt;li&gt;The sample is three repeats per group on one synthetic task.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;p&gt;Two pieces are now public:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Method package:&lt;/strong&gt; &lt;a href="https://github.com/janzong/agent-lab-method" rel="noopener noreferrer"&gt;https://github.com/janzong/agent-lab-method&lt;/a&gt; (MIT, commit &lt;code&gt;6cc70156facb37baf23fe5fe57dad93d43502b91&lt;/code&gt;) — schema, synthetic GenMentor adapter, fixtures, tests;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trust layer:&lt;/strong&gt; &lt;a href="https://github.com/janzong/agent-lab-trust" rel="noopener noreferrer"&gt;https://github.com/janzong/agent-lab-trust&lt;/a&gt; (MIT, release &lt;code&gt;v0.1.0-rc2&lt;/code&gt;) — local-first run validation, hashed reports, and a synthetic deletion proof.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The full protocol and the archived runs stay private. The result is intentionally boring: a negative result with caps, hashes, and every failed run preserved.&lt;/p&gt;

&lt;p&gt;I am also looking for &lt;strong&gt;three independent reproductions by non-authors&lt;/strong&gt;. The trust layer guide expects &lt;code&gt;13 passed&lt;/code&gt; under both &lt;code&gt;TZ=UTC&lt;/code&gt; and &lt;code&gt;TZ=Asia/Shanghai&lt;/code&gt;, &lt;code&gt;report_hash&lt;/code&gt; &lt;code&gt;a841b192981fd7e7&lt;/code&gt;, and deletion &lt;code&gt;audit_hash&lt;/code&gt; &lt;code&gt;4f0193abbd49a0f9&lt;/code&gt;. If you run it and any hash differs, that is the most useful reply I can get. If you have a task where you believe more agents should win, that is the experiment I want to run next.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>llm</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How common is AGENTS.md, really? I sampled GitHub: 6.2% of active repos, 1.0% of all repos</title>
      <dc:creator>Janz</dc:creator>
      <pubDate>Sat, 19 Sep 2026 08:30:25 +0000</pubDate>
      <link>https://dev.to/janzong/how-common-is-agentsmd-really-i-sampled-github-62-of-active-repos-10-of-all-repos-1175</link>
      <guid>https://dev.to/janzong/how-common-is-agentsmd-really-i-sampled-github-62-of-active-repos-10-of-all-repos-1175</guid>
      <description>&lt;p&gt;&lt;strong&gt;Short version:&lt;/strong&gt; Every rate in my previous two posts had a denominator I picked myself — the 558 repos&lt;br&gt;
that already had an &lt;code&gt;AGENTS.md&lt;/code&gt;. That is a fine way to describe a corpus and a terrible way to answer&lt;br&gt;
"how common is this?". So I sampled GitHub two different ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;6.2%&lt;/strong&gt; of &lt;em&gt;active&lt;/em&gt; repos (pushed in the last 90 days, not a fork, not archived) contain an &lt;code&gt;AGENTS.md&lt;/code&gt;
— 51 of 817, 95% CI [4.8, 8.1]&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1.0%&lt;/strong&gt; of &lt;em&gt;all&lt;/em&gt; public repos do — 9 of 924, 95% CI [0.5, 1.8]&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same file, same counting rule, two numbers that differ by 6×. Which one you quote depends entirely on&lt;br&gt;
the question you are asking, and I had been quietly dodging that choice.&lt;/p&gt;

&lt;p&gt;Three things I did not expect, in order of how much they changed my mind:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;CLAUDE.md&lt;/code&gt; is at 5.4% of active repos.&lt;/strong&gt; Statistically indistinguishable from &lt;code&gt;AGENTS.md&lt;/code&gt;. If you
assumed one format won, it has not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;93% of public repos have not been pushed in 90 days&lt;/strong&gt;, 29% are forks, and &lt;strong&gt;8.3% are completely
empty&lt;/strong&gt;. "GitHub" as a population is mostly a graveyard, which is why the stock rate is so low.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Most &lt;code&gt;AGENTS.md&lt;/code&gt; files in the wild are tombstones.&lt;/strong&gt; Of the 9 files my population sample found,
&lt;strong&gt;7 were in repos that have not been pushed in three months.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;
  
  
  Why the denominator was missing
&lt;/h2&gt;

&lt;p&gt;Here is the trap I was in. To build the corpus I searched GitHub for repos containing an &lt;code&gt;AGENTS.md&lt;/code&gt;, then&lt;br&gt;
labeled what I found. Every percentage since — "85.7% of files prohibit things", "13.6% record a&lt;br&gt;
gotcha" — has a denominator of &lt;em&gt;files that already exist&lt;/em&gt;. Those numbers are real and I stand by them,&lt;br&gt;
but they cannot answer the question every reader actually has: &lt;strong&gt;should I write one of these?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For that you need a sample of repos drawn &lt;strong&gt;independently of whether they have the file&lt;/strong&gt;. That is a&lt;br&gt;
different sampling problem, and it needs two different frames, which I kept conflating:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;quantity&lt;/th&gt;
&lt;th&gt;question it answers&lt;/th&gt;
&lt;th&gt;frame&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;stock rate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"is this mainstream?"&lt;/td&gt;
&lt;td&gt;every public repo&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;active rate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"is this what working projects do?"&lt;/td&gt;
&lt;td&gt;repos pushed in the last 90 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;trend&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"is it spreading?"&lt;/td&gt;
&lt;td&gt;rates by repo creation year&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Mixing them produces confident nonsense, because the stock population is dominated by abandoned&lt;br&gt;
one-off repos and the active population is not.&lt;/p&gt;

&lt;p&gt;I wrote the decision rule down &lt;strong&gt;before&lt;/strong&gt; running anything, so I could not move the goalposts after&lt;br&gt;
seeing the number: &lt;strong&gt;&amp;lt;1% ⇒ describe it as an early-adopter curiosity; 1–5% ⇒ "early but measurable",&lt;br&gt;
and every rate must be labeled as in-corpus or ecosystem-wide; &amp;gt;10% ⇒ "standard practice".&lt;/strong&gt; It landed&lt;br&gt;
in the middle band, slightly high — so: not a curiosity, not a standard either.&lt;/p&gt;
&lt;h2&gt;
  
  
  Method, briefly
&lt;/h2&gt;

&lt;p&gt;Both frames use &lt;strong&gt;one tree call per repo&lt;/strong&gt; —&lt;br&gt;
&lt;code&gt;GET /repos/{owner}/{repo}/git/trees/HEAD?recursive=1&lt;/code&gt; — and check every path with a case-insensitive&lt;br&gt;
match on the filename. Recursive matters: a root-only check would miss &lt;code&gt;docs/AGENTS.md&lt;/code&gt; and friends and&lt;br&gt;
under-count. Roughly 1,900 repos, all responses cached, seed &lt;code&gt;20260918&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frame A — active rate, by creation cohort.&lt;/strong&gt; For each year 2010–2026 I picked one random slice of&lt;br&gt;
creation time, queried&lt;br&gt;
&lt;code&gt;created:&amp;lt;slice&amp;gt; pushed:&amp;gt;2026-06-20 fork:false archived:false&lt;/code&gt;, &lt;strong&gt;pulled every result&lt;/strong&gt; (rather than&lt;br&gt;
taking the top page, which is ranked by GitHub's relevance and would bias toward popular repos), then&lt;br&gt;
randomly sampled 50 repos from the slice. A 2026 week contains ~62,000 active repos, which exceeds the&lt;br&gt;
1,000-result search cap, so recent cohorts narrowed to a random day and then a random hour.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frame B — stock rate, unweighted.&lt;/strong&gt; GitHub search cannot give you a random sample of the population —&lt;br&gt;
there is no random sort, and ranking favors stars and activity. So I enumerated the ID space instead:&lt;br&gt;
binary-searched the current maximum repo ID (1,375,203,308), then sampled IDs uniformly and asked&lt;br&gt;
&lt;code&gt;GET /repositories/{id}&lt;/code&gt;. &lt;strong&gt;Only 35% of IDs correspond to an existing public repo&lt;/strong&gt; (the rest are&lt;br&gt;
deleted, private, or never existed), so 1,000 usable repos cost 2,920 probes.&lt;/p&gt;
&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Active repos (Frame A, n=817):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AGENTS.md                        6.2%   [4.8, 8.1]
CLAUDE.md                        5.4%   [4.0, 7.2]
.github/copilot-instructions.md  1.1%   [0.6, 2.1]
.cursorrules / .cursor/rules     0.7%   [0.3, 1.6]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;All public repos (Frame B, n=924):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;all public repos                 1.0%   [0.5, 1.8]
non-fork, non-archived           0.5%   [0.2, 1.4]
AGENTS.md ∩ CLAUDE.md             18 repos
CLAUDE.md only, no AGENTS.md      26 repos
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The number I got wrong twice
&lt;/h2&gt;

&lt;p&gt;In my corpus, &lt;strong&gt;59.1%&lt;/strong&gt; of repos that have an &lt;code&gt;AGENTS.md&lt;/code&gt; also have a &lt;code&gt;CLAUDE.md&lt;/code&gt;. In the wild it is&lt;br&gt;
&lt;strong&gt;35.3%&lt;/strong&gt;. Both are correct; they are answered by different populations, and only one of them is&lt;br&gt;
"typical".&lt;/p&gt;

&lt;p&gt;The reason for the gap is a selection effect I should have predicted: a repo that has one agent&lt;br&gt;
instruction file is already a repo whose author cares about agent tooling, so it is much more likely to&lt;br&gt;
have several. My corpus is a sample of the &lt;em&gt;enthusiastic&lt;/em&gt; end, and it over-represents multi-tool setups&lt;br&gt;
by about 1.7×. If I had quoted 59.1% as a base rate, I would have been describing my sample, not the&lt;br&gt;
world.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trend measurement failed, and I am not going to dress it up
&lt;/h2&gt;

&lt;p&gt;I wanted to show adoption rising by cohort and I could not measure it. In Frame A,&lt;br&gt;
&lt;code&gt;p_2026 / p_≤2022 = 0.67×&lt;/code&gt; — if anything, &lt;em&gt;older&lt;/em&gt; active repos are more likely to have the file.&lt;/p&gt;

&lt;p&gt;That number is not evidence that adoption is flat, because creation year and repo age are perfectly&lt;br&gt;
confounded. A repo created in 2010 that is still receiving pushes in 2026 is a &lt;strong&gt;survivor&lt;/strong&gt; — a project&lt;br&gt;
that lived long enough to accumulate conventions. A repo created in 2026 is mostly somebody's first&lt;br&gt;
weekend project. The cohort axis is really an age axis, and age predicts having-writers and having-time.&lt;/p&gt;

&lt;p&gt;Separating those would require reading commit history to find when each file was &lt;em&gt;added&lt;/em&gt;. I did not do&lt;br&gt;
that, so the honest deliverable here is "not measured", not "not spreading".&lt;/p&gt;

&lt;h2&gt;
  
  
  A GitHub API gotcha worth knowing
&lt;/h2&gt;

&lt;p&gt;Half a day of this project went into a bug that was not a bug. &lt;code&gt;403&lt;/code&gt; from the REST API has &lt;strong&gt;three&lt;/strong&gt;&lt;br&gt;
distinct meanings, and you have to read the body to tell them apart:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Primary rate limit&lt;/strong&gt; — &lt;code&gt;X-RateLimit-Remaining: 0&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Secondary rate limit&lt;/strong&gt; — burst/concurrency. GitHub's docs say it plainly: &lt;em&gt;make requests for a
single user serially&lt;/em&gt;. Three threads was enough to get me permanently throttled, while serial
requests with connection reuse ran at ~8/s without complaint.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A single repo blocked by GitHub&lt;/strong&gt; —
&lt;code&gt;{"message":"Repository access blocked","block":{"reason":"tos"}}&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I had classified (3) as (2), so my code kept sleeping and retrying the same blocked repo, forever. If&lt;br&gt;
you write a bulk GitHub crawler, put that string in your error handling; it is stable, it is not a rate&lt;br&gt;
limit, and it will silently eat your retry budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this changes about my own claims
&lt;/h2&gt;

&lt;p&gt;I have to soften things I said earlier:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"&lt;code&gt;AGENTS.md&lt;/code&gt; is the emerging standard"&lt;/strong&gt; — no. &lt;strong&gt;6.2%&lt;/strong&gt; of active repos is a real practice, not a
standard. It &lt;em&gt;is&lt;/em&gt; roughly 6–9× more common than the vendor-specific alternatives, which is a
measurable reason to keep the word "format" instead of a vendor name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"&lt;code&gt;CLAUDE.md&lt;/code&gt; is just a companion file"&lt;/strong&gt; — that was a conditional rate described as if it were a
base rate. Unconditionally, the two are neck and neck.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"adoption is spreading"&lt;/strong&gt; — unmeasured, see above.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What survives: &lt;strong&gt;it is a habit of active projects&lt;/strong&gt; (6.2%) &lt;strong&gt;rather than something the population does&lt;/strong&gt;&lt;br&gt;
(1.0%). Those two sentences imply completely different advice, and I could not tell them apart until I&lt;br&gt;
sampled for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment I have not run
&lt;/h2&gt;

&lt;p&gt;There is still a hole underneath all of this, and it is the one that matters: &lt;strong&gt;does any of it work?&lt;/strong&gt;&lt;br&gt;
Every number in this project — mine and everyone else's — is descriptive. Nobody has shown that a repo&lt;br&gt;
with an &lt;code&gt;AGENTS.md&lt;/code&gt; produces better outcomes than the same repo without one, because that requires a&lt;br&gt;
controlled task, a blind judge, and an effect size, not a sample.&lt;/p&gt;

&lt;p&gt;I wrote the design for that experiment (three arms, including a placebo arm that gets an equal-length&lt;br&gt;
unrelated document, so "more context" and "this document" can be told apart) but I deliberately did not&lt;br&gt;
run it yet. The sample sizes are brutal and the honest outcome is probably "we could not detect it".&lt;/p&gt;

&lt;p&gt;If you have actually noticed a charter changing a decision — a rule that stopped you from doing&lt;br&gt;
something you would otherwise have done — that is the data I cannot generate myself, and it is worth&lt;br&gt;
more to me than a star.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;agent-charters              &lt;span class="c"&gt;# the CLI&lt;/span&gt;
&lt;span class="c"&gt;# the sampling code is in the repo, not the package:&lt;/span&gt;
git clone https://github.com/janzong/agent-charters
&lt;span class="nb"&gt;cd &lt;/span&gt;agent-charters
.venv/bin/python work/prevalence.py active &lt;span class="nt"&gt;--per-gen&lt;/span&gt; 50   &lt;span class="c"&gt;# ~10 min, search-rate-limited&lt;/span&gt;
.venv/bin/python work/prevalence.py report
.venv/bin/python work/prevalence.py stock  &lt;span class="nt"&gt;--n&lt;/span&gt; 1000       &lt;span class="c"&gt;# ~40 min&lt;/span&gt;
.venv/bin/python work/prevalence.py report-stock
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seed &lt;code&gt;20260918&lt;/code&gt;, every response cached, so &lt;code&gt;report&lt;/code&gt; is instant and costs nothing after the first run.&lt;br&gt;
The full write-up with every caveat I could think of — including the two frames disagreeing (6.2% vs&lt;br&gt;
1.8% on a 57-repo active sub-sample; the honest answer is a range of roughly 2–6%) — is in&lt;br&gt;
&lt;code&gt;work/audit/prevalence.md&lt;/code&gt; in the repo.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/janzong/agent-charters" rel="noopener noreferrer"&gt;https://github.com/janzong/agent-charters&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>github</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I looked at 558 AGENTS.md files: here's a 5-minute check for yours</title>
      <dc:creator>Janz</dc:creator>
      <pubDate>Mon, 14 Sep 2026 13:25:06 +0000</pubDate>
      <link>https://dev.to/janzong/i-looked-at-558-agentsmd-files-heres-a-5-minute-check-for-yours-5cih</link>
      <guid>https://dev.to/janzong/i-looked-at-558-agentsmd-files-heres-a-5-minute-check-for-yours-5cih</guid>
      <description>&lt;p&gt;&lt;strong&gt;Short version:&lt;/strong&gt; I labeled 558 public &lt;code&gt;AGENTS.md&lt;/code&gt; files against a 9-category taxonomy. The measured&lt;br&gt;
base rates say something boring and useful — almost every file &lt;strong&gt;prohibits&lt;/strong&gt; things (85.7%) and lists&lt;br&gt;
&lt;strong&gt;build/test commands&lt;/strong&gt; (82.8%), while almost none of them record a &lt;strong&gt;gotcha&lt;/strong&gt; (13.6%). Two of the nine&lt;br&gt;
slots are nearly empty across the whole corpus: &lt;code&gt;gotchas&lt;/code&gt; and &lt;code&gt;agent_meta&lt;/code&gt; (rules about the agent itself,&lt;br&gt;
25.8%).&lt;/p&gt;

&lt;p&gt;Then I ran the same ruler over two big, well-maintained files. Both missed &lt;code&gt;gotchas&lt;/code&gt;. So here is a&lt;br&gt;
five-minute check you can run on your own file, and the exact numbers behind it.&lt;/p&gt;
&lt;h2&gt;
  
  
  The base rates
&lt;/h2&gt;

&lt;p&gt;Measured on 516 substantive files (558 collected, the rest were one-line pointers):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;boundaries     85.7%   what must never be done
build_test     82.8%   the commands CI runs
workflow       67.1%   commit format, branches, release steps
structure      59.1%   layout, where new code belongs
style          54.5%   naming, formatting — or a pointer to the config that enforces it
environment    45.0%   toolchain versions, required env vars
overview       32.2%   one paragraph: what this is, what it deliberately is not
agent_meta     25.8%   rules about the agent: tone, when to ask first
gotchas        13.6%   pitfalls that are NOT derivable from the code
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The shape is not surprising once you see it as a genre: an &lt;code&gt;AGENTS.md&lt;/code&gt; is usually written &lt;em&gt;defensively&lt;/em&gt;,&lt;br&gt;
as a list of things not to break. The file that would actually save you time is the one almost nobody&lt;br&gt;
writes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two receipts
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;compare&lt;/code&gt; prints your file's coverage next to the corpus baseline. Two real examples from the corpus:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;file&lt;/th&gt;
&lt;th&gt;size&lt;/th&gt;
&lt;th&gt;sections&lt;/th&gt;
&lt;th&gt;coverage&lt;/th&gt;
&lt;th&gt;missing&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;langchain-ai/deepagents&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10 KB&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7/9&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;overview&lt;/code&gt;, &lt;code&gt;gotchas&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;openai/openai-agents-python&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;34 KB&lt;/td&gt;
&lt;td&gt;27&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;6/9&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;style&lt;/code&gt;, &lt;code&gt;agent_meta&lt;/code&gt;, &lt;code&gt;gotchas&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both are good files. The 34 KB one is one of the more thorough agent-instruction files in the corpus —&lt;br&gt;
27 sections, 19 separate boundary markers. It still has nothing in it that you could only learn by&lt;br&gt;
running the thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why gotchas are rare (and why that is not laziness)
&lt;/h2&gt;

&lt;p&gt;You can only write a gotcha &lt;em&gt;after&lt;/em&gt; being bitten by it — and by the time you have been bitten, the&lt;br&gt;
temptation is to &lt;strong&gt;fix the thing&lt;/strong&gt; rather than write the sentence down. The fix is visible in the code;&lt;br&gt;
the sentence is a liability nobody wants to maintain.&lt;/p&gt;

&lt;p&gt;There is a second, worse failure mode. I sampled 347 entries from the &lt;code&gt;Gotchas&lt;/code&gt; / &lt;code&gt;Common Pitfalls&lt;/code&gt; /&lt;br&gt;
&lt;code&gt;Troubleshooting&lt;/code&gt; sections in the corpus (an earlier snapshot, 507 files) and hand-labeled 120 of them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;58%&lt;/strong&gt; are readable from the repo itself (interface contracts, platform limits, build requirements)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;34%&lt;/strong&gt; are not pitfalls at all — they are generic advice ("remember to install dependencies", "don't
commit &lt;code&gt;.env&lt;/code&gt;"), the same sentence you would write for any project&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;8%&lt;/strong&gt; are genuinely experience-only: upstream/third-party behaviour, past incidents, and the places
where the docs disagree with the code&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the section is rare, and a third of what does live there is filler. The 8% is the part worth&lt;br&gt;
handing to an agent, and it cannot be generated from a reading of the repository. It has to come from&lt;br&gt;
a person who was there.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five-minute check
&lt;/h2&gt;

&lt;p&gt;No tool needed. Ask these five questions about your own file:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Are the commands copy-pasteable?&lt;/strong&gt; Not "run the tests" — the actual command CI runs, with the
working directory. If your README says one port and production uses another, say so (that mistake
is in the corpus, in a file that otherwise looks complete).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does it name what must never be committed or never touched?&lt;/strong&gt; This is the one thing the corpus
does well (85.7%) — check that yours names the &lt;em&gt;tempting&lt;/em&gt; case, not the obvious one. "Don't commit
secrets" is obvious; "don't hand-edit the production database to fix a row, use the backfill script"
is a boundary that will actually stop someone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does it say anything about the agent's own behaviour?&lt;/strong&gt; Only 25.8% do. Tone, when to stop and ask,
which actions need explicit approval, what must not leave the machine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is there at least one sentence that is not derivable from the code?&lt;/strong&gt; If every line in your file
could have been written by reading the repo, the file is documentation, not a charter. This is the
gotcha test.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do the paths it points at exist?&lt;/strong&gt; Measured: 49% of files route to another file, and 15% point at a
knowledge store or rules directory. A pointer to a file that moved is worse than no pointer — an
agent will go looking, and will read whatever it finds there as authoritative.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  If you want the baseline instead of the feeling
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;agent-charters                 &lt;span class="c"&gt;# 0.4.1, on PyPI — no runtime deps, ~110 KB wheel&lt;/span&gt;
agent-charters compare path/to/AGENTS.md   &lt;span class="c"&gt;# coverage vs the 558-file baseline, plus gaps&lt;/span&gt;
agent-charters brief                       &lt;span class="c"&gt;# the checklist + a paste-ready prompt&lt;/span&gt;
agent-charters refs path/to/AGENTS.md      &lt;span class="c"&gt;# external pointers and dangling references&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;(Updated 2026-09-19 — when this was first posted the only way in was a clone, and this block said so.&lt;br&gt;
It is on PyPI now: verified in a clean virtualenv on a default index, install takes a couple of&lt;br&gt;
seconds, because the corpus ships as gzipped JSONL and &lt;code&gt;pandas&lt;/code&gt; / &lt;code&gt;pyarrow&lt;/code&gt; are optional extras.&lt;br&gt;
Source, if you would rather read or pin it: github.com/janzong/agent-charters — CN mirror&lt;br&gt;
gitee.com/janzong/agent-charters.)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;compare&lt;/code&gt; is the one that answers question 4 in aggregate. It also does something I did not expect:&lt;br&gt;
when I used &lt;code&gt;brief&lt;/code&gt;'s prompt to write a charter for a real project, &lt;code&gt;compare&lt;/code&gt; flagged coverage I had&lt;br&gt;
skipped — and one of the nine slots it missed was the &lt;em&gt;name of the slot itself&lt;/em&gt;, which is a bug in my&lt;br&gt;
taxonomy, not in the file. That is the kind of thing a rule-based labeler gives you: you can point at&lt;br&gt;
the pattern that fired and argue with it.&lt;/p&gt;

&lt;p&gt;There is no LLM in the labeling loop. Every label is recomputable and arguable, which is the point —&lt;br&gt;
if you disagree with a label, you can find the rule that produced it and overrule it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is not
&lt;/h2&gt;

&lt;p&gt;I do not want to oversell the numbers, so:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The classifier scores &lt;strong&gt;92% precision / 70% recall&lt;/strong&gt; on a 55-file held-out English set, and
&lt;strong&gt;88% / 73%&lt;/strong&gt; on 50 held-out Chinese files. The recall number is the honest one: it misses roughly
&lt;strong&gt;three in ten&lt;/strong&gt; of the labels it should have produced. &lt;code&gt;gotchas&lt;/code&gt; and &lt;code&gt;agent_meta&lt;/code&gt; are the weakest
slots in both languages.&lt;/li&gt;
&lt;li&gt;The held-out sets were labeled by &lt;strong&gt;one person (me)&lt;/strong&gt;. No second annotator, no inter-annotator
agreement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Coverage is a process metric, not a quality metric.&lt;/strong&gt; In a 3-repo test, a checklist that names all
nine slots pushed a generator from 4–5 categories to 9/9 — and filling all nine slots is not the same
as writing a good file. It is a prompt for the questions, not a grade.&lt;/li&gt;
&lt;li&gt;The labels and the rates come from public files and a rule-based classifier, not from a language
model. The one LLM in this story is the generator in the 3-repo test, which is why that number is
reported as n=3.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The ask
&lt;/h2&gt;

&lt;p&gt;The weakest part of this project is that the only person who has ever tested it is its author. If you&lt;br&gt;
have an &lt;code&gt;AGENTS.md&lt;/code&gt; (or a &lt;code&gt;CLAUDE.md&lt;/code&gt;, or a &lt;code&gt;.cursorrules&lt;/code&gt;) on a real project, run &lt;code&gt;compare&lt;/code&gt; on it and&lt;br&gt;
tell me what it gets wrong — the file, the label, or the baseline rate. A wrong label on your file is&lt;br&gt;
worth more to me than a star.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/janzong/agent-charters" rel="noopener noreferrer"&gt;https://github.com/janzong/agent-charters&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I labeled 558 AGENTS.md files. Here's what they say — and what almost nobody writes down</title>
      <dc:creator>Janz</dc:creator>
      <pubDate>Mon, 14 Sep 2026 12:23:13 +0000</pubDate>
      <link>https://dev.to/janzong/i-labeled-558-agentsmd-files-heres-what-they-say-and-what-almost-nobody-writes-down-34gb</link>
      <guid>https://dev.to/janzong/i-labeled-558-agentsmd-files-heres-what-they-say-and-what-almost-nobody-writes-down-34gb</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — I collected &lt;strong&gt;558 &lt;code&gt;AGENTS.md&lt;/code&gt; files&lt;/strong&gt; from public repos and labeled each one against a 9-category&lt;br&gt;
taxonomy with a &lt;strong&gt;rule-based&lt;/strong&gt; classifier (no LLM in the loop, so it is auditable and recomputable). Then I&lt;br&gt;
blind-labeled held-out samples and compared: &lt;strong&gt;92% precision / 70% recall&lt;/strong&gt; on 55 English files,&lt;br&gt;
&lt;strong&gt;88% / 73%&lt;/strong&gt; on 50 Chinese files. The most common categories are prohibitions (&lt;strong&gt;85.7%&lt;/strong&gt;) and build/test&lt;br&gt;
commands (&lt;strong&gt;82.8%&lt;/strong&gt;). The rarest: &lt;strong&gt;gotchas (13.6%)&lt;/strong&gt; and instructions about how the agent itself should behave&lt;br&gt;
(&lt;strong&gt;25.8%&lt;/strong&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Why bother
&lt;/h2&gt;

&lt;p&gt;Almost every discussion about &lt;code&gt;AGENTS.md&lt;/code&gt; is anecdote-led: &lt;em&gt;my&lt;/em&gt; repo's file works, &lt;em&gt;my&lt;/em&gt; agent ignores it,&lt;br&gt;
a good one is a model upgrade, a bad one is worse than nothing. All of that may be true — but nobody&lt;br&gt;
seems to have the distribution. So I built it: snapshot of 558 files from 558 public repos&lt;br&gt;
(2026-09-10, 5.3 MB, &lt;strong&gt;516 usable for statistics&lt;/strong&gt;), labeled, versioned, and published with the tooling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Method, in one paragraph
&lt;/h2&gt;

&lt;p&gt;Nine categories: &lt;code&gt;boundaries&lt;/code&gt;, &lt;code&gt;build_test&lt;/code&gt;, &lt;code&gt;workflow&lt;/code&gt;, &lt;code&gt;structure&lt;/code&gt;, &lt;code&gt;style&lt;/code&gt;, &lt;code&gt;environment&lt;/code&gt;, &lt;code&gt;overview&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;agent_meta&lt;/code&gt; (rules about the AI itself), &lt;code&gt;gotchas&lt;/code&gt;. Labeling is done by pattern rules over headings and&lt;br&gt;
body text — deliberately, because a rule set can be read, argued with, and re-run, and every number below&lt;br&gt;
can be recomputed from the released dataset. I then measured how well the rules match a human reading:&lt;br&gt;
&lt;strong&gt;100 files in-sample&lt;/strong&gt; (upper bound, 90%/75%) and two held-out sets I had never tuned against —&lt;br&gt;
55 English (92%/70%) and 50 Chinese (88%/73%). Held-out numbers use the &lt;em&gt;conservative&lt;/em&gt; reading&lt;br&gt;
(items I was unsure about count as classifier errors).&lt;/p&gt;

&lt;h2&gt;
  
  
  Five things the numbers say
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Two categories dominate — and they are tied
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;category&lt;/th&gt;
&lt;th&gt;share of 516 files&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;boundaries&lt;/code&gt; (what you must never do)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;85.7%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;build_test&lt;/code&gt; (install/build/test/CI commands)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;82.8%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;workflow&lt;/code&gt; (branching, commits, review, release)&lt;/td&gt;
&lt;td&gt;67.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;structure&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;59.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;style&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;54.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;environment&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;45.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;overview&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;32.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;agent_meta&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;25.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gotchas&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;13.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 2.9 pp gap between the top two is &lt;em&gt;smaller&lt;/em&gt; than the known false-positive rate (~3%) of the&lt;br&gt;
prohibition pattern — so the honest statement is &lt;strong&gt;tied for first&lt;/strong&gt;, not "prohibitions beat build commands".&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Nobody writes down their scars
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;gotchas&lt;/code&gt; is dead last at 13.6%. Worse: when people &lt;em&gt;do&lt;/em&gt; open a "known issues" section, a third of it&lt;br&gt;
isn't a gotcha. I hand-read 120 items from those sections:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;58%&lt;/strong&gt; were readable straight from the repo (config, code, README mismatch),&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;34%&lt;/strong&gt; were not gotchas at all (generic advice: "remember to install dependencies"),&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;8%&lt;/strong&gt; needed experience or the outside world (OS behavior, an upstream outage, yesterday's incident).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That 8% is the part an agent can never derive from the code — and it is exactly the part that is&lt;br&gt;
almost never written down.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The slot language models skip is &lt;code&gt;workflow&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;In a controlled experiment (11 repos × 3 prompt styles), prompts that listed topics explicitly produced&lt;br&gt;
&lt;strong&gt;9/9 categories&lt;/strong&gt;, while prompts that left the slots implicit skipped &lt;code&gt;workflow&lt;/code&gt; in &lt;strong&gt;11 out of 11&lt;/strong&gt; files.&lt;br&gt;
Point at &lt;code&gt;workflow&lt;/code&gt; by name and it appears &lt;strong&gt;3/3&lt;/strong&gt; times, with real content. The gap is not knowledge,&lt;br&gt;
it is &lt;em&gt;questions&lt;/em&gt; — which is why I turned the corpus distribution into a checklist tool.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Your weakest category depends on the language
&lt;/h3&gt;

&lt;p&gt;English files fail differently from Chinese ones. English: &lt;code&gt;gotchas&lt;/code&gt; recall 32–38% — the classifier&lt;br&gt;
misses casual "watch out" prose. Chinese: &lt;code&gt;agent_meta&lt;/code&gt; recall &lt;strong&gt;26%&lt;/strong&gt; — Chinese files express agent rules&lt;br&gt;
in the second person ("you are the dispatcher, not the executor"), and the body-pattern rules for that&lt;br&gt;
category are entirely English, so the whole style is invisible to them. File-level exact agreement&lt;br&gt;
(9/9 categories identical) is &lt;strong&gt;12%&lt;/strong&gt; in both languages.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Half of these files are entry points, not documentation
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;49%&lt;/strong&gt; point to some other file; &lt;strong&gt;15%&lt;/strong&gt; route to a knowledge or rules directory. That's a structural&lt;br&gt;
fact about the format, and it means "does this repo have an &lt;code&gt;AGENTS.md&lt;/code&gt;?" is a much weaker question than&lt;br&gt;
"what is actually in it".&lt;/p&gt;

&lt;h2&gt;
  
  
  The tool
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;agent-charters

agent-charters brief      &lt;span class="c"&gt;# checklist of the 9 slots + a paste-ready prompt&lt;/span&gt;
agent-charters compare your-AGENTS.md   &lt;span class="c"&gt;# your coverage vs the 558-file baseline&lt;/span&gt;
agent-charters refs your-AGENTS.md      &lt;span class="c"&gt;# does your file point at paths that exist&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Honest note: &lt;code&gt;compare&lt;/code&gt; is a &lt;strong&gt;checklist, not an oracle&lt;/strong&gt;. It warned me that one of the nine categories was&lt;br&gt;
missing from a file I wrote myself — it was actually present, but the heading used the tool's own slot name&lt;br&gt;
instead of natural language. That is documented in the repo (along with the exact experiment) rather than&lt;br&gt;
quietly patched, because a tool that tells you "you're missing X" should be checked by a human.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is not
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Rule-based labels, not per-file human labels.&lt;/strong&gt; Precision/recall above are the honest measures; the
per-category numbers in the dataset carry that error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not a representative sample of GitHub.&lt;/strong&gt; Repos were found through AI/agent topics and Chinese keyword
search; the Chinese set came out 97% Chinese by construction, which says nothing about GitHub's language mix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single annotator.&lt;/strong&gt; The blind labeling was done by one model-driven annotator, not multiple raters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A snapshot, not a trend.&lt;/strong&gt; A baseline of file hashes is stored so that a future re-crawl can measure
how these files change — that measurement doesn't exist yet.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Code + tooling: &lt;a href="https://github.com/janzong/agent-charters" rel="noopener noreferrer"&gt;https://github.com/janzong/agent-charters&lt;/a&gt; (mirror: &lt;a href="https://gitee.com/janzong/agent-charters" rel="noopener noreferrer"&gt;https://gitee.com/janzong/agent-charters&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Dataset &lt;strong&gt;v0.5&lt;/strong&gt; release + methodology, limitations and every number above: see &lt;code&gt;LIMITATIONS.md&lt;/code&gt; in the repo&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The ask
&lt;/h2&gt;

&lt;p&gt;I'm looking for &lt;strong&gt;2–3 people who are not me&lt;/strong&gt; to run &lt;code&gt;compare&lt;/code&gt; on an &lt;code&gt;AGENTS.md&lt;/code&gt; they actually maintain and&lt;br&gt;
tell me where it's wrong — missing a category you clearly have, or claiming one you don't. That is the one&lt;br&gt;
piece of evidence this project doesn't have yet: an external user. Issues and comments are both fine.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
      <category>data</category>
    </item>
  </channel>
</rss>
