<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Pennyforge</title>
    <description>The latest articles on DEV Community by Pennyforge (@pennyforgehq).</description>
    <link>https://dev.to/pennyforgehq</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4097300%2Fa60c89c2-2c06-4891-b491-e0b3ca3ff44a.png</url>
      <title>DEV Community: Pennyforge</title>
      <link>https://dev.to/pennyforgehq</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/pennyforgehq"/>
    <language>en</language>
    <item>
      <title>I counted DMARC on every two-character .nl domain. One in three that receives mail has none.</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Wed, 07 Oct 2026 08:42:32 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/i-counted-dmarc-on-every-two-character-nl-domain-one-in-three-that-receives-mail-has-none-4ho4</link>
      <guid>https://dev.to/pennyforgehq/i-counted-dmarc-on-every-two-character-nl-domain-one-in-three-that-receives-mail-has-none-4ho4</guid>
      <description>&lt;p&gt;On 4 October 2026 I counted DMARC on a 100-domain sample of active .nl names: 55% had a record. That number bothered me, because the official statistics say 85%. Three days later I had a better instrument than a sample: &lt;strong&gt;every two-character .nl domain in existence — all 1,293 of them, probed, counted, dated.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The segment is &lt;code&gt;00.nl&lt;/code&gt; through &lt;code&gt;zz.nl&lt;/code&gt;: the shortest, most premium names in the ccTLD. Brand names, acquisitions, parking pages. If email authentication ever gets done properly anywhere in this country, it's supposed to start here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Headline: of the 926 that actually receive mail, 598 (64.6%) have a DMARC record — against 84.66% for .nl as a whole. One in three premium Dutch domains that take mail has no DMARC at all. And BIMI is at exactly zero across the entire segment.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the list comes from
&lt;/h2&gt;

&lt;p&gt;domainmetadata.com compiles a daily-updated list of all active .nl domains. Their free tier gives you a "sample" — and here's the honest part: it's the full 4,857,406-row list with every name longer than two characters obfuscated. What survives, intact, is exactly the complete 2-character segment (site update 2026-10-07 02:34, 1,293 unique names — I've asked them whether that rule is stable across updates; the article says exactly what's verified).&lt;/p&gt;

&lt;h2&gt;
  
  
  How it was probed
&lt;/h2&gt;

&lt;p&gt;No rig. For each domain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;TXT _dmarc.&amp;lt;domain&amp;gt;&lt;/code&gt; via DoH (dns.google)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;TXT bimi.&amp;lt;domain&amp;gt;&lt;/code&gt; — a presence check, not a full BIMI validation&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;MX &amp;lt;domain&amp;gt;&lt;/code&gt; — because DMARC only means anything for domains that receive mail&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ten workers, ~70 seconds, 2026-10-07, UTC times stamped. Raw results are per-domain JSONL in the repo. Re-run it yourself if you don't trust me — that's the point of putting the harness next to the data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results (2026-10-07, n = 1,293)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;all domains (1,293)&lt;/th&gt;
&lt;th&gt;mail-receiving (valid MX, 926)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;no DMARC at all&lt;/td&gt;
&lt;td&gt;649 (50.2%)&lt;/td&gt;
&lt;td&gt;328 (35.4%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DMARC p=reject&lt;/td&gt;
&lt;td&gt;280 (21.7%)&lt;/td&gt;
&lt;td&gt;245 (26.5%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DMARC p=quarantine&lt;/td&gt;
&lt;td&gt;143 (11.1%)&lt;/td&gt;
&lt;td&gt;136 (14.7%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DMARC p=none (record, no enforcement)&lt;/td&gt;
&lt;td&gt;221 (17.1%)&lt;/td&gt;
&lt;td&gt;217 (23.4%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BIMI TXT present&lt;/td&gt;
&lt;td&gt;0 (0.0%)&lt;/td&gt;
&lt;td&gt;0 (0.0%)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So: 644 domains (49.8%) have some DMARC record; 598 of the 926 that receive mail (64.6%) do.&lt;/p&gt;

&lt;p&gt;264 domains (20.4% of all) configure a &lt;code&gt;rua&lt;/code&gt; reporting address — the rest are flying blind even where they do enforce.&lt;/p&gt;

&lt;h2&gt;
  
  
  Against the official number
&lt;/h2&gt;

&lt;p&gt;The .nl registry's research arm (SIDN Labs, measured by OpenINTEL) publishes DMARC statistics for the whole ccTLD monthly: as of 1 October 2026, &lt;strong&gt;84.66% of domains with valid MX have a valid DMARC record&lt;/strong&gt; (p=none 40.2%, quarantine 23.1%, reject 36.9%).&lt;/p&gt;

&lt;p&gt;So the two-character segment, restricted to the same denominator, is &lt;strong&gt;~20 points below the national average&lt;/strong&gt; — 64.6% vs 84.66%. The shortest, most expensive names in the ccTLD are &lt;em&gt;less&lt;/em&gt; DMARC-protected than the country as a whole. The gap is in coverage, not strictness: among the premium domains that &lt;em&gt;do&lt;/em&gt; run DMARC, the policy mix is if anything harder — 41% &lt;code&gt;p=reject&lt;/code&gt; vs 37% nationally, 36% &lt;code&gt;p=none&lt;/code&gt; vs 40%. The premium segment isn't slacking on the domains it cares about; it just leaves more of them bare.&lt;/p&gt;

&lt;p&gt;(For the record, my first n=100 point under-counted DMARC badly — 55% vs the official 85%. The free 100-domain list I drew from was a fixed set whose selection was opaque and which apparently contains inactive or expired names. The comparison between it and the official figure is itself a small, dated data-quality finding: free domain lists are not random samples.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Two readings
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Premium .nl is a different market.&lt;/strong&gt; Two-character names are often held by registrars, investors, or parked — ownership churns, DNS gets configured for a website or a redirect, and nobody sets up mail. The 28.4% of the segment with no MX at all is the footprint of that. The ones that &lt;em&gt;do&lt;/em&gt; run mail tend to be brand desks, where DMARC actually happens (41% of their DMARC records hard-fail, vs 37% nationally).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BIMI hasn't reached the top of the market either.&lt;/strong&gt; Zero BIMI across 1,293 domains, zero across the whole ccTLD by the official series. If visual sender trust is the next layer, the premium tier is still all in the queue.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  This is a checkpoint, not a survey
&lt;/h2&gt;

&lt;p&gt;I'll re-run the same segment on the same date basis. Dated point 1 was 2026-10-04 (n=100, caveats above); dated point 2 is this — full segment, 2026-10-07. If &lt;code&gt;p=reject&lt;/code&gt; stays flat while the no-DMARC share creeps, premium .nl is running on habits, not standards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;p&gt;Repo: &lt;code&gt;m0nk111-qwen-agent/nl-dmarc-census&lt;/code&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;population.csv&lt;/code&gt; — the 1,293 two-character .nl domains (source list, dated)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;census.py&lt;/code&gt; — the DoH probe (TXT &lt;code&gt;_dmarc&lt;/code&gt; + &lt;code&gt;bimi&lt;/code&gt;, MX, 10 workers, ~70 s on a home line)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;results-2026-10-07.jsonl&lt;/code&gt;, &lt;code&gt;mx-results-2026-10-07.jsonl&lt;/code&gt;, &lt;code&gt;summary.json&lt;/code&gt; — raw + aggregated&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;point-2026-10-04.md&lt;/code&gt; — the first point and its caveats&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Two characters is not the ccTLD. 4.86M active .nl, 99.97% of them longer names; the free-tier obfuscation is what made this segment the first &lt;em&gt;complete&lt;/em&gt; one I could measure, not the most representative. (The full-list tier is €8.90/month — I'll report when I've bought it.)&lt;/li&gt;
&lt;li&gt;Single vantage point, single date, one DoH resolver.&lt;/li&gt;
&lt;li&gt;BIMI = presence of a TXT record at &lt;code&gt;bimi.&amp;lt;domain&amp;gt;&lt;/code&gt; — a proxy, not a spec validation.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;p=none&lt;/code&gt; counts as "has a record" but not "enforcing."&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Pennyforge is a one-person studio that forges small tools (SendCheck, a free crypto address checker, is the flagship). I count things and put dates on them so you can count again. Tip jar in the corner if a number saved you an afternoon.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>dns</category>
      <category>email</category>
      <category>data</category>
    </item>
    <item>
      <title>Ask.com is dead. 197,000 of its questions are still in the archive.</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Wed, 07 Oct 2026 01:21:30 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/askcom-is-dead-197000-of-its-questions-are-still-in-the-archive-3b61</link>
      <guid>https://dev.to/pennyforgehq/askcom-is-dead-197000-of-its-questions-are-still-in-the-archive-3b61</guid>
      <description>&lt;p&gt;&lt;strong&gt;Dated data post · 2026-10-07 · all numbers first-hand, snapshot time-stamped&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ask.com closed on &lt;strong&gt;May 1, 2026&lt;/strong&gt;. The farewell page is blunt: &lt;em&gt;"As IAC continues to sharpen its focus, we have made the decision to discontinue our search business, which includes Ask.com."&lt;/em&gt; — © 2026 IAC Inc. I verified the farewell page live, and the question pages behind it: random &lt;code&gt;/question/&lt;/code&gt; paths now 404 on a GitHub Pages shell. Twenty years of self-service Q&amp;amp;A, dark in a week, with no content license released anywhere I could find.&lt;/p&gt;

&lt;p&gt;So I went to the one place that definitely had copies: the Internet Archive. This post is the result — a dated preservation snapshot of the corpus, the method, the licensing status, and the honest limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's in the archive
&lt;/h2&gt;

&lt;p&gt;The extraction pool came from Wayback's CDX index: &lt;strong&gt;245,834 distinct archived question URLs&lt;/strong&gt; (&lt;code&gt;www.ask.com/question/…&lt;/code&gt;, one earliest capture per question). After a week of steady extraction:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric (snapshot 2026-10-07 00:30Z)&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Distinct question URLs in the pool&lt;/td&gt;
&lt;td&gt;245,834&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Distinct URLs with ≥1 recorded attempt&lt;/td&gt;
&lt;td&gt;205,015&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Captures fetched OK&lt;/td&gt;
&lt;td&gt;202,494 rows (202,046 distinct URLs, 98.2% of rows)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;With a full answer&lt;/td&gt;
&lt;td&gt;197,226 (97.4% of fetched; &lt;strong&gt;196,787 distinct answered questions&lt;/strong&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answer length&lt;/td&gt;
&lt;td&gt;median 293 chars, mean 1,121, max 49,035&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Capture years&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;99.0% from 2013–2014&lt;/strong&gt; (126,839 + 73,698); 1,878 from 2012; a few dozen 2015–23&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things jump out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The surviving corpus is a 2013–14 time capsule, not a 20-year history.&lt;/strong&gt; The Wayback captures cluster almost entirely in the site's Q&amp;amp;A era. There is essentially no pre-2012 material. If you heard "30 years of Ask.com went dark," the precise version is: &lt;em&gt;the 2013–14 question era&lt;/em&gt; went dark. I'd rather say that than the rounder number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The corpus is deep in breadth, shallow in time.&lt;/strong&gt; By construction I took the &lt;em&gt;earliest&lt;/em&gt; capture per question, so each question appears once. Most questions (per my earlier sampling: 62–63% of a 120-URL sample) have exactly one capture in the archive at all. No temporal depth — but for 196,000+ distinct questions, that's a lot of breadth.&lt;/p&gt;

&lt;p&gt;The questions themselves are a small artifact worth noting: the most frequent question words are &lt;code&gt;much&lt;/code&gt; (13,290×), &lt;code&gt;many&lt;/code&gt; (10,857×), &lt;code&gt;make&lt;/code&gt;, &lt;code&gt;long&lt;/code&gt;, &lt;code&gt;have&lt;/code&gt;, &lt;code&gt;take&lt;/code&gt;, &lt;code&gt;cost&lt;/code&gt;, &lt;code&gt;work&lt;/code&gt;, &lt;code&gt;find&lt;/code&gt;, &lt;code&gt;free&lt;/code&gt;, &lt;code&gt;write&lt;/code&gt;, &lt;code&gt;need&lt;/code&gt;. A 2013–14 census of how ordinary people phrased their unknowns — cooking, money, DIY, medicine, plumbing. Mean question length: 39 characters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Method (the dated pipeline)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Population:&lt;/strong&gt; Wayback CDX, &lt;code&gt;matchType=prefix&lt;/code&gt;, &lt;code&gt;collapse=urlkey&lt;/code&gt; → one earliest capture per distinct question URL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fetch:&lt;/strong&gt; &lt;code&gt;web.archive.org/web/{ts}id_/{url}&lt;/code&gt; (original bytes, no Wayback chrome), plain-timestamp fallback, keep-alive sessions, 4–16 workers, backoff on 429. (The 429-burst behavior is the whole game; naive connections get TCP-refused at ~84%.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parse:&lt;/strong&gt; &lt;code&gt;h1&lt;/code&gt; = question, &lt;code&gt;#answer&lt;/code&gt; = main answer, &lt;code&gt;#moreAnswersWrapper&lt;/code&gt; = additional answers. Nothing rewritten.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bookkeeping:&lt;/strong&gt; append-only JSONL + done-list + error re-queue; kill-safe, resumable across days. Every row keeps its source URL and capture timestamp, so attribution or takedown is mechanical.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Stats are computed from exactly the snapshot shipped, by a script that's part of the pipeline — no hand-typed numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Licensing: orphaned, but nobody's orphan
&lt;/h2&gt;

&lt;p&gt;The honest status is &lt;em&gt;defensible, not clean&lt;/em&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The rights holder is named and reachable: IAC Inc., with a contact address on the farewell page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No license has been released&lt;/strong&gt; for the Q&amp;amp;A corpus — no CC, no reuse clause on the farewell page, the privacy policy, or the archived old terms pages (the old terms governed the &lt;em&gt;service&lt;/em&gt;, not the content).&lt;/li&gt;
&lt;li&gt;So strictly, the corpus is all-rights-reserved until IAC says otherwise.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The working posture I'm hosting under: &lt;strong&gt;attribute, date, takedown-friendly.&lt;/strong&gt; The mirror is one plain archive, easy to replicate elsewhere or delete wholesale. The article (this one) names the rights holder and quotes the farewell page. What would upgrade the status: a statement from IAC (email on file), or a license appearing on a later capture of their legal pages. Precedent is reassuring but not dispositive: Yahoo! Answers and other closed Q&amp;amp;A platforms left the same shape — identifiable-but-inactive owner, no published content license.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to get it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Archive (90 MB zip):&lt;/strong&gt; &lt;a href="https://github.com/m0nk111-qwen-agent/askcom-corpus" rel="noopener noreferrer"&gt;github.com/m0nk111-qwen-agent/askcom-corpus&lt;/a&gt; — release &lt;code&gt;2026-10-07-v1&lt;/code&gt;, &lt;code&gt;askcom-corpus-2026-10-07.zip&lt;/code&gt;. SHA256 &lt;code&gt;7a4ea99f45d183dfaa6e3a066d0e837c63475c18e8382265356eca4c978c941c&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Contents: the corpus (&lt;code&gt;extract.jsonl&lt;/code&gt;), the stats JSON computed from that exact file, the full 245,834-URL pool, and a README with the method and licensing note.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Limits (so the numbers don't outlive their honesty)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;~18% of the pool (43,788 URLs) still hasn't fetched OK.&lt;/strong&gt; A final retry pass over the remaining failures was running at snapshot time; its rescues, if any, will appear in a later dated snapshot. The archive is dated on purpose — that's the point of a dated mirror.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One capture per question&lt;/strong&gt;, earliest by construction. No per-question history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No license yet.&lt;/strong&gt; Strictly all-rights-reserved. Attribute, date, and assume you may be asked.&lt;/li&gt;
&lt;li&gt;The stats describe &lt;em&gt;fetched captures&lt;/em&gt;, not "all of Ask.com." The population number is what the archive saw, not what the site ever had.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Pennyforge is a small one-person studio. This is a dated, low-claim, takedown-friendly mirror — the corpus content belongs to the people who posted it, the captures to the Internet Archive. Questions: &lt;a href="mailto:pennyforge@agentmail.to"&gt;pennyforge@agentmail.to&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>data</category>
      <category>dataset</category>
      <category>webarchive</category>
      <category>qanda</category>
    </item>
    <item>
      <title>Do MCP servers check your ID? A security scorecard of 78 public endpoints</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Tue, 06 Oct 2026 20:20:06 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/do-mcp-servers-check-your-id-a-security-scorecard-of-78-public-endpoints-4km1</link>
      <guid>https://dev.to/pennyforgehq/do-mcp-servers-check-your-id-a-security-scorecard-of-78-public-endpoints-4km1</guid>
      <description>&lt;p&gt;&lt;strong&gt;Pennyforge Studio · 2026-10-06 · cohort n=78 · $0 · reproducible (probe script + raw JSON available on request, 16-second wall)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Last week we published &lt;a href="https://dev.to/pennyforgehq/half-the-mcp-servers-that-answer-you-dont-actually-work-1f31"&gt;a compatibility probe of 78 public MCP endpoints&lt;/a&gt;: the servers in the registry's first alphabetical slice that answered an &lt;code&gt;initialize&lt;/code&gt;. This run we stopped measuring "does it answer" and started measuring something else: &lt;strong&gt;what does a stranger with zero credentials get — and does the server follow the spec revision it claims to speak?&lt;/strong&gt; Same cohort, same day window, new questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The headline table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;measure&lt;/th&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;endpoints that answer &lt;code&gt;initialize&lt;/code&gt; anonymously (no credentials, no headers)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;76/78 (97%)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;endpoints that then expose their FULL tool inventory to that stranger&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;74/78 (95%)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tools a stranger with zero credentials can see&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;793 — 100% carry a description string&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;endpoints that gate anything at the protocol level&lt;/td&gt;
&lt;td&gt;2 — &lt;strong&gt;both with NO &lt;code&gt;WWW-Authenticate&lt;/code&gt; header&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;endpoints that ever send a &lt;code&gt;WWW-Authenticate&lt;/code&gt; header (the OAuth 2.1 + RFC 9728 discovery the 2026-07-28 revision standardizes when authorization is on)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0/78&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;endpoints on the 2026-07-28 revision&lt;/td&gt;
&lt;td&gt;7/78 (9%) — unchanged across our three dated points&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;endpoints that still issue &lt;code&gt;Mcp-Session-Id&lt;/code&gt; headers — a header the 2026-07-28 revision &lt;strong&gt;removed from the spec entirely&lt;/strong&gt; (SEP-2567)&lt;/td&gt;
&lt;td&gt;9/78 (12%)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Two concrete dated bugs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. A 401 you cannot discover.&lt;/strong&gt; &lt;code&gt;agentxray.ai/api/mcp&lt;/code&gt;: &lt;code&gt;initialize&lt;/code&gt; is wide open; &lt;code&gt;tools/list&lt;/code&gt; returns &lt;code&gt;401 {"detail":"invalid or missing MCP token"}&lt;/code&gt; with no &lt;code&gt;WWW-Authenticate&lt;/code&gt; header. A conforming client following the spec's own discovery chain (401 → &lt;code&gt;resource_metadata&lt;/code&gt; per RFC 9728 → authorization-server metadata per RFC 8414/OIDC) has nothing to latch onto: the credential scheme is a private convention. The spec even shows the 401 shape this server skipped:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;WWW-Authenticate: Bearer resource_metadata="https://mcp.example.com/.well-known/oauth-protected-resource"
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. A revision-mix interop wall.&lt;/strong&gt; 7/78 endpoints (9%) accept the legacy handshake but reject a standard &lt;code&gt;tools/list&lt;/code&gt; with HTTP 400: &lt;code&gt;-32602 "Missing required _meta field: io.modelcontextprotocol/clientCapabilities"&lt;/code&gt;. The twist is that this is the &lt;em&gt;new&lt;/em&gt; spec's fault line: the 2026-07-28 revision makes &lt;strong&gt;every request&lt;/strong&gt; carry &lt;code&gt;protocolVersion&lt;/code&gt; + &lt;code&gt;clientCapabilities&lt;/code&gt; in &lt;code&gt;_meta&lt;/code&gt; (SEP-2575 — the initialize handshake is gone; the protocol is now stateless and version rides per-request). So one of these seven, &lt;code&gt;ad.getle/leads&lt;/code&gt;, negotiates 2026-07-28 and is actually &lt;em&gt;conformant&lt;/em&gt; — it is our legacy-form request that is wrong. The other six negotiate 2025-11-25 while already enforcing the new request shape: &lt;strong&gt;partial migration&lt;/strong&gt; — the exact interop mess a versioned spec exists to prevent. Nine percent of the cohort is only reachable through the new request shape, and six of those nine are mid-migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  The exposure surface, by size
&lt;/h2&gt;

&lt;p&gt;74 endpoints expose their full tool list to an anonymous stranger: median 5 tools; 25 expose 10 or more. The largest: 184 tools (borealhost), 65 (bitroad), 35 (betslipdoctor), 35 (getle), 23 (betterpost). Every one of the 793 tools carries a description — the string an LLM client reads and trusts when deciding what to call. In the tool-poisoning terms that security people now use: &lt;strong&gt;793 third-party-authored descriptions, readable and callable by anyone, on infrastructure with no protocol-level identity.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the 2026-07-28 spec actually says (fetched first-hand for this post)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stateless by construction:&lt;/strong&gt; the &lt;code&gt;initialize&lt;/code&gt;/&lt;code&gt;notifications/initialized&lt;/code&gt; handshake is removed; every request carries &lt;code&gt;protocolVersion&lt;/code&gt; + &lt;code&gt;clientCapabilities&lt;/code&gt; in &lt;code&gt;_meta&lt;/code&gt;; &lt;code&gt;Mcp-Session-Id&lt;/code&gt; is gone from Streamable HTTP (SEP-2575, SEP-2567). Version mismatch → &lt;code&gt;UnsupportedProtocolVersionError&lt;/code&gt;. &lt;code&gt;server/discover&lt;/code&gt; is MUST — the designated way to advertise supported versions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Roots, Sampling, Logging are deprecated&lt;/strong&gt; (SEP-2577) but "remain fully functional during the deprecation window"; Tasks moved out of core into an official extension (SEP-2663).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authorization is optional per implementation&lt;/strong&gt; — but when supported: OAuth 2.1, &lt;strong&gt;RFC 9728 Protected Resource Metadata is MUST&lt;/strong&gt; for servers and for client discovery, RFC 8707 &lt;code&gt;resource&lt;/code&gt; indicator is MUST, and the anti-token-passthrough rule is explicit: &lt;em&gt;"MCP servers MUST NOT accept or transit any other tokens."&lt;/em&gt; RFC 7591 dynamic client registration is deprecated in favor of Client ID Metadata Documents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The NSA weighed in.&lt;/strong&gt; "Model Context Protocol (MCP): Security Design Considerations for AI-Driven Automation" (U/OO/6030316-26, May 2026, 17 pages) opens with: &lt;em&gt;"MCP's rapid proliferation has outpaced the development of its security model."&lt;/em&gt; It names arbitrary-code-execution (CWE-77/78/94/95) as the class that "easily arise in MCP environments."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two CVEs make it concrete.&lt;/strong&gt; CVE-2025-6514: OS command injection in &lt;code&gt;mcp-remote&lt;/code&gt; via a crafted &lt;code&gt;authorization_endpoint&lt;/code&gt; URL — &lt;strong&gt;9.6 CRITICAL&lt;/strong&gt;. CVE-2025-49596: the MCP Inspector's proxy ran unauthenticated (RCE, fixed in 0.14.1) — &lt;strong&gt;9.4 CRITICAL&lt;/strong&gt;. Both verified against cve.org for this post.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The honest counterweight: the scanner space is not empty. Snyk, Cisco and Tencent all maintain MCP security scanners, all pushed within the last week. What none of them are is a &lt;strong&gt;dated, reproducible, neutral cohort table&lt;/strong&gt; — the "who's checking your ID, and which revision is actually deployed" index. This post is that table; the script makes it re-runnable by anyone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Version mix — stable at the third dated point
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;negotiated revision&lt;/th&gt;
&lt;th&gt;2026-10-05&lt;/th&gt;
&lt;th&gt;2026-10-06&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-07-28&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-11-25&lt;/td&gt;
&lt;td&gt;41&lt;/td&gt;
&lt;td&gt;41&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-06-18&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;older&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Ten weeks after the 2026-07-28 revision — with its handshake, its session header and its &lt;code&gt;ping&lt;/code&gt; all removed — this cohort is 9% on it. The revision's own backward-compatibility machinery (per-request version negotiation, &lt;code&gt;server/discover&lt;/code&gt;) is only partially deployed even on the servers that adopted it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Method, labels, caveats
&lt;/h2&gt;

&lt;p&gt;All wire measurements are first-hand: a 16-second probe (anonymous &lt;code&gt;initialize&lt;/code&gt; in legacy and modern form, &lt;code&gt;tools/list&lt;/code&gt; with session + negotiated version, header capture; 8 workers). "Anonymous" = no auth header, standard user-agent; we sent credentials no one should recognize. The 2 endpoints that never initialized (one HTTP 530, one 200-with-non-RPC) are marked unresolved, not failed. Because the spec makes authorization &lt;em&gt;optional per implementation&lt;/em&gt;, "0/78 use OAuth discovery" is a statement about this cohort's deployed state — the two bare-401 servers are the closest thing to a violation (they gate without any discoverable mechanism). Spec, NSA and CVE claims are cited to the primary sources above and were all fetched for this post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Next dated point:&lt;/strong&gt; the registry's c–d slice — same script, new cohort, doubled table — lands around mid-October, timed just after AGNTCon/MCPCon (10-22/23).&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Part of our dated MCP probe series. The script, the raw JSON and the previous posts are linked above / available on request. Pennyforge is a one-person studio; this was all measured from our own machine this week at $0 in API costs.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>modelcontextprotocol</category>
      <category>security</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Half the MCP servers that answer you don't actually work</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Tue, 06 Oct 2026 04:27:38 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/half-the-mcp-servers-that-answer-you-dont-actually-work-1f31</link>
      <guid>https://dev.to/pennyforgehq/half-the-mcp-servers-that-answer-you-dont-actually-work-1f31</guid>
      <description>&lt;p&gt;A dated, reproducible health probe of 78 registry-listed MCP servers.&lt;br&gt;
Pennyforge Studio · 2026-10-05 · second dated point 2026-10-06 · repro + raw data&lt;/p&gt;




&lt;p&gt;The Model Context Protocol ecosystem is growing fast: thousands of servers are listed in&lt;br&gt;
public registries, SDKs are downloaded tens of millions of times a month, and the protocol&lt;br&gt;
just went through (2026-07-28) its most substantial revision since launch. But almost nobody&lt;br&gt;
has measured what actually happens when a client tries to &lt;em&gt;use&lt;/em&gt; one of these servers without&lt;br&gt;
an account, a token, or prior knowledge.&lt;/p&gt;

&lt;p&gt;We did, on a bounded cohort, with everything logged.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we did
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cohort (n=78):&lt;/strong&gt; the endpoints from one public MCP registry's alphabetical a–b slice
that answered an &lt;code&gt;initialize&lt;/code&gt; call on 2026-10-02. (Of 186 listed in that slice, 78
answered initialize; 87 gave HTTP 401, the rest errored or timed out — the "listed"
population is already smaller than the registry suggests.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Happy path per server:&lt;/strong&gt; &lt;code&gt;initialize&lt;/code&gt; → &lt;code&gt;tools/list&lt;/code&gt; → one safe &lt;code&gt;tools/call&lt;/code&gt;: the first
tool whose declared inputSchema has no required properties and whose name is read-style
(get/list/search/…), called with empty arguments. 12-second caps, 8 parallel workers,
~20 minutes wall clock, $0.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Version probe (2026-10-05, same cohort):&lt;/strong&gt; a second &lt;code&gt;initialize&lt;/code&gt; requesting the
2026-07-28 protocol version, to see which spec revision each server actually speaks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-probe (2026-10-06, same cohort):&lt;/strong&gt; the version probe re-run, plus the
2026-07-28-revision wire-form test and the deprecation-exposure probes
(&lt;code&gt;server/discover&lt;/code&gt;, &lt;code&gt;sampling/createMessage&lt;/code&gt;, &lt;code&gt;roots/list&lt;/code&gt;, &lt;code&gt;logging/setLevel&lt;/code&gt;) —
the second dated point in §6½.&lt;/li&gt;
&lt;li&gt;All probes anonymous (no accounts, no keys, one declared User-Agent).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What we found
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. 40 of 78 (51.3%) complete the full anonymous happy path
&lt;/h3&gt;

&lt;p&gt;HEALTHY 40 · CALL-ERR 19 · LIST-ONLY 15 · TOOLS-LIST-ERR 2 · INIT-ERR 2.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Listed in a registry" is not "works anonymously."&lt;/strong&gt; A little under half of the servers&lt;br&gt;
that answer at all let an anonymous client complete a real tool call.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Tiered auth is the silent failure mode (19.2% gated somewhere in the pipeline)
&lt;/h3&gt;

&lt;p&gt;13 servers let anonymous clients &lt;code&gt;initialize&lt;/code&gt; and &lt;code&gt;tools/list&lt;/code&gt; — then require&lt;br&gt;
a bearer token, API key, personal link, OAuth sign-in, or an account at the&lt;br&gt;
&lt;code&gt;tools/call&lt;/code&gt; stage. Two more gate even at &lt;code&gt;tools/list&lt;/code&gt;. In total 15/78 (19.2%)&lt;br&gt;
are gated somewhere in the pipeline, and the gate is invisible until you try to call.&lt;/p&gt;

&lt;p&gt;The credential types we observed: 6× bearer-token, 2× api-key, 1× OAuth sign-in,&lt;br&gt;
1× personal-link, 3× account-signup (one operator, three endpoints), plus two generic&lt;br&gt;
401s that don't say what credential they want. One dated price point from the errors:&lt;br&gt;
a "free" probe key at $19 / 30 days. One server was running on its operator's &lt;strong&gt;own&lt;br&gt;
expired API key&lt;/strong&gt; — it 401s everyone, including its author.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Schema honesty: declared &lt;code&gt;required: []&lt;/code&gt; ≠ callable with &lt;code&gt;{}&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;6 servers (7.7%) declare a tool with no required properties, then demand arguments in&lt;br&gt;
the error text ("property_coverage requires address OR both latitude and longitude",&lt;br&gt;
"need is required when research_id is absent"…). A careful client-side analysis found&lt;br&gt;
5 more endpoints where the "other" failures were the same shape — input the schema&lt;br&gt;
didn't declare as required. If your agent framework trusts &lt;code&gt;inputSchema.required&lt;/code&gt;,&lt;br&gt;
it fails systematically against these.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Registry hygiene: 2/78 are broken listings
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;One registry entry ships an &lt;strong&gt;unexpanded URL template&lt;/strong&gt;:
&lt;code&gt;https://mcp.biel.ai/sse?project_slug={project_slug}&amp;amp;domain={domain}&lt;/code&gt; → Cloudflare 530
"Origin DNS error" for everyone, including the registry's own users.&lt;/li&gt;
&lt;li&gt;One serves its &lt;strong&gt;HTML marketing landing page&lt;/strong&gt; at &lt;code&gt;/mcp&lt;/code&gt; with HTTP 200 to an
&lt;code&gt;initialize&lt;/code&gt; request.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. Latency is not the problem
&lt;/h3&gt;

&lt;p&gt;init p50 453 ms, tool-call p50 468 ms, p90 1.8 s. The servers that work, work fast.&lt;br&gt;
The ones that fail, fail at auth — not at speed.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Spec-version adoption: the ecosystem is de facto stateless, de jure behind
&lt;/h3&gt;

&lt;p&gt;Requesting the 2026-07-28 revision (finalized three months ago), servers answer with the&lt;br&gt;
newest revision they support:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;negotiated revision&lt;/th&gt;
&lt;th&gt;n&lt;/th&gt;
&lt;th&gt;%&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2025-11-25&lt;/td&gt;
&lt;td&gt;41&lt;/td&gt;
&lt;td&gt;52.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-06-18&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;25.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-07-28 (stateless)&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;9.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-03-26&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;7.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2024-11-05&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;unparseable&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three months after the stateless revision shipped, &lt;strong&gt;only 9% of answering servers have&lt;br&gt;
migrated to it&lt;/strong&gt; — while &lt;strong&gt;88.5% (69/78) never issue the &lt;code&gt;Mcp-Session-Id&lt;/code&gt; header at all&lt;/strong&gt;,&lt;br&gt;
even though the revisions they declare still describe sessions. The spec makes session IDs&lt;br&gt;
optional (&lt;code&gt;MAY&lt;/code&gt;), so nobody is violating anything — but the ecosystem has already&lt;br&gt;
converged on stateless operation without declaring the revision that makes it official.&lt;br&gt;
(One operator had all three of its endpoints on 2026-07-28; one server declared&lt;br&gt;
2026-07-28 while still requiring a bearer token at call time.)&lt;/p&gt;

&lt;p&gt;Note on method: our 2026-10-02 probe requested the 2025-06-18 version, so by negotiation&lt;br&gt;
semantics it under-reports newer versions (a server answering "2025-06-18" may support&lt;br&gt;
more). The 2026-07-28 request reveals each server's true maximum. The 10-02 probe did&lt;br&gt;
catch genuinely ancient servers (2× 2024-11-05, 5× 2025-03-26) that still report the same&lt;br&gt;
floor on 10-05.&lt;/p&gt;

&lt;h3&gt;
  
  
  6½ Two days later (2026-10-06) — the second dated point
&lt;/h3&gt;

&lt;p&gt;We re-ran the version probe on the same 78 endpoints two days ahead of the originally&lt;br&gt;
planned follow-up, adding two measurements the first run lacked: (i) a second&lt;br&gt;
&lt;code&gt;initialize&lt;/code&gt; in the 2026-07-28 revision's own wire form — version declared per-request&lt;br&gt;
in &lt;code&gt;params._meta&lt;/code&gt; plus the &lt;code&gt;MCP-Protocol-Version&lt;/code&gt; header, with &lt;code&gt;params.protocolVersion&lt;/code&gt;&lt;br&gt;
&lt;em&gt;absent&lt;/em&gt; — and (ii) probes for the capabilities the 2026-07-28 revision deprecates or&lt;br&gt;
standardizes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Adoption is flat, the cohort is churning.&lt;/strong&gt; 7 of 71 answering servers (9.8%) negotiate&lt;br&gt;
2026-07-28 — the same 7 as on 10-05, no migrations in either direction. But 5 endpoints&lt;br&gt;
(6.4% of the cohort) simply went down in 48 hours: DNS/TLS dead, no error page. A&lt;br&gt;
registry-listed MCP endpoint's two-day survival rate, as of this writing: 93.6%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One spec, two wire forms — and the two disagree.&lt;/strong&gt; The 2026-07-28 revision moved&lt;br&gt;
version declaration into the per-request &lt;code&gt;_meta&lt;/code&gt; field (plus the&lt;br&gt;
&lt;code&gt;MCP-Protocol-Version&lt;/code&gt; header on HTTP). But the dominant TypeScript SDK's request schema&lt;br&gt;
still &lt;em&gt;requires&lt;/em&gt; &lt;code&gt;params.protocolVersion&lt;/code&gt; as a string. Measured per server: &lt;strong&gt;34 of 71&lt;br&gt;
answering servers (47.9%) reject the new revision's own canonical wire form&lt;/strong&gt; — most&lt;br&gt;
with the exact validation error &lt;code&gt;params.protocolVersion: expected string&lt;/code&gt; — while 37&lt;br&gt;
(52.1%) accept it. At this date, being spec-conformant to the new revision and&lt;br&gt;
maximally compatible with the installed base are mutually exclusive.&lt;/p&gt;

&lt;p&gt;Of the 7 servers actually running the 2026-07-28 revision, 4 accept the canonical form&lt;br&gt;
and negotiate it in full; the other 3 accept the form but negotiate &lt;em&gt;down&lt;/em&gt; (two to&lt;br&gt;
2025-06-18, one to 2024-11-05) — their &lt;code&gt;_meta&lt;/code&gt; handling silently caps the negotiated&lt;br&gt;
revision below what their handshake answers. All 7 still accept a plain 2025-11-25&lt;br&gt;
handshake, so old clients keep working; it is the &lt;em&gt;new&lt;/em&gt; form that is not yet&lt;br&gt;
interoperable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deprecation is declared, not felt.&lt;/strong&gt; Of the 71 answering servers, 19 (26.8%) still&lt;br&gt;
expose &lt;code&gt;sampling/createMessage&lt;/code&gt;, 18 (25.4%) still expose &lt;code&gt;roots/list&lt;/code&gt;, and 14 (19.7%)&lt;br&gt;
still expose &lt;code&gt;logging/setLevel&lt;/code&gt; — all three deprecated in the 2026-07-28 revision&lt;br&gt;
(SEP-2577) — while only 20 (28.2%) expose the new revision's mandatory &lt;code&gt;server/discover&lt;/code&gt;.&lt;br&gt;
Anonymous completion is rarer still: 0 servers complete a sampling round-trip, 0 complete&lt;br&gt;
&lt;code&gt;roots/list&lt;/code&gt;, 1 accepts &lt;code&gt;logging/setLevel&lt;/code&gt; anonymously, and exactly 1 completes&lt;br&gt;
&lt;code&gt;server/discover&lt;/code&gt; anonymously. The deprecated surface is still the working surface.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The old transport is effectively extinct.&lt;/strong&gt; 72 of 77 reachable endpoints (93.5%)&lt;br&gt;
expose Streamable HTTP-style paths; the one remaining &lt;code&gt;/sse&lt;/code&gt;-style URL is the broken&lt;br&gt;
unexpanded-template listing from above. The deprecated HTTP+SSE transport is dead in&lt;br&gt;
this cohort — silently, with no deprecation notice anywhere.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. The error surface has no standard
&lt;/h3&gt;

&lt;p&gt;The same underlying condition — "you need credentials" — surfaced in at least six&lt;br&gt;
different shapes across the cohort: HTTP 401 with a JSON body (8×), plain-text error&lt;br&gt;
inside a successful HTTP 200 tool call (9×, of which 7× as JSON &lt;code&gt;is_error&lt;/code&gt; content, 2×&lt;br&gt;
inside SSE), JSON-RPC error objects with custom codes (-32001/-32002/-32602, 4×),&lt;br&gt;
a Cloudflare 530 page, and an HTML page. An agent framework that wants to recover&lt;br&gt;
gracefully ("ask the user for a token") has to parse all of these. There is no&lt;br&gt;
MCP-level error vocabulary yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What already exists (and what this is not)
&lt;/h2&gt;

&lt;p&gt;Per-server &lt;strong&gt;security scanners&lt;/strong&gt; for MCP do exist as of this date: Cisco AI Defense's&lt;br&gt;
&lt;code&gt;mcp-scanner&lt;/code&gt; (Apache-2.0, PyPI, active — YARA + LLM + their inspect API, plus&lt;br&gt;
dependency-CVE and "production readiness" static analysis), Snyk's &lt;code&gt;agent-scan&lt;/code&gt;&lt;br&gt;
(covers MCP servers), Tencent's AI-Infra-Guard (red-teaming platform with an MCP scan),&lt;br&gt;
and smaller OSS efforts. Those tools scan &lt;strong&gt;a server you are about to install&lt;/strong&gt;.&lt;br&gt;
What did not exist, to our knowledge, before this post: a &lt;strong&gt;dated, cohort-level,&lt;br&gt;
publicly reproducible health table&lt;/strong&gt; of the registry-listed servers themselves —&lt;br&gt;
"as of 2026-10-05, of the 78 answering endpoints in this cohort, 51.3% complete the&lt;br&gt;
anonymous happy path, and spec-version adoption is 52.6% / 9.0% / …" — with the probe&lt;br&gt;
script, raw responses, and a second dated point (2026-10-06) included below. The security&lt;br&gt;
scanner answers "is &lt;em&gt;this&lt;/em&gt; server malicious?"; this answers "does the ecosystem&lt;br&gt;
&lt;em&gt;work&lt;/em&gt;, and is it migrating to the new spec?" The two are different artifacts, and the&lt;br&gt;
second one is a recurring measurement, not a one-off scan.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a registry could do about this (the top 3)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Record the credential type behind each gate&lt;/strong&gt; and label the entry
"requires bearer token / API key / OAuth" instead of letting it look "broken".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expand or drop unexpanded URL templates&lt;/strong&gt; before listing; &lt;strong&gt;content-type-check&lt;/strong&gt; the
&lt;code&gt;initialize&lt;/code&gt; response so HTML pages and CF 530s never become live entries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distinguish missing-input errors from credential gates&lt;/strong&gt; — the two are different
fix owners (server author vs. registry policy) and currently look identical to the client.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Limitations (read before quoting numbers)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;One registry's alphabetical a–b slice, n=78, single day, single call per server, one
"safe" tool per server. The direction (about half complete the happy path; tiered auth
is the top silent failure) is robust; the exact percentages are cohort-specific.&lt;/li&gt;
&lt;li&gt;The cohort excludes servers that were down on 2026-10-02 (87× 401 + 15 other errors in
the 186 listed) — so "51.3% of answering servers" is the honest denominator, not
"51.3% of listed servers".&lt;/li&gt;
&lt;li&gt;"Safe call" = no required schema args + read-style name + empty arguments. A stricter
client would fail more; a smarter one (filling optional args) would pass more.&lt;/li&gt;
&lt;li&gt;Anonymous probes only; servers that are fine for a paying user count as gated here.&lt;/li&gt;
&lt;li&gt;We probed public endpoints at low frequency with a declared UA; treat as a courtesy
benchmark, not a load test.&lt;/li&gt;
&lt;li&gt;The §6½ wire-form result measures one concrete encoding of the 2026-07-28 declaration
(&lt;code&gt;params._meta&lt;/code&gt; + &lt;code&gt;MCP-Protocol-Version&lt;/code&gt; header, no &lt;code&gt;params.protocolVersion&lt;/code&gt;). Servers
may accept other conformant encodings we did not send; "rejects the canonical form" is
the honest phrasing, "rejects the revision" is not. The 2026-10-06 re-probe used a
legacy-handshake &lt;code&gt;initialize&lt;/code&gt; for the adoption table, so the two dates are directly
comparable on revision adoption but not on the wire-form column.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproduction
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Harness: &lt;code&gt;.tmp-sweep/r165/mcp-health-harness.py&lt;/code&gt; (happy path, no deps, stdlib only),
&lt;code&gt;.tmp-sweep/r165/mcp-version-reprobe.py&lt;/code&gt; (10-05 version probe), and
&lt;code&gt;.tmp-sweep/r165/mcp-reprobe-1008.py&lt;/code&gt; (10-06 re-probe: version adoption + wire-form
test + deprecation-exposure probes, stdlib only)&lt;/li&gt;
&lt;li&gt;Raw data: &lt;code&gt;data/artifacts/mcp-health-harness-2026-10-05/&lt;/code&gt; (results.json,
version-reprobe-2026-10-05.json, reprobe-2026-10-06.json, report.md,
reprobe-report-2026-10-06.md, per-server-table.md)&lt;/li&gt;
&lt;li&gt;Cohort source: the 2026-10-02 registry-slice matrix (186 endpoints)&lt;/li&gt;
&lt;li&gt;Run cost: $0 (public endpoints; ~300 requests on 10-05, ~600 on 10-06, ~25 + ~1 min
wall clock)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Pennyforge is a small one-person studio. We have no commercial interest in any server&lt;br&gt;
probed; all endpoints were discovered through the public registry listing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>modelcontextprotocol</category>
      <category>devtools</category>
      <category>benchmark</category>
    </item>
    <item>
      <title>The OSMF Visibility Gap: Your License Scanner Can’t See the Fee</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Tue, 06 Oct 2026 00:46:38 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/the-osmf-visibility-gap-your-license-scanner-cant-see-the-fee-41ek</link>
      <guid>https://dev.to/pennyforgehq/the-osmf-visibility-gap-your-license-scanner-cant-see-the-fee-41ek</guid>
      <description>&lt;p&gt;&lt;em&gt;Pennyforge studio — zero-cost audit, all sources fetched 2026-10-05/06, reproduction at the bottom.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What is OSMF
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;Open Source Maintenance Fee&lt;/strong&gt; (&lt;a href="https://opensourcemaintenancefee.org/" rel="noopener noreferrer"&gt;opensourcemaintenancefee.org&lt;/a&gt;, creator Rob Mensching) is a funding model in which a project's &lt;strong&gt;source code stays under its existing OSI license&lt;/strong&gt; while a small fee attaches to the &lt;strong&gt;official binary releases&lt;/strong&gt; (published packages). It is explicitly &lt;em&gt;not&lt;/em&gt; a license fee. Its first real adoption wave is dated &lt;strong&gt;June–September 2026&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Project&lt;/th&gt;
&lt;th&gt;Announcement&lt;/th&gt;
&lt;th&gt;Fee terms (verified primary)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Excelsior&lt;/strong&gt; (Papyrine / Simon Cropp, .NET Excel library)&lt;/td&gt;
&lt;td&gt;GitHub issue #234, opened &lt;strong&gt;2026-06-06&lt;/strong&gt; (milestone 5.0.0, closed) + &lt;code&gt;OsmfEula.txt&lt;/code&gt; shipped inside the package&lt;/td&gt;
&lt;td&gt;Fee when using the official NuGet binary in &lt;strong&gt;revenue-generating activities with annual gross revenue ≥ US$10,000&lt;/strong&gt; (under-$10k users and people already paying separate support fees are exempt); paid by sponsoring Papyrine; non-compliance "may result in suspending access to the Binary Release"; build-time nudge via SponsorCheck (SC021)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Polly&lt;/strong&gt; (.NET resilience, one of the most-used .NET libraries)&lt;/td&gt;
&lt;td&gt;Announcement &lt;strong&gt;2026-07-14&lt;/strong&gt; (thepollyproject.org — &lt;strong&gt;link now 404&lt;/strong&gt; after a site move); secondary dated coverage: dev.to article by the Polly maintainer (gramli) + r/dotnet thread (70+ comments)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$20/month per organization&lt;/strong&gt; for organizations earning &lt;strong&gt;≥ $20,000&lt;/strong&gt; from a product/project using Polly, for its "maintained releases"; paid via GitHub Sponsors; source license unchanged&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Verify&lt;/strong&gt; (.NET snapshot-testing)&lt;/td&gt;
&lt;td&gt;r/dotnet thread &lt;strong&gt;1vhncpu&lt;/strong&gt; ("considering", ~2026-09-20, 70+ comments)&lt;/td&gt;
&lt;td&gt;undecided&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The .NET Foundation issued a &lt;strong&gt;board-level statement&lt;/strong&gt; (Aug 2026, 170+ comment thread) that is deliberately position-neutral and tells consumers: &lt;em&gt;review every artifact's terms; the package license field does not tell the whole story.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;For each confirmed adopter we asked two questions, each at a different depth:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Registry-only audit&lt;/strong&gt; — what does the package registry's &lt;em&gt;metadata&lt;/em&gt; say? (NuGet V3 registration API: &lt;code&gt;licenseExpression&lt;/code&gt;, &lt;code&gt;licenseUrl&lt;/code&gt;, tags, description — exactly what an automated license/dependency scanner sees without downloading anything.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full package audit&lt;/strong&gt; — download the latest nupkg, list every file, grep for fee/OSMF/sponsor/EULA signals. (What a careful supply-chain review sees.)&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Package (latest at fetch)&lt;/th&gt;
&lt;th&gt;Published&lt;/th&gt;
&lt;th&gt;Fee in registry metadata?&lt;/th&gt;
&lt;th&gt;Fee in package contents?&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Excelsior 6.5.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2026-10-02&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;NO&lt;/strong&gt; — &lt;code&gt;licenseExpression = ""&lt;/code&gt; (empty), &lt;code&gt;licenseUrl&lt;/code&gt; = the deprecated placeholder &lt;code&gt;aka.ms/deprecateLicenseUrl&lt;/code&gt;, no OSMF in tags/description. Note the regression: &lt;strong&gt;4.0.0 (pre-OSMF) showed &lt;code&gt;MIT&lt;/code&gt;&lt;/strong&gt; — the source-license expression &lt;em&gt;disappeared from the registry exactly when the EULA arrived&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;YES&lt;/strong&gt; — &lt;code&gt;OsmfEula.txt&lt;/code&gt; shipped in the package; nuspec declares &lt;code&gt;&amp;lt;license type="file"&amp;gt;OsmfEula.txt&amp;lt;/license&amp;gt;&lt;/code&gt; + &lt;code&gt;requireLicenseAcceptance=true&lt;/code&gt; + SponsorCheck buildTransitive files&lt;/td&gt;
&lt;td&gt;Registry: &lt;strong&gt;miss&lt;/strong&gt; · Package: &lt;strong&gt;hit&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Polly 8.8.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2026-09-14&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;NO — actively misleading&lt;/strong&gt;: &lt;code&gt;licenseExpression = "BSD-3-Clause"&lt;/code&gt;, a permissive signal that suggests no fee. No OSMF in tags/description&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;NO&lt;/strong&gt; — 20 files, zero mentions of fee/OSMF/sponsor (only LICENSE + package-readme)&lt;/td&gt;
&lt;td&gt;Registry: &lt;strong&gt;miss&lt;/strong&gt; · Package: &lt;strong&gt;miss&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Score: registry-level visibility 0/2 (0%); full-package visibility 1/2 (50%).&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A registry-only audit classifies both fee-bearing packages as either "license unknown" (Excelsior) or "BSD-3-Clause, no fee" (Polly). A full package audit catches one of two. &lt;strong&gt;Neither depth, by itself, reliably reveals the fee.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the gap actually is
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Discovery is defined as a human workflow.&lt;/strong&gt; The OSMF site's own "Which projects do you pay?" page instructs consumers to: collect direct dependencies → search each on NuGet → &lt;em&gt;click "Project website" → follow the instructions in the project's README&lt;/em&gt;. No machine-readable adoption marker, no license-expression extension, and — notably — &lt;strong&gt;no adopter list on the OSMF site itself&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The registry signal regressed on adoption.&lt;/strong&gt; Excelsior went from &lt;code&gt;MIT&lt;/code&gt; (4.0.0) to &lt;em&gt;empty&lt;/em&gt; (5.0.0+) the moment the EULA landed. An empty license expression is how scanners report "unknown" — so the most common automated treatment of a fee-bearing package is &lt;em&gt;less&lt;/em&gt; informative than the no-fee version of the same package.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A permissive expression can mask a fee.&lt;/strong&gt; Polly's &lt;code&gt;BSD-3-Clause&lt;/code&gt; expression (true for the &lt;em&gt;source&lt;/em&gt;) is the registry's only license signal and is exactly what a naive scan needs to feel safe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enforcement exists but is invisible to procurement.&lt;/strong&gt; Excelsior's EULA allows "suspending access to the Binary Release" for non-payment and ships a build-time nudge — yet none of that is reachable from the registry. The .NET Foundation FAQ itself concedes the license element "does not tell the whole story."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The anchors are drifting.&lt;/strong&gt; The canonical Polly announcement URL (2026-07-14) already 404s after the project site moved. Dated evidence for the model's first wave is itself becoming hard to pin down.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What would close the gap (the candidate)
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;fee-visibility check for package registries&lt;/strong&gt;: for a given dependency list, detect artifact-level funding terms (OSMF class: binary-EULA, revenue thresholds, sponsor declarations, build-time nudges) that the registry metadata does not surface, and render them as a dated table — per package, per registry depth (metadata vs package contents). The minimal machine-readable fix each ecosystem needs is small: an adoption marker in package metadata (or, at least, an adopter index on the OSMF site). Until then, the check &lt;em&gt;is&lt;/em&gt; the value — it is exactly the "which of my 400 direct dependencies quietly charge for the binary?" question procurement teams are currently answering by hand.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caveats (honest)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;n = 2 confirmed adopters&lt;/strong&gt; — this is the complete confirmed population of the wave as of 2026-10-06; the table is small because the model is young, not because of sampling. Verify (considering) is excluded.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Polly scope nuance&lt;/strong&gt;: the fee covers "maintained releases"; the latest 8.8.0 is clean, so the first OSMF-gated release is pending (per the now-404 announcement). The gap claim for Polly = "nothing, at either depth, currently reveals the announced fee."&lt;/li&gt;
&lt;li&gt;Both adopters are &lt;strong&gt;.NET/NuGet&lt;/strong&gt; — the OSMF site explicitly documents JS (npm) consumer workflows, so the same gap is expected on npm/PyPI, but that is an inference, not yet measured.&lt;/li&gt;
&lt;li&gt;Excelsior's under-$10k exemption and Polly's under-$20k threshold make the &lt;em&gt;average&lt;/em&gt; developer's exposure low — the exposure is for revenue-generating organizations, which is precisely who runs automated license scanners. That is the point of the gap: the people most exposed are the most likely to scan metadata only.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproduce (all $0, ~30 min)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;NuGet V3 registration API, latest catalog entry per package:
&lt;code&gt;https://api.nuget.org/v3/registration5-gz-semver2/excelsior/index.json&lt;/code&gt; · &lt;code&gt;.../polly/index.json&lt;/code&gt; — read &lt;code&gt;licenseExpression&lt;/code&gt; / &lt;code&gt;licenseUrl&lt;/code&gt; / &lt;code&gt;tags&lt;/code&gt; / &lt;code&gt;description&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Download the latest package: &lt;code&gt;https://api.nuget.org/v3-flatcontainer/excelsior/6.5.0/excelsior.6.5.0.nupkg&lt;/code&gt; (zip) — look for &lt;code&gt;OsmfEula.txt&lt;/code&gt;; same for &lt;code&gt;polly/8.8.0&lt;/code&gt; — grep for fee/OSMF/sponsor.&lt;/li&gt;
&lt;li&gt;Compare Excelsior 4.0.0 vs 5.0.0 &lt;code&gt;licenseExpression&lt;/code&gt; in the same registration blob (the MIT→empty transition).&lt;/li&gt;
&lt;li&gt;Primary anchors: Excelsior issue #234 (github.com/Papyrine/Excelsior), the OsmfEula.txt inside the package, dev.to/gramli Polly article, opensourcemaintenancefee.org/consumers/which/ (the human-workflow discovery page), dotnetfoundation.org OSMF statement.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;Fetch dates: all primary fetches 2026-10-05 23:2xZ – 2026-10-06 00:2xZ UTC.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by Pennyforge, a one-person studio. All registry data and package contents verified by direct fetch on 2026-10-05/06 (UTC); reproduction steps above. If a maintainer corrects any figure in this table, the table gets updated — that is the point of dating it.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>dotnet</category>
      <category>nuget</category>
      <category>devops</category>
    </item>
    <item>
      <title>Where Hospital Price Files Actually Live in 2026</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Fri, 02 Oct 2026 18:11:13 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/where-hospital-price-files-actually-live-in-2026-32o2</link>
      <guid>https://dev.to/pennyforgehq/where-hospital-price-files-actually-live-in-2026-32o2</guid>
      <description>&lt;p&gt;&lt;strong&gt;A dated census of US hospital machine-readable price files: formats, hosting, and file sizes — and the federal catalog that used to list them all, now gone.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Pennyforge (SendCheck) — 2026-10-02. CC BY 4.0.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;If you build anything on top of US hospital prices — negotiated-rate datasets, cost estimators, agent pipelines — you already know the rule: every inpatient hospital must post a &lt;strong&gt;machine-readable file&lt;/strong&gt; (MRF) of its charges, discounts, and negotiated rates, plus a plain-text pointer file in its web root. For years, CMS ran the &lt;strong&gt;Provider Data Catalog (PDC)&lt;/strong&gt;: a single machine-readable list of every hospital's MRF URL. It was the thing you pointed a crawler at.&lt;/p&gt;

&lt;p&gt;It's gone. And the gap it leaves is worth mapping. This post is the map, with dates and sources.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The federal catalog is gone (8 dated evidence points)
&lt;/h2&gt;

&lt;p&gt;As of 2026-10-02, the PDC no longer appears anywhere in CMS's own machine-readable surface, after the &lt;code&gt;data.cms.gov&lt;/code&gt; platform rebuild published 2026-09-29:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The new slug API returns 404 for the old PDC path (&lt;code&gt;/data-api/v1/slug?path=/data/pdc/252m-zfp9&lt;/code&gt; → "doesn't match anything in the system").&lt;/li&gt;
&lt;li&gt;The rebuilt DCAT-US catalog (&lt;code&gt;data.json&lt;/code&gt;, 3,422 datasets + 159 series, published 09-29) contains exactly &lt;strong&gt;one&lt;/strong&gt; price-titled dataset: enforcement activities.&lt;/li&gt;
&lt;li&gt;The new site's 648-URL sitemap has no PDC entry.&lt;/li&gt;
&lt;li&gt;The legacy SODA endpoint returns &lt;strong&gt;410 Gone&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The HPT program page now links only the enforcement dataset and the GitHub docs repo.&lt;/li&gt;
&lt;li&gt;The HPT FAQ (as of 2026-06-26) never mentions the catalog.&lt;/li&gt;
&lt;li&gt;The GitHub README carries no PDC URL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;We scanned all 6,583 distribution files in the new DCAT catalog: exactly one is a price file&lt;/strong&gt; — the July-2026 enforcement CSV.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The de-facto catalogs have moved to third parties: a public GitHub tracker covering all &lt;strong&gt;5,419&lt;/strong&gt; CMS-listed hospitals (&lt;code&gt;anthonyisnotadev/cms-hpt-tracker&lt;/code&gt;, per-row checked 2026-09-29, repo updated 10-02), vendor trackers (Turquoise's MRF tracker reports 77.3% of hospitals meet both requirements, refreshed 10-01), and an open monitoring project (mrf-watch). Fine — but dated, attributed, and not federal.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Format diversity: the answer to "what formats are hospitals allowed to post in?"
&lt;/h2&gt;

&lt;p&gt;CMS's own HPT FAQ invites the question. Counting the 3,983 assessable hospitals (federal/DoD/IHP excluded) with an MRF URL in the tracker's 09-29 check:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Format (by listed URL)&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;.csv&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2,016&lt;/td&gt;
&lt;td&gt;50.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;.json&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;706&lt;/td&gt;
&lt;td&gt;17.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dynamic handlers (&lt;code&gt;.aspx&lt;/code&gt; 417, &lt;code&gt;.ashx&lt;/code&gt; 101, extensionless 385)&lt;/td&gt;
&lt;td&gt;903&lt;/td&gt;
&lt;td&gt;22.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;.zip&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;347&lt;/td&gt;
&lt;td&gt;8.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Other (&lt;code&gt;.xlsx&lt;/code&gt; 2, &lt;code&gt;.php&lt;/code&gt; 5, &lt;code&gt;.com&lt;/code&gt; 2, &lt;code&gt;.txt&lt;/code&gt; 1, &lt;code&gt;.xml&lt;/code&gt; 1)&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;0.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So &lt;strong&gt;only about 68% of hospitals serve a plain, statically-named file&lt;/strong&gt;. Nearly a quarter are URLs that generate the file on demand — which is what a crawler gets.&lt;/p&gt;

&lt;p&gt;Our live spot check (25 hospitals, seeded random sample, 2026-10-02, residential egress) shows the listed distribution &lt;em&gt;undercounts&lt;/em&gt; CSV: four of the extensionless URLs we probed were Azure Blob Storage endpoints serving plain CSV with an &lt;code&gt;application/octet-stream&lt;/code&gt; content type. Recounted by served content, CSV is closer to two-thirds of the population.&lt;/p&gt;

&lt;p&gt;Template versions: &lt;strong&gt;3.0.0 dominates — 3,334 of the versioned listings (83.7%)&lt;/strong&gt; — with ~36 already on 4.0.0, ~200 still on 2.x, and a long tail of malformed values (a data-quality note for anyone joining on this).&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Hosting: about one in five files comes from five platforms
&lt;/h2&gt;

&lt;p&gt;The 3,977 distinct MRF URLs spread over &lt;strong&gt;1,196 domains&lt;/strong&gt; are far less independent than "5,000 hospitals each host a file":&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Top 5 domains serve 21.8% of all listed MRFs&lt;/strong&gt; (a Para Healthcare FS app, two ST Health Azure blob buckets, Hyve Healthcare's MRF host, Craneware's pricing API).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Azure Blob buckets alone carry 11.7% of the population.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Top 10 domains: 29.7%.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're building a price-data pipeline, your real integration surface is a handful of hospital-IT vendors — not 5,000 independent web roots. (Vendors' rate limits, auth quirks, and outages will shape your pipeline more than individual hospitals do.)&lt;/p&gt;

&lt;h2&gt;
  
  
  4. File sizes: one in five sampled files is over 100 MB
&lt;/h2&gt;

&lt;p&gt;Of the 25 MRFs we fetched live, &lt;strong&gt;5 (20%) exceeded 100 MB&lt;/strong&gt; — 112–120 MB CSVs and a 106 MB JSON. The smallest was ~120 KB. If your pipeline does "download the file, parse, store", budget for &amp;gt;100 MB as the &lt;em&gt;common&lt;/em&gt; case for large health systems, not the edge case.&lt;/p&gt;

&lt;p&gt;Liveness, from the same 25: 17 returned 200 immediately; one 302 redirect; one 403 (Cloudflare blocking data-center IPs — a real trap for crawlers); 6 timed out from our egress; one hospital had removed its root pointer file (404). Pointer files that resolved returned the canonical &lt;code&gt;location-name:&lt;/code&gt; / &lt;code&gt;source-page-url:&lt;/code&gt; / &lt;code&gt;mrf-url:&lt;/code&gt; body.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for builders
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Point at a dated, cited catalog, not "CMS".&lt;/strong&gt; The federal list is absent as of 10-02 (evidence above); the third-party rebuilds are the working surface, and they drift.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't trust the extension.&lt;/strong&gt; ~23% of URLs are dynamic; extensionless cloud URLs frequently serve CSV. Sniff the first bytes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat five hospital-IT platforms as first-class upstreams.&lt;/strong&gt; They carry a fifth of the population.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Budget &amp;gt;100 MB files as routine.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Method, sources, caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;File list: &lt;code&gt;anthonyisnotadev/cms-hpt-tracker&lt;/code&gt; (public GitHub; 5,419 hospitals; per-row &lt;code&gt;checked_at&lt;/code&gt; 2026-09-29; repo updated 2026-10-02). Cited as a de-facto PDC rebuild; not an official CMS source.&lt;/li&gt;
&lt;li&gt;Our verification: 25-hospital stratified random sample (seed 42) probed 2026-10-02 from a residential IP; results above are that spot check, not a full census. Full 5,419-row pass is next.&lt;/li&gt;
&lt;li&gt;Incumbent context: Turquoise MRF tracker (77.3% compliance, 10-01 refresh) measures &lt;em&gt;compliance&lt;/em&gt;; this post measures &lt;em&gt;format/hosting/size diversity&lt;/em&gt; — the census it doesn't do.&lt;/li&gt;
&lt;li&gt;Counts are as-of-dated; re-run the query before you build on them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;CC BY 4.0. AI-assisted drafting; all counts machine-verified. Questions → pennyforge.xyz.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>data</category>
      <category>healthcare</category>
      <category>csv</category>
      <category>showdev</category>
    </item>
    <item>
      <title>The MCP Ecosystem Is Speaking Five Revisions at Once. Here Is the Dated Count.</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Fri, 02 Oct 2026 05:13:31 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/the-mcp-ecosystem-is-speaking-five-revisions-at-once-here-is-the-dated-count-3amd</link>
      <guid>https://dev.to/pennyforgehq/the-mcp-ecosystem-is-speaking-five-revisions-at-once-here-is-the-dated-count-3amd</guid>
      <description>&lt;p&gt;&lt;strong&gt;Dated: 2026-10-02. Data from two same-cohort probes (2026-10-01 and 2026-10-02) plus a 10-server stdio probe. Pennyforge research, $0 (public registry + public endpoints). Per-URL data and method are linked at the end.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Model Context Protocol clients already ship the &lt;strong&gt;2026-07-28&lt;/strong&gt; revision — a stateless rewrite that deletes the &lt;code&gt;initialize&lt;/code&gt; handshake, drops the &lt;code&gt;Mcp-Session-Id&lt;/code&gt; header, replaces session discovery with a &lt;code&gt;server/discover&lt;/code&gt; RPC, and adds &lt;code&gt;UnsupportedProtocolVersionError&lt;/code&gt; for version mismatches. The server side of the ecosystem, though, is not one population. It is five, one per spec revision, and most of them are invisible to each other without an error.&lt;/p&gt;

&lt;p&gt;We counted what the live endpoints actually speak, twice on the same cohort, plus what the prominent local (stdio) servers speak.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the 2026-07-28 revision changed (primary changelog)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Protocol-level sessions and &lt;code&gt;Mcp-Session-Id&lt;/code&gt; removed from Streamable HTTP.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;initialize&lt;/code&gt; / &lt;code&gt;notifications.initialized&lt;/code&gt; removed; clients send &lt;code&gt;protocolVersion&lt;/code&gt; + capabilities per request in &lt;code&gt;_meta&lt;/code&gt;; new &lt;code&gt;server/discover&lt;/code&gt; RPC.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;subscriptions/listen&lt;/code&gt; replaces HTTP GET + &lt;code&gt;resources/subscribe&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ping&lt;/code&gt;, &lt;code&gt;logging/setLevel&lt;/code&gt;, &lt;code&gt;roots/list_changed&lt;/code&gt; removed.&lt;/li&gt;
&lt;li&gt;Tasks promoted to the official &lt;code&gt;io.modelcontextprotocol/tasks&lt;/code&gt; extension.&lt;/li&gt;
&lt;li&gt;MRTR (multi-request/multi-response) replaces server-initiated requests.&lt;/li&gt;
&lt;li&gt;Error-code allocation policy updated (−32020…−32099 reserved).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Source: &lt;code&gt;modelcontextprotocol.io/specification/2026-07-28/changelog&lt;/code&gt; (accessed 2026-10-01).&lt;/p&gt;

&lt;h2&gt;
  
  
  Point 1 — 186 remote endpoints, wire census (2026-10-01 05:27Z, re-probed 06:06Z)
&lt;/h2&gt;

&lt;p&gt;Cohort: the official MCP registry's v0 REST API, first 900 listings, deduplicated by endpoint URL, then the contiguous slice of hostnames beginning &lt;strong&gt;a–b&lt;/strong&gt; = 186 unique endpoints (a slice, not a random sample — see caveats). Probe: &lt;code&gt;POST initialize&lt;/code&gt; offering &lt;code&gt;protocolVersion: "2026-07-28"&lt;/code&gt;, browser User-Agent, classified by the negotiated response version and session header.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Class&lt;/th&gt;
&lt;th&gt;n&lt;/th&gt;
&lt;th&gt;Share of 186&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Auth-gated (HTTP 401) — invisible to anonymous probes&lt;/td&gt;
&lt;td&gt;86&lt;/td&gt;
&lt;td&gt;46%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answered old wire: 2025-11-25&lt;/td&gt;
&lt;td&gt;36&lt;/td&gt;
&lt;td&gt;19%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answered old wire: 2025-06-18&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;td&gt;11%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answered old wire: 2025-03-26&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answered old wire: 2024-11-05&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Answered new wire: 2026-07-28&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transient/errors&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;td&gt;15%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Of the 72 endpoints that returned any protocol version, &lt;strong&gt;7 (9.7%) spoke the new wire&lt;/strong&gt; — and the two-probe decomposition matters: 6 of the 7 were provably negotiating (they upgraded when offered 2026-07-28 and answered old when offered old), 1 was pinned. A re-probe 39 minutes later returned the identical class shape.&lt;/p&gt;

&lt;p&gt;The 7 are public registry entries and three of them share one operator — an operator-cluster, not seven independent early adopters: &lt;code&gt;mcp.getle.ad/mcp&lt;/code&gt;, &lt;code&gt;www.hood.ag/api/mcp&lt;/code&gt;, &lt;code&gt;mcp.bev-buyer.ai/mcp&lt;/code&gt;, &lt;code&gt;api.aislabs.ai/{,cve/,recorder/}mcp&lt;/code&gt;, &lt;code&gt;bankrolled.ai/mcp&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Point 2 — the prominent local servers (stdio, 2026-10-01)
&lt;/h2&gt;

&lt;p&gt;Ten prominent npm MCP servers, spawned via &lt;code&gt;npx -y&lt;/code&gt;, probed twice: &lt;code&gt;initialize&lt;/code&gt; (legacy handshake) and &lt;code&gt;server/discover&lt;/code&gt; (the new-revision call).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Server&lt;/th&gt;
&lt;th&gt;initialize&lt;/th&gt;
&lt;th&gt;server/discover&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;official &lt;code&gt;server-filesystem&lt;/code&gt; (0.2.0, 2026-08-31)&lt;/td&gt;
&lt;td&gt;2025-06-18&lt;/td&gt;
&lt;td&gt;−32601 Method not found&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;official &lt;code&gt;server-everything&lt;/code&gt; (2.0.0)&lt;/td&gt;
&lt;td&gt;2025-06-18&lt;/td&gt;
&lt;td&gt;−32601 Method not found&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;@upstash/context7-mcp&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2025-06-18&lt;/td&gt;
&lt;td&gt;−32601 Method not found&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;tavily-mcp&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2025-06-18&lt;/td&gt;
&lt;td&gt;−32601 Method not found&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;playwright-mcp&lt;/code&gt; (microsoft)&lt;/td&gt;
&lt;td&gt;2025-06-18&lt;/td&gt;
&lt;td&gt;−32601 Method not found&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;@modelcontextprotocol/server&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;no executable to run&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mcp-remote&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;OAuth-discovery hang (45s)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;All five answering servers speak only the 2025-06-18 wire and reject the new call. The flagship npm packages are the oldest wire in the ecosystem&lt;/strong&gt; — and the reference &lt;code&gt;@modelcontextprotocol/server&lt;/code&gt; package currently doesn't even ship a default executable (&lt;code&gt;npm error could not determine executable to run&lt;/code&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Point 3 — the same 186 endpoints, one week later, tested for the new call (2026-10-02 00:45Z)
&lt;/h2&gt;

&lt;p&gt;Same cohort, new probe design: &lt;code&gt;initialize&lt;/code&gt; requesting the &lt;strong&gt;floor&lt;/strong&gt; version (2025-06-18), then &lt;code&gt;server/discover&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Class&lt;/th&gt;
&lt;th&gt;count&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AUTH-401/403&lt;/td&gt;
&lt;td&gt;88 (47%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answered, no &lt;code&gt;server/discover&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;96&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Answered, accepts &lt;code&gt;server/discover&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DNS/5xx/redirects&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 3 acceptors (&lt;code&gt;mcp.getle.ad/mcp&lt;/code&gt;, &lt;code&gt;analyticslegends.ai/mcp&lt;/code&gt;, &lt;code&gt;api.askmiles.ai/mcp&lt;/code&gt;) all return the new-era discover payload. &lt;strong&gt;Combined across both cohorts: 8 of 109 answering endpoints (7.3%) accept the new-revision call&lt;/strong&gt; while major clients already ship it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The version-echo lesson (the methodology finding that will outlive the numbers)
&lt;/h2&gt;

&lt;p&gt;Our 10-01 census reported 43 servers "speaking" 2025-11-25 or 2026-07-28. On 10-02 we requested the &lt;strong&gt;floor&lt;/strong&gt; version and all 43 echoed the floor back — none "regressed". The servers &lt;strong&gt;negotiate&lt;/strong&gt;: they echo any requested version within their supported range and report their maximum for requests above it. Verified on four servers that day:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;api.aislabs.ai/mcp&lt;/code&gt;, &lt;code&gt;bankrolled.ai/mcp&lt;/code&gt; — permissive: accept every version 2024-11-05 → 2026-07-28&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;mcp.getle.ad/mcp&lt;/code&gt; — floor 2025-03-26&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;advisorsai.ai/mcp&lt;/code&gt; — &lt;strong&gt;capped at 2025-11-25&lt;/strong&gt; (a 2026-07-28 offer gets a 2025-11-25 answer)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So a server's "reported protocolVersion" is only a capability test when you request the &lt;em&gt;highest&lt;/em&gt; version you care about and read the answer: echo = speaks it; lower echo = caps below; &lt;code&gt;UnsupportedProtocolVersionError&lt;/code&gt; = hard floor above you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure mode
&lt;/h2&gt;

&lt;p&gt;A reader hard-coded to one era sees half the population as missing — and nothing errors. The identical lesson landed on the x402 side this week (payment requirements in the 402 &lt;strong&gt;body&lt;/strong&gt; in v1, base64 in the &lt;code&gt;PAYMENT-REQUIRED&lt;/code&gt; &lt;strong&gt;header&lt;/strong&gt; in v2: one operator's body-only parser marked 87 of 183 doors "no payTo" when all 87 had one). Across both protocols the same shape: &lt;strong&gt;the migration changes where the data lives, not whether it lives, and only the reader's vantage point decides what is invisible.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Caveats (honest)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The 186-URL cohort is a hostname slice of the registry's default ordering, not a random sample.&lt;/li&gt;
&lt;li&gt;~47% of the cohort is auth-gated — the public surface is the visible half.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;server/discover&lt;/code&gt; is defined in the 2026-07-28 Final spec; the project blog and spec disagree on its optionality, so the 3 acceptors might be 2025-11-25 servers.&lt;/li&gt;
&lt;li&gt;Pinned ≠ behaviorally old: a pinned 2025-11-25 server may still accept 2026 clients.&lt;/li&gt;
&lt;li&gt;Two dated points, one week apart — a direction, not a trend.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Next
&lt;/h2&gt;

&lt;p&gt;The weekly re-probe of this exact cohort now requests &lt;strong&gt;2026-07-28&lt;/strong&gt; and reads echo/cap/error, which finally separates "pure new wire" from "old handshake, new version" and gives the series its first real migration curve. Next point: 2026-10-08.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Pennyforge (one-person studio). Method: &lt;code&gt;mcp_census.py&lt;/code&gt; + &lt;code&gt;probe-remote.py&lt;/code&gt;; per-URL classes and the dated probe notes (10-01 wire census, 10-02 remote probe) are kept as research artifacts next to the method scripts. Data: public registry + public endpoints, $0.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>api</category>
      <category>showdev</category>
      <category>web3</category>
    </item>
    <item>
      <title>The EU's CRA guidance is 84 pages long and never says "SBOM" — a field report</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Fri, 02 Oct 2026 01:05:38 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/the-eus-cra-guidance-is-84-pages-long-and-never-says-sbom-a-field-report-43jp</link>
      <guid>https://dev.to/pennyforgehq/the-eus-cra-guidance-is-84-pages-long-and-never-says-sbom-a-field-report-43jp</guid>
      <description>&lt;p&gt;&lt;strong&gt;Date:&lt;/strong&gt; 2026-10-02 · &lt;strong&gt;Method:&lt;/strong&gt; primary sources only (EUR-Lex + European Commission, both retrieved 2026-10-02) · &lt;strong&gt;Framing:&lt;/strong&gt; what a 5-person connected-device firm actually finds when it "follows the guidance"&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;The Cyber Resilience Act (Regulation (EU) 2024/2847) entered into force 10 December 2024. Two dated obligations matter right now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reporting obligations are ALREADY in force — since 11 September 2026&lt;/strong&gt; (three weeks ago). Manufacturers must report actively exploited vulnerabilities and security incidents to the national coordination centre (ENISA) on 24-hour/14-day clocks.&lt;/li&gt;
&lt;li&gt;The main requirements apply from &lt;strong&gt;11 December 2027&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On &lt;strong&gt;27 July 2026&lt;/strong&gt; the Commission published its flagship non-binding guidance — Communication C(2026) 5252 with an &lt;strong&gt;84-page annex&lt;/strong&gt; — explicitly pitched at "manufacturers, developers, and businesses of all sizes", with 67 practical examples and "particular attention" to microenterprises and SMEs. That annex is the document a small device maker is supposed to follow.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the regulation itself says about SBOMs
&lt;/h2&gt;

&lt;p&gt;EUR-Lex text of Regulation (EU) 2024/2847 (retrieved 2026-10-02):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Art. 3(39): defines "software bill of materials".&lt;/li&gt;
&lt;li&gt;Recital 24: the Commission &lt;strong&gt;may&lt;/strong&gt; specify "the format and elements of the software bill of materials" by implementing act. (None has been issued as of 2026-10-02 — the format remains unspecified in law.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Annex I, Part II, point 1&lt;/strong&gt; (documentation requirements, applicable to all products with digital elements): manufacturers shall "identify and document vulnerabilities and components contained in products with digital elements, &lt;strong&gt;including by drawing up a software bill of materials in a commonly used and machine-readable format&lt;/strong&gt; covering at the very least the top-level dependencies of the products."&lt;/li&gt;
&lt;li&gt;"bill of materials" occurs &lt;strong&gt;7 times&lt;/strong&gt; in the regulation.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What the Commission's own 84-page guidance says
&lt;/h2&gt;

&lt;p&gt;I extracted the full text of the C(2026) 5252 annex (84 pages, ~251k characters) and counted:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;term&lt;/th&gt;
&lt;th&gt;hits in 84-page guidance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"bill of materials" / "SBOM" / "BOM"&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"CycloneDX"&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"SPDX"&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"firmware"&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (and it is a spare-parts exemption example, not an update-cycle requirement)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"machine-readable"&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (in the security-fix-sharing context of Art. 13(6), not the SBOM)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The guidance discusses Annex I heavily ("Annex I": 35 hits, "Part II": 12 hits, "technical documentation": 13 hits) and the exact regulatory phrase "identify and document" appears &lt;strong&gt;0 times&lt;/strong&gt;. It covers scope, substantial modification, support periods, reporting mechanics and risk assessment in depth — and leaves the SBOM, the one deliverable the regulation names with a format requirement, to… the reader.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for the 5-person device firm
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The law tells you to produce an SBOM "in a commonly used and machine-readable format." The Commission's guidance never names which formats are commonly used. You must choose CycloneDX vs SPDX yourself (industry default: CycloneDX for component SBOMs, SPDX for license metadata) with no Commission blessing either way.&lt;/li&gt;
&lt;li&gt;The "simplified technical documentation form targeted at the needs of micro- and small enterprises" promised on the Commission's own MSME page is still a &lt;strong&gt;may&lt;/strong&gt; — not issued as of its 31 July 2026 update.&lt;/li&gt;
&lt;li&gt;The 10 funded EU projects listed on the same page (OCCTET, CONFIRMATE, CRACY, CYBERFORT, CURIUM, OSCRAT, CRA-AI, SECURE, STAN4CR, CYBERSTAND) are mostly research-grade tools — none is the "the" SME SBOM answer.&lt;/li&gt;
&lt;li&gt;The reporting obligation that is already in force (since 2026-09-11) is the least-discussed dated hook: most small firms do not know their 24h/14d vulnerability-reporting duty started three weeks before the guidance even shipped.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Competing view / caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;"0 hits" is a search over extracted PDF text (pypdf); the term could appear in a figure I did not OCR. The "firmware: 1 hit" context was manually verified, which makes the extraction trustworthy, but a human should re-verify the SBOM count against the PDF before this is published.&lt;/li&gt;
&lt;li&gt;One could argue the guidance deliberately defers SBOM detail to the future implementing act (Recital 24). That is plausible — but then the gap is even larger: the law demands it now (in 14 months), the format is unspecified, and the flagship guidance stays silent.&lt;/li&gt;
&lt;li&gt;ENISA's old CRA topic page (enisa.europa.eu/topics/cyber-resilience-act) 404s after a site restructure; ENISA's MSME angle currently points to its pre-CRA "Secure by Design and Default Playbook".&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;EUR-Lex, Regulation (EU) 2024/2847 (OJ version): &lt;a href="https://eur-lex.europa.eu/eli/reg/2024/2847/oj" rel="noopener noreferrer"&gt;https://eur-lex.europa.eu/eli/reg/2024/2847/oj&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Commission CRA policy page (last update 7 September 2026): &lt;a href="https://digital-strategy.ec.europa.eu/en/policies/cyber-resilience-act" rel="noopener noreferrer"&gt;https://digital-strategy.ec.europa.eu/en/policies/cyber-resilience-act&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Commission guidance announcement, 27 July 2026: &lt;a href="https://digital-strategy.ec.europa.eu/en/library/commission-publishes-new-guidance-support-timely-cyber-resilience-act-implementation" rel="noopener noreferrer"&gt;https://digital-strategy.ec.europa.eu/en/library/commission-publishes-new-guidance-support-timely-cyber-resilience-act-implementation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Guidance annex PDF (C(2026) 5252): &lt;a href="https://ec.europa.eu/newsroom/dae/redirection/document/131456" rel="noopener noreferrer"&gt;https://ec.europa.eu/newsroom/dae/redirection/document/131456&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;MSME card (last update 31 July 2026): &lt;a href="https://digital-strategy.ec.europa.eu/en/policies/cra-msmes" rel="noopener noreferrer"&gt;https://digital-strategy.ec.europa.eu/en/policies/cra-msmes&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Pennyforge (one-person studio) · full text extractions + probe data available on request · AI-assisted research, studio-owned&lt;/em&gt;&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>eu</category>
      <category>cybersecurity</category>
      <category>compliance</category>
    </item>
    <item>
      <title>Who's actually blocking AI crawlers in Europe? A 185-site census (first dated point)</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Thu, 01 Oct 2026 23:55:52 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/whos-actually-blocking-ai-crawlers-in-europe-a-185-site-census-first-dated-point-3a59</link>
      <guid>https://dev.to/pennyforgehq/whos-actually-blocking-ai-crawlers-in-europe-a-185-site-census-first-dated-point-3a59</guid>
      <description>&lt;p&gt;&lt;strong&gt;Dated: 2026-10-01, 17:15–17:22 UTC — single crawl pass, fixed panel, fixed method.&lt;/strong&gt;&lt;br&gt;
Pennyforge research artifact. $0 data (public robots.txt + &lt;code&gt;/.well-known/ai-crawler&lt;/code&gt;, plain HTTP with a browser user-agent). Companion data: &lt;code&gt;.tmp-sweep/r141/tdm-census-v2-2026-10-01.json&lt;/code&gt; (per-site results, full 12-UA signal table, method).&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this question now
&lt;/h2&gt;

&lt;p&gt;Every AI lab says it respects &lt;code&gt;robots.txt&lt;/code&gt; for its training crawlers (GPTBot, ClaudeBot, Bytespider) — but &lt;code&gt;robots.txt&lt;/code&gt; is a voluntary convention (RFC 9309, 2022), written for a web that had one kind of crawler. In 2026 there are at least three distinct jobs: &lt;strong&gt;training crawlers&lt;/strong&gt; (bulk dataset collection), &lt;strong&gt;search-index crawlers&lt;/strong&gt; (citation/answer engines), and &lt;strong&gt;on-demand fetchers&lt;/strong&gt; (one user, one page). Publishers increasingly want to allow one and block the others — and the European Commission's open copyright×AI consultation (closes &lt;strong&gt;2026-11-03&lt;/strong&gt;) is asking precisely what a text-data-mining opt-out signal should be. Meanwhile the web is growing unpoliced "manifest" conventions — &lt;code&gt;llms.txt&lt;/code&gt;, &lt;code&gt;ai.txt&lt;/code&gt; builders, &lt;code&gt;human.json&lt;/code&gt; (mocked as "a blogroll wearing a protocol's costume" in a May 2026 essay), and a &lt;code&gt;.well-known/ai-crawler&lt;/code&gt; file nobody standardised.&lt;/p&gt;

&lt;p&gt;So: who has actually written any rules? And does anything exist beyond &lt;code&gt;robots.txt&lt;/code&gt;? We measured, on a fixed panel, with a fixed method, on one dated pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  The panel (185 sites)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;23 major NL/DE retail shops + 8 Dutch public-interest sites (kept from an earlier 31-site slice for comparability)&lt;/li&gt;
&lt;li&gt;66 EU retail shops (FR/IT/ES/PT/PL/AT/SE/BE + more NL/DE)&lt;/li&gt;
&lt;li&gt;24 EU + national government portals (EU institutions, UK, FR, DE, PL, IT, ES, AT, IE, SE, FI, CZ, LT, LV, SK, SI, HR, GR, PT, RO, HU, EE)&lt;/li&gt;
&lt;li&gt;30 EU media/publishers (DE/FR/IT/ES/PT/PL/GR/BE/AT/FI)&lt;/li&gt;
&lt;li&gt;34 tech, AI and scholarly-infrastructure sites (OpenAI, Anthropic, Perplexity, Mistral, Hugging Face, arXiv, GitHub, the major publishers' platforms, Crossref, Zenodo…)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For each site: (1) parse &lt;code&gt;robots.txt&lt;/code&gt; for a fixed &lt;strong&gt;12-UA list&lt;/strong&gt; (GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot, CCBot, Google-Extended, Applebot-Extended, Bytespider, Meta-ExternalAgent, MistralAI-SearchBot); (2) probe &lt;code&gt;/.well-known/ai-crawler&lt;/code&gt; and classify the response by content type (a 200 with &lt;code&gt;text/html&lt;/code&gt; is an SPA catch-all, &lt;strong&gt;not&lt;/strong&gt; a file).&lt;/p&gt;

&lt;h2&gt;
  
  
  The findings
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Two-thirds publish &lt;code&gt;robots.txt&lt;/code&gt;; most of it says nothing about AI.&lt;/strong&gt; 121/185 sites (65.4%) serve a robots.txt; &lt;strong&gt;77 of those (64%) contain zero entries for any of the 12 AI crawlers.&lt;/strong&gt; 39 sites (21%) answer &lt;strong&gt;403&lt;/strong&gt; to the file (CDN/bot walls — zalando.de, ad.nl, sciencedirect.com, mdpi.org, several big retailers); by the protocol's own convention a 403 means "temporarily unobtainable", so their policy is effectively invisible to polite crawlers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. 35 of 185 sites (18.9%) have a hard AI-crawler opt-out — and every one of them is a full &lt;code&gt;Disallow: /&lt;/code&gt;.&lt;/strong&gt; No partial path-level opt-outs anywhere on the panel. The hard opt-outs cluster violently by sector:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Sector&lt;/th&gt;
&lt;th&gt;Sites with hard opt-out&lt;/th&gt;
&lt;th&gt;Rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;EU media/publishers&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20 of 30&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;66.7%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dutch public-interest (orig.)&lt;/td&gt;
&lt;td&gt;2 of 8&lt;/td&gt;
&lt;td&gt;25.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NL/DE retail (orig. panel)&lt;/td&gt;
&lt;td&gt;4 of 23&lt;/td&gt;
&lt;td&gt;17.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tech / AI / scholarly&lt;/td&gt;
&lt;td&gt;4 of 34&lt;/td&gt;
&lt;td&gt;11.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EU retail (new 66)&lt;/td&gt;
&lt;td&gt;4 of 66&lt;/td&gt;
&lt;td&gt;6.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Government portals&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1 of 24&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4.2%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Publishers are blocking; governments and shops are mostly silent. The near-total opt-outs (10 of 12 UAs) are: &lt;strong&gt;amazon.de, faz.net, derstandard.at, rtbf.be, lalibre.be&lt;/strong&gt; — with spiegel.de, ilsole24ore.com and corriere.it at 8/12. Government is almost uniformly quiet; the only gov hit is &lt;strong&gt;gov.uk blocking Meta-ExternalAgent&lt;/strong&gt; alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. CCBot is the most-blocked crawler on the panel (26 sites), with zero explicit allows.&lt;/strong&gt; The fear is Common Crawl's role as a training-data conduit, not any single lab: GPTBot 22, ClaudeBot 18, Google-Extended 17, Bytespider 17, Meta-ExternalAgent 16 full disallows — but CCBot is the only one with &lt;strong&gt;0&lt;/strong&gt; sites explicitly allowing it. The opt-in side is thin overall: only &lt;strong&gt;9 of 185&lt;/strong&gt; sites carry an explicit &lt;code&gt;Allow&lt;/code&gt; for at least one AI user-agent, and most of them are consumer-electronics retailers (mediamarkt.nl/.at, saturn.de, mediaworld.it, mediaexpert.pl, morele.net) — the visible-in-AI-search incentive, not a content policy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. &lt;code&gt;MistralAI-SearchBot&lt;/code&gt; has no rule anywhere: 0 of 185.&lt;/strong&gt; The newest major search bot on our list has no allow, no disallow, no mention on the entire panel. Policy is written for yesterday's crawlers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. &lt;code&gt;.well-known/ai-crawler&lt;/code&gt; has zero real adoption.&lt;/strong&gt; 20 sites returned HTTP 200 for the path — and &lt;strong&gt;all 20 were &lt;code&gt;text/html&lt;/code&gt; SPA catch-alls&lt;/strong&gt; (content-type verified; e.g. dm.de, government.fr, digi.ee, springer.com). &lt;strong&gt;0 of 185 serve an actual file.&lt;/strong&gt; The path is a convention, not a protocol: nothing requires it, nothing specifies its format, and on this panel nobody serves one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means
&lt;/h2&gt;

&lt;p&gt;The de facto standard for TDM opt-out in Europe is a &lt;strong&gt;fragmented, voluntary robots.txt with per-UA lines&lt;/strong&gt; — enforced only if the crawler chooses to check. The strongest signal in the data is the &lt;strong&gt;66.7% of EU media&lt;/strong&gt; that have drawn a hard line (all full-disallow), against &lt;strong&gt;4.2% of government portals&lt;/strong&gt; that have drawn any. The emerging "manifest file" family (&lt;code&gt;llms.txt&lt;/code&gt;/&lt;code&gt;ai.txt&lt;/code&gt;/&lt;code&gt;human.json&lt;/code&gt;/&lt;code&gt;.well-known/ai-crawler&lt;/code&gt;) adds intent documentation but no enforcement and — for &lt;code&gt;.well-known/ai-crawler&lt;/code&gt; at least — &lt;strong&gt;zero adoption&lt;/strong&gt; on a 185-site EU-heavy panel. That is exactly the gap a standardised, machine-readable TDM opt-out would fill, and it is what the EC consultation is debating while the window is open.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest caveats
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Single pass, single date, single (browser) user-agent; a 403 wall reads as "no policy visible", not "no policy exists".&lt;/li&gt;
&lt;li&gt;Brand-biased panel (major retailers, national media, EU portals) — an indicative cross-section, not a random sample; long-tail EU websites are not represented.&lt;/li&gt;
&lt;li&gt;Naive group-based parser (per-UA group heuristic; the RFC's longest-path-match precedence is not fully implemented — affects at most the one "partial" case found).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;robots.txt&lt;/code&gt; is voluntary: a disallow is an expression of intent, not a technical block.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What does NOT exist yet
&lt;/h2&gt;

&lt;p&gt;A dated, multi-country, multi-sector census of TDM opt-out signals with a fixed panel and method. This artifact is the first dated point; the series continues quarterly with the same method, and the next pass extends the panel past 300 sites and adds the allow-side (opt-in) signal as its own measure.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Pennyforge (one-person studio) · per-site evidence in the companion JSON · re-run script on request · AI-assisted research, studio-owned&lt;/em&gt;&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>ai</category>
      <category>web</category>
      <category>data</category>
    </item>
    <item>
      <title>What the EU AI Act actually requires of a two-person studio in 2026 — and the database problem nobody mentions</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:41:16 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/what-the-eu-ai-act-actually-requires-of-a-two-person-studio-in-2026-and-the-database-problem-36lk</link>
      <guid>https://dev.to/pennyforgehq/what-the-eu-ai-act-actually-requires-of-a-two-person-studio-in-2026-and-the-database-problem-36lk</guid>
      <description>&lt;p&gt;&lt;strong&gt;Dated: 2026-10-01. Every date below verified against primary or first-tier sources that day; sources at the end.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you build or deploy a high-risk AI system in the EU in 2026, you are caught in a gap that has a date on every edge of it. This post lays out the gap end-to-end, because the pieces are spread across a regulation, two Commission documents, a service-desk message and a press report.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cliff in one paragraph
&lt;/h2&gt;

&lt;p&gt;A small studio that ships an AI system falling under Annex III (e.g. AI-assisted CV screening, point 4(b)) can escape the full high-risk regime &lt;strong&gt;if&lt;/strong&gt; it qualifies for the Article 6(3) filter — but using that filter requires &lt;strong&gt;a written four-part self-assessment before launch and registration in a central EU database that, by the Commission's own words, is not yet open&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The dated chain
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;2026-02-02&lt;/strong&gt; — statutory deadline for the Commission's Article 6(5) guidelines on classifying high-risk systems.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-05-19&lt;/strong&gt; — the guidelines are published. As draft. ~3.5 months late.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-07-23&lt;/strong&gt; — the draft's last update (the EC library page still marks it draft; a targeted consultation is running).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-07-09&lt;/strong&gt; — the Commission's own AI Act service desk tells inquirers: &lt;em&gt;"This database is not yet open and operational."&lt;/em&gt; (the Article 71 EU database, via a Rapporteur/Euractiv report of 2026-07-27; launch expected Q3 2027.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;December 2027&lt;/strong&gt; — when, per the service desk, the registration duty actually starts applying, aligned with the Digital Omnibus "adjusted timeline" for high-risk systems.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;…or 2026-08-02&lt;/strong&gt; — because the Omnibus did &lt;em&gt;not&lt;/em&gt; move the standalone article that creates the registration duty, a 2026 start date is legally plausible. The Commission declined to confirm.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So a studio following the draft guidelines (step 3) is told to register a system (step 4's database) before the database exists (steps 4–6) — with the duty-start date genuinely ambiguous between "next week" and "15 months from now".&lt;/p&gt;

&lt;h2&gt;
  
  
  What the draft guidelines actually say (the part people skip)
&lt;/h2&gt;

&lt;p&gt;From the draft's Annex III section (148 pages, §2.7, ¶84–117):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The Article 6(3) filter is a &lt;strong&gt;self-assessment by the provider&lt;/strong&gt; — there is no independent check, and the four conditions (narrow procedural task; improving a previously completed human activity; detecting decision-making patterns without replacing human assessment; preparatory task) are exhaustive, alternative, and "must be interpreted narrowly".&lt;/li&gt;
&lt;li&gt;But using the filter triggers &lt;strong&gt;Article 6(4)&lt;/strong&gt;: you must &lt;strong&gt;(i) document the assessment before the system is placed on the market&lt;/strong&gt; and &lt;strong&gt;(ii) register the system in the Article 71 EU database&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The documentation must contain four things: the intended purpose; why the system would be high-risk under 6(2); which 6(3) condition(s) apply and why; and &lt;strong&gt;why the system does not perform profiling&lt;/strong&gt; (profiling kills the filter).&lt;/li&gt;
&lt;li&gt;It must be producible &lt;strong&gt;at any time&lt;/strong&gt; on a market-surveillance authority's request — and Article 80 gives those authorities power to re-evaluate the classification and, on misclassification, to impose Article 99 penalties.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is the practical shape of the cliff: not 40 pages of new obligations, but &lt;strong&gt;four paragraphs of legal self-assessment plus a registration in a database that doesn't exist yet&lt;/strong&gt;, enforceable from day one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means if you are a two-person studio
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;If you're &lt;strong&gt;outside Annex III&lt;/strong&gt; (most B2B SaaS, content tools): nothing high-risk applies to you in 2026 beyond the already-in-force Article 50 transparency duties (mark AI-generated text, disclose chatbots).&lt;/li&gt;
&lt;li&gt;If you're &lt;strong&gt;inside Annex III and can document a 6(3) condition&lt;/strong&gt;: write the four-part assessment now (it's a day of work, not a project), and track the database — your registration is a line item that must exist before launch, wherever "launch" is dated.&lt;/li&gt;
&lt;li&gt;If you're &lt;strong&gt;inside Annex III and can't&lt;/strong&gt; (you do profile people, you make the decision): the full Chapter III regime is the plan, and the Omnibus timeline (not the old 2026-08-02 date) is what your compliance calendar should follow — but the registration-article ambiguity is an exception to that comfort.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Honest caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The guidelines are a &lt;strong&gt;draft&lt;/strong&gt;; the final version can move the filter's conditions or documentation requirements.&lt;/li&gt;
&lt;li&gt;The "2026-08-02 vs December 2027" conflict rests on the Omnibus not amending one standalone article; the Commission has not ruled, and a small clarifying act could close it either way.&lt;/li&gt;
&lt;li&gt;I am one studio reading the documents, not a law firm; this is a dated map, not advice.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources (all checked 2026-10-01)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;EC draft guidelines, Annex III section (PDF, 148 pp) — digital-strategy.ec.europa.eu library page, published 2026-05-19, last update 2026-07-23 (¶84–117 = the filter section).&lt;/li&gt;
&lt;li&gt;Rapporteur (Euractiv) 2026-07-27: "EU's high-risk AI database pushed back to mid-to-late 2027" — service desk message of 2026-07-09, Q3-2027 launch, Dec-2027 duty start, un-moved registration article.&lt;/li&gt;
&lt;li&gt;AI Act Service Desk, Article 71 explainer (ai-act-service-desk.ec.europa.eu) — database design per the regulation.&lt;/li&gt;
&lt;li&gt;EUR-Lex, Regulation (EU) 2024/1689, Art. 6 + Annex III (text checked in an earlier pass of this series).&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;This work was done by Pennyforge, a one-person studio: the analysis was AI-assisted with human editorial review, as we practice it across all our research posts.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>ai</category>
      <category>law</category>
      <category>regulation</category>
    </item>
    <item>
      <title>iDEAL Wero: how far has the merchant migration really gone? (first dated point)</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:22:27 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/ideal-wero-how-far-has-the-merchant-migration-really-gone-first-dated-point-505k</link>
      <guid>https://dev.to/pennyforgehq/ideal-wero-how-far-has-the-merchant-migration-really-gone-first-dated-point-505k</guid>
      <description>&lt;p&gt;&lt;strong&gt;Dated: 2026-10-01 — census v1 07:12Z (static HTML) + v2 09:53–09:54Z (headless browser); incumbent checks 10:05–10:20Z.&lt;/strong&gt;&lt;br&gt;
Pennyforge research artifact. $0 data (public shop pages, primary PSP/EPI/association pages). Per-shop results and the method are available on request; the re-run script too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context (primary sources, verified 2026-10-01)
&lt;/h2&gt;

&lt;p&gt;iDEAL is the Netherlands' dominant online bank-payment method; Wero (European Payments Initiative) is its successor. The iDEAL brand owner (ideal.nl/naar-wero) says the migration runs &lt;strong&gt;until end-2027&lt;/strong&gt; ("Uiterlijk eind 2027 is iedere iDEAL-acceptant overgestapt naar Wero"), hedged "onder voorbehoud van verdere afstemming met de toezichthouder". Dated milestones (CM.com merchant guide + EPI confirmation 2026-07-16, with Thuiswinkel.org / eCommerce Europe / Dutch Payments Association):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;2026-01-08 media campaign · 2026-01-15 bank consumer communications&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026-01-29 → 2026-03-31: co-branded "iDEAL | Wero" logo rollout&lt;/strong&gt; (the window is closed — ~2 months ago)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Oct 2026: all Dutch issuing banks on Wero&lt;/strong&gt; (issuer milestone)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;31 Dec 2027: full migration, iDEAL phased out&lt;/strong&gt;; purchase protection full coverage targeted 2028-01-01&lt;/li&gt;
&lt;li&gt;Third-party (Clearing Post tracker, updated Sep 2026): &lt;strong&gt;NL online payments "Planned Dec 2026"&lt;/strong&gt;; EPI claims 60M registered users and &lt;strong&gt;48,000+ merchants have processed Wero transactions (2026-09-16)&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The census (24 major NL/DE shops)
&lt;/h2&gt;

&lt;p&gt;Method v1: curl of homepage + payment/help pages, substring &lt;code&gt;wero&lt;/code&gt;/&lt;code&gt;ideal&lt;/code&gt; on static HTML — &lt;strong&gt;a JS-rendered absence is NOT an absence&lt;/strong&gt;. Method v2: headless Chromium (local) DOM text on the shops v1 could not see.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Class&lt;/th&gt;
&lt;th&gt;v1 (static)&lt;/th&gt;
&lt;th&gt;v2 (after browser pass)&lt;/th&gt;
&lt;th&gt;Shops (v2)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;WERO + IDEAL (dual branding)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;5 (21%)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;coolblue.nl, mediamarkt.nl, action.nl, lidl.de, &lt;strong&gt;bol.com&lt;/strong&gt; (browser-verified on its official payment page: "Betaal via iDEAL | Wero, achteraf betalen, creditcard of een cadeaubon.")&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;IDEAL-ONLY&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;apple.com/nl, aboutyou.nl, etos.nl, apple.com/de, c-and-a.com&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;NEITHER in static HTML&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;wehkamp.nl (Cloudflare wall, &amp;gt;10s), dm.de (JS SPA shell), decathlon.nl, vink.com, hema.nl, amazon.de, zalando.de, otto.de, mediamarkt.de, rossmann.de, aboutyou.de, kaufland.de&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;UNCHECKED&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;albertheijn.com (JS redirect shell to /lander), h&amp;amp;m.com/de (fetch ERR)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Reading:&lt;/strong&gt; with the co-branding window closed 2026-03-31, &lt;strong&gt;5 of the 24 largest NL/DE shops (including the largest NL marketplace, bol.com) already show dual branding, and 5 are still iDEAL-only&lt;/strong&gt; — the migration is visibly in progress but far from uniform. The 12 "neither-in-static-HTML" shops need deeper checkout passes (most render payment logos via JS or require login); that is the next census pass, not a verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  The incumbent tension (verified 2026-10-01)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mollie's Wero page&lt;/strong&gt; (mollie.com/payments/wero): status "&lt;strong&gt;Coming soon&lt;/strong&gt;", available in &lt;strong&gt;DE, FR, BE — NL not listed&lt;/strong&gt;, NL pricing already shown (€0.32/txn), sidebar links "iDEAL | Wero" to the iDEAL page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Adyen&lt;/strong&gt; (knowledge hub, 2026-07-06): "one of the first PSPs to offer Wero" (FR/BE/LU/DE/&lt;strong&gt;NL&lt;/strong&gt;/AT), but "adoption is still in its early stages" and issuer coverage "hasn't yet been reached consistently".&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The iDEAL brand owner says the migration has already started; a major NL PSP still lists the Netherlands as not-yet-available. Dated primary-adjacent quotes, both retrievable today.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does NOT exist yet (the refined negative claim)
&lt;/h2&gt;

&lt;p&gt;PSP marketing pages, EPI/association news, and one &lt;strong&gt;country-level&lt;/strong&gt; tracker (Clearing Post) exist — but &lt;strong&gt;no dated, third-party, per-merchant migration census with a fixed public method&lt;/strong&gt;. This artifact is the first dated point of exactly that: a fixed shop list, a fixed method, a timestamp. The series continues quarterly (same 24 shops + replacements).&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest caveats
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;24 shops is an indicative panel, not a population; DE shops are included because the Wero rollout is pan-EU.&lt;/li&gt;
&lt;li&gt;Static-HTML substring + single browser pass ≠ checkout-level truth; "dual branding visible" is a branding signal, not a contract signal.&lt;/li&gt;
&lt;li&gt;SERP/incumbent pages are self-reported by the vendors with a stake.&lt;/li&gt;
&lt;li&gt;The "end-2027" deadline carries the iDEAL owner's own hedge (regulator alignment).&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;This work was done by Pennyforge, a one-person studio: the analysis was AI-assisted with human editorial review, as we practice it across all our research posts. Per-shop evidence and the re-run script on request.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>payments</category>
      <category>fintech</category>
      <category>netherlands</category>
    </item>
    <item>
      <title>Do published reference lists still cite retracted papers? A dated OpenAlex audit</title>
      <dc:creator>Pennyforge</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:20:46 +0000</pubDate>
      <link>https://dev.to/pennyforgehq/do-published-reference-lists-still-cite-retracted-papers-a-dated-openalex-audit-16me</link>
      <guid>https://dev.to/pennyforgehq/do-published-reference-lists-still-cite-retracted-papers-a-dated-openalex-audit-16me</guid>
      <description>&lt;p&gt;&lt;strong&gt;Dated: 2026-10-01 (gate 05:13Z, old-lists 07:10Z, seeded control 09:50Z — same day, same pipeline version).&lt;/strong&gt;&lt;br&gt;
Pennyforge research artifact. $0 data (PubMed eutils + OpenAlex polite pool + Crossref production API). All data and the per-DOI results are available on request; the whole pipeline is four small scripts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question
&lt;/h2&gt;

&lt;p&gt;Every paper's reference list is a small, checkable corpus of claims that other works exist and carry authority. A retracted paper still cited in a published list = silent authority leakage. We tested whether &lt;strong&gt;OpenAlex's &lt;code&gt;is_retracted&lt;/code&gt; flag is good enough to audit real published reference lists&lt;/strong&gt; — first on a gold set (gate), then on live 2019–2025 lists, and finally with a seeded control that proves the pipeline catches a real case by construction.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Gate: OpenAlex vs a Retraction Watch gold set
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gold set:&lt;/strong&gt; Retraction Watch dataset via Crossref (72,790 rows, generated 2026-09-30, daily-updated). Sample: &lt;strong&gt;120 retractions, 10 per publication year 2015–2026&lt;/strong&gt;, deterministic stride.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAlex:&lt;/strong&gt; 120/120 found, &lt;strong&gt;115/120 flagged retracted → 95.8% recall&lt;/strong&gt; on the sample.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Crossref (authoritative source, &lt;code&gt;updated-by&lt;/code&gt; retraction events):&lt;/strong&gt; 117/120 found, &lt;strong&gt;111/120 carry a retraction event → 7.5% of even the gold-set retractions are missing from the authoritative source itself&lt;/strong&gt; (inherited gaps the audit cannot fix).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agreement OpenAlex vs Crossref (117 both found): 116/117 = 99.1%.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Practical note:&lt;/strong&gt; OpenAlex exposes only a boolean &lt;code&gt;is_retracted&lt;/code&gt; — &lt;strong&gt;no retraction date&lt;/strong&gt;. Dated notices come from Crossref &lt;code&gt;updated-by&lt;/code&gt; (retraction-watch records, production schema since 2025-01).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. Live lists: 0 hits, honestly framed
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;2025 lists (run 138): 0 retracted references in 322 DOIs.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;2019–2020 lists (run 139): 8 published systematic-review reference lists (PubMed), 812 references, 606 unique DOIs, 599 resolved (98.8%) → 0 flagged retracted.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Honest frame:&lt;/strong&gt; a retraction base rate near ~1% makes 0/928 an upper bound (~0.3% at 95% CI), not evidence of absence. The live-list pass is a &lt;em&gt;negative within its confidence interval&lt;/em&gt; — exactly what a dated audit should report.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Seeded control: the pipeline finds a real case by construction
&lt;/h2&gt;

&lt;p&gt;To prove the pipeline is not blind, we seeded the hardest case: a retracted, well-cited review that must still appear in a recent list.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Seed:&lt;/strong&gt; 10.1038/s41598-023-28418-1 — "RETRACTED ARTICLE: Exogenous melatonin alleviates neuropathic pain-induced affective disorders…" (Sci Rep 2023, cited_by 50).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Citer:&lt;/strong&gt; 10.1007/s11064-026-04725-7 — Mol Neurobiol 2026 (PMID 41824110); its published reference list &lt;strong&gt;contains the seed's DOI&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Result:&lt;/strong&gt; 29 refs / 28 DOIs, 28/28 resolved; &lt;strong&gt;seed flagged &lt;code&gt;is_retracted: true&lt;/code&gt;&lt;/strong&gt; and &lt;strong&gt;Crossref dated retraction notice 2026-08-10&lt;/strong&gt;; &lt;strong&gt;0 false positives&lt;/strong&gt; in the same list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost:&lt;/strong&gt; confirmation on the &lt;strong&gt;21st (seed, citer) attempt&lt;/strong&gt; — the pipeline's confirmation rate is ~1-in-20 because it only sees a retraction when the citer's reference text carries the seed's DOI string.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What this is worth
&lt;/h2&gt;

&lt;p&gt;A reproducible, $0, API-only audit of any published reference list: &lt;strong&gt;gate 95.8% recall / 99.1% agreement, dated notices via Crossref, and a constructive proof the pipeline fires.&lt;/strong&gt; It converts "is this citation still trustworthy?" from a manual hunt into a batch job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations (the ones that matter)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;DOI-in-text only:&lt;/strong&gt; the audit flags a retraction only if the cited DOI appears in the reference list text. ~1-in-20 (seed, citer) confirmation rate (21 attempts to one hit).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No dates at OpenAlex:&lt;/strong&gt; dates require the Crossref &lt;code&gt;updated-by&lt;/code&gt; cross-check (adds a request per flagged DOI).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inherited gaps:&lt;/strong&gt; the authoritative source (Crossref &lt;code&gt;updated-by&lt;/code&gt;) itself misses 7.5% of the gold-set retractions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero-hit passes are upper bounds,&lt;/strong&gt; not absences (0/928 across both live passes).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pipeline, not product:&lt;/strong&gt; per-list latency and PubMed-coverage bias (non-indexed papers absent) apply.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Run the audit on a fixed weekly batch (new review lists per PubMed date) → dated series, first point = this artifact.&lt;/li&gt;
&lt;li&gt;Extend seed sourcing to OpenAlex's own retracted, high-citation pool (done — it's what produced the passing seed).&lt;/li&gt;
&lt;li&gt;Optional: a public "retracted-citation finder" form (paste a DOI, get its list audited).&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;This work was done by Pennyforge, a one-person studio: the analysis was AI-assisted with human editorial review, as we practice it across all our research posts. All APIs public; per-DOI results available on request.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>api</category>
      <category>python</category>
      <category>research</category>
    </item>
  </channel>
</rss>
