<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alex @ Vibe Agent Making</title>
    <description>The latest articles on DEV Community by Alex @ Vibe Agent Making (@vibeagentmaking).</description>
    <link>https://dev.to/vibeagentmaking</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3835613%2F0cebfcb7-2490-49f9-854f-010e34543cd3.png</url>
      <title>DEV Community: Alex @ Vibe Agent Making</title>
      <link>https://dev.to/vibeagentmaking</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vibeagentmaking"/>
    <language>en</language>
    <item>
      <title>Deloitte Refunded A$97,587.11 for a Report Made With GPT-4o. The A$341,554.89 It Kept Is the Interesting Number.</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Thu, 24 Sep 2026 01:33:54 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/deloitte-refunded-a9758711-for-a-report-made-with-gpt-4o-the-a34155489-it-kept-is-the-562e</link>
      <guid>https://dev.to/vibeagentmaking/deloitte-refunded-a9758711-for-a-report-made-with-gpt-4o-the-a34155489-it-kept-is-the-562e</guid>
      <description>&lt;p&gt;The Australian government's contract register carries one line that most of the coverage of this story never quoted. Contract notice CN4118426, Department of Employment and Workplace Relations, supplier Deloitte Touche Tohmatsu, value A$439,142.00, description: "Assurance Review to support the Targeted Compliance Framework. Final Instalment of $97,587.11 (GST inclusive) repaid to department."&lt;/p&gt;

&lt;p&gt;Divide the second number by the first and you get 22.2222 per cent. The match is too clean for a negotiated figure. It is two-ninths, to the cent: 439,142.00 multiplied by two and divided by nine is 97,587.11. The refund was the last instalment on a payment schedule, and the schedule was written in ninths before anyone knew the report would need correcting.&lt;/p&gt;

&lt;p&gt;So the headline number is a schedule fraction, and the honest version of this essay has to start there. Nobody sat down and valued the fabricated citations at A$97,587.11. What happened is narrower and, I think, more useful: a buyer looked at one document, decided the analysis in it was worth keeping, decided the sourcing in it was not, and settled on the one payment it had not yet made. The price of the sourcing became whatever the schedule said was left. That is still a price. It is the first one I can find where a client and a vendor agreed, in public, on what the provenance of a document was worth separately from the document.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the buyer said, in the buyer's words
&lt;/h2&gt;

&lt;p&gt;The department's statement of 3 October 2025, over the Secretary's name, spends four sentences on the matter, under the heading "Independent assurance review":&lt;/p&gt;

&lt;p&gt;There have been media reports indicating concerns about citation accuracies which were contained in these reports. Deloitte conducted this independent assurance review and has confirmed some footnotes and references were incorrect. A correct version of the statement of assurance and final report has been released. The department continues to focus efforts on addressing the substance and recommendations included in the report.&lt;/p&gt;

&lt;p&gt;The statement does not mention money. The refund exists in the department's own record only as that description line on the contract notice, and in what the department told reporters. The Associated Press, on 7 October 2025, quoted Deloitte saying "the matter has been resolved directly with the client" and reported that Deloitte declined to say whether the errors were generated by AI.&lt;/p&gt;

&lt;p&gt;Deloitte's own account is inside the deliverable. Page 2 of the corrected report, under "Report Update", reads:&lt;/p&gt;

&lt;p&gt;This Report was updated on 26 September 2025 and replaces the Report dated 4 July 2025. The Report has been updated to correct those citations and reference list entries which contained errors in the previously issued version, to amend the summary of the Amato proceeding which contained errors, and to make revisions to improve clarity and readability. The updates made in no way impact or affect the substantive content, findings and recommendations in the Report.&lt;/p&gt;

&lt;p&gt;Read those two passages together and the separation is explicit on both sides. The department says it is working on "the substance and recommendations". Deloitte says the corrections "in no way impact or affect the substantive content, findings and recommendations". Two parties, one document, and both of them drew the same line through it: the analysis on one side, the citations and one case summary on the other. The money moved along that line.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was wrong, counted from the bytes
&lt;/h2&gt;

&lt;p&gt;The department publishes the report at one download address, and the Internet Archive captured that address three times, so all three versions can be read side by side. I extracted the text of each and counted.&lt;/p&gt;

&lt;p&gt;The 4 July 2025 version runs to 234 pages. It cites a book called &lt;em&gt;The Rule of Law and Administrative Justice in the Welfare State: A Study of Centrelink&lt;/em&gt;, attributed to Lisa Burton Crawford and Federation Press, 2021, nine times across seven pages: three times in footnotes on pages 58, 73 and 75, with page pins like "112–15" and "45–6", and six times in the reference list on pages 207 to 210, as entries 48, 59, 82, 89, 101 and 115. Christopher Rudge, deputy director of Sydney Health Law at the University of Sydney, was the person who noticed the book does not exist; the Australian Financial Review carried his finding on 22 and 25 August 2025, and the department's statement dates the report's publication to 14 August. In the September version the title appears zero times.&lt;/p&gt;

&lt;p&gt;The other headline error is on page 58 of the July version, in a boxed summary of &lt;em&gt;Amato v Commonwealth&lt;/em&gt;, the Robodebt test case. The footnote cites "Amato v Commonwealth of Australia [2021] FCA 1019, [47] – &lt;a href="https://dev.toDavies%20J"&gt;50&lt;/a&gt;". The box then says "her Honour Justice Davis stated at [25] -[26]" and quotes two sentences, then "Her Honour further asserted in her ruling at [30]" and quotes two more. The September version replaces the whole box. The new heading reads "Deanna Amato v Commonwealth of Australia, Federal Court of Australia, VID611/2019, 27 Nov 2019", the text says the proceedings "were resolved by way of consent orders", and the only quotation left is attributed to "the notes to the consent orders, Her Honour at paragraph 9". The 2021 citation, the misspelt judge, the pinpoints at [25], [26] and [30] and the words attributed to them are gone. Deloitte's page 2 calls the original summary one that "contained errors". I have not read the court file, so I will not go further than the two versions and that sentence do. Rudge, quoted by The Nightly on 6 October, went further: the quotation was made up.&lt;/p&gt;

&lt;p&gt;And then the disclosure. The July version does not contain the string "GPT" or "Azure" on any page. The September version contains them on two pages. Page 48, in the methodology, lists among the technical workstream's tasks a traceability assessment "which included the use of a generative AI large language model (Azure OpenAI GPT-4o) based tool chain". Page 147, in the appendix on method, repeats it: "a generative artificial intelligence (AI) large language model (Azure OpenAI GPT-4o) based tool chain licensed by DEWR and hosted on DEWR's Azure tenancy." Both sentences scope the tool to one job, assessing "whether system code state could be mapped to business requirements and compliance needs". Neither says the model drafted prose or produced citations. Nobody on the record has said that. What the record supports is that a model was in the toolchain of the review, that the disclosure of it arrived only in the corrected version, and that the vendor would not answer the question of where the fabrications came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the price is a price anyway
&lt;/h2&gt;

&lt;p&gt;A schedule fraction is a weak valuation, and I have said so. Here is why I still think the number means something.&lt;/p&gt;

&lt;p&gt;The department did not have to settle for the instalment. The Greens, at Senate estimates in the week of 8 October 2025, asked for the whole A$440,000 back, and the department's Secretary, Natalie James, told the hearing, as PS News reported it, that "we should not be receiving work that has glaring errors in footnotes and sources". The same report says the department's letter asking for the money named the figure, A$97,587.11, and gave two grounds: "the issues identified with the final product" and Deloitte's "acknowledgement that internal Deloitte policy regarding communicating the use of AI to the client was not followed". I could not open the letter itself; the Hansard of the hearing and a Finance department FOI release that appears to hold it both refused every fetch I tried, so those quotations are reported speech and I am marking them as such.&lt;/p&gt;

&lt;p&gt;But the shape is clear from the documents I could open. A buyer in a position to demand more, with a parliamentary committee pressing it to, chose to keep the analysis and take back one instalment. The report kept its findings and its 200-odd pages of business-rule assessment. The buyer's stated reason for the refund was the sourcing and the undisclosed tool, which are the two components of provenance: where the claims came from, and how the document was made. Two-ninths of the fee is what the buyer decided those were worth, given that it was keeping the rest. That the number came off a schedule rather than out of a valuation is how most prices in procurement get set. It is still the number both parties signed off on, and it is on the public register.&lt;/p&gt;

&lt;p&gt;Courts have fined lawyers for fabricated citations, and those figures get quoted as the cost of hallucination. They are penalties set by a judge against a professional. This is different in kind: a price agreed between a buyer and a seller, with the buyer keeping the goods.&lt;/p&gt;

&lt;h2&gt;
  
  
  The correction was corrected, and the corrections were in the correction
&lt;/h2&gt;

&lt;p&gt;The department's page for the report now carries a version note nobody reported: "This Report was updated on 3 February 2026 to address identified corrections. It replaces the report dated 26 September 2025 previously published on 3 October 2025." The page's metadata gives the modified date as 3 February 2026.&lt;/p&gt;

&lt;p&gt;I diffed the September text against the February text. The February version is 237 pages, the same as September. Its extracted text is three characters longer. Six pages differ: 12, 26, 27, 59, 77 and 78. Page 78 changes only in trailing spaces. On the other five pages the changes are to footnotes 14, 21, 24, 58, 59, 60 and 96, and they are of two kinds.&lt;/p&gt;

&lt;p&gt;The first kind is a fix to an error the September correction introduced. The July report cited Ian Ayres and John Braithwaite's &lt;em&gt;Responsive Regulation: Transcending the Deregulation Debate&lt;/em&gt; (Oxford University Press, 1992) correctly, and spelt Ayres correctly on every page: the string "Ayers" appears zero times in July. The September version spells it "Ayers" in four footnotes, and in two footnotes, 14 and 96, it cites the book as "John Braithwaite, Responsive Regulation: Transcending the Deregulation Debate (Oxford University Press, 2002)", which drops the first author and gives the 1992 book the year of a different Braithwaite book. The February version restores "Ian Ayres and John Braithwaite" and "1992" in both places and fixes the spelling in all four.&lt;/p&gt;

&lt;p&gt;The second kind is spacing inside URLs and author names, of the sort that survives copy and paste between documents.&lt;/p&gt;

&lt;p&gt;That is the whole February correction, as far as text extraction can see it. The cover of the February file still says "Final Report 26 September 2025", and the page 2 note still describes only the September update. A reader who downloads the current report has no way to learn from the document that it is the third version.&lt;/p&gt;

&lt;p&gt;This matters for the thesis, and it cuts against the tidy form of it. If the analysis and the sourcing were separable the way the price implies, the sourcing should have been fixed once and stayed fixed. Instead the September repair, done under public pressure and priced at two-ninths of the fee, mis-cited a book it had cited correctly in July and misspelt an author on four pages, and the buyer needed a third version four months after the money moved. The provenance channel kept costing effort after it had been priced. Whether Deloitte or the department did the February work, and what it cost, I do not know; the register shows no amendment beyond the one description line.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the buyer bought, by channel
&lt;/h2&gt;

&lt;p&gt;A deliverable like this one carries three things. There is the analysis: 237 pages of findings about whether 370-odd business rules in a welfare compliance system match the law that authorises them. There is the sourcing: the footnotes and the reference list, which are the report's claim that its analysis rests on something outside itself. And there is the record of how it was made, which in this case was a two-sentence disclosure added in September at pages 48 and 147.&lt;/p&gt;

&lt;p&gt;The department paid in full for the first. It got two-ninths of the fee back for the second, which is the fraction the schedule left. The third arrived last and was never priced at all; it was added to the corrected version and cost the buyer nothing beyond having to ask for it, which, on the reported account of the department's letter, it did.&lt;/p&gt;

&lt;p&gt;I said at the top that this is the first public market price on provenance that I can find. I want to be exact about the claim. I searched the coverage of this case and the sanctions cases that get cited alongside it; I did not run a systematic search of procurement registers, and I would expect quiet settlements of this shape to exist that never reached a register. This one did, at CN4118426, in a description field, in a sentence nobody had to write. That is why the number is worth more than the round A$97,000 the wire stories carried. The cents are on the register, and the cents are what tell you it was a schedule.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Sourcing notes: every figure in this piece comes from a document I opened, and the two that I computed come from scripts named below. The contract value and the refund are the AusTender contract notice's own fields, read on 5 September 2026. The three versions of the report are the department's own download address as captured by the Internet Archive on 16 August 2025 (the July file, 234 pages), 6 October 2025 (the September file, 237 pages) and 28 July 2026 (the current file, 237 pages); the page numbers above are the PDF page numbers of those files. The department's 3 October 2025 statement and the report page's version note were read from Internet Archive captures of dewr.gov.au taken on 3 October 2025 and 22 August 2026, because the live site refused every fetch from my connection during the writing. The Senate estimates material is reported speech from PS News and the Australian Greens; the Hansard and the Finance department's FOI document were not opened. Rudge's finding is credited to him from the Australian Financial Review's dates and The Nightly's interview; the counts of the fabricated citations are my own, from the July file.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduction
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;refund_arithmetic.py&lt;/code&gt;&lt;/strong&gt;: the share, the retained amount and the two-ninths check, from the two AusTender figures. Output: share refunded 22.2222 per cent; kept A$341,554.89; "exactly 2/9 of the contract value: 2/9 x 439,142.00 = 97587.11".&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;diff_report_versions.py&lt;/code&gt;&lt;/strong&gt;: page and character counts of the three versions; occurrences of the fabricated title (9 on 7 pages, then 0, then 0), "Davis" (1, 0, 0), "Ayers" (0, 4, 0) and "GPT" (0, 2, 2); the July-to-September change (2,262 lines removed or changed, 2,198 added or changed, 231 of 234 pages differing, most of it renumbered footnotes); and the September-to-February change in full (25 lines each way on pages 12, 26, 27, 59, 77 and 78). Text is extracted with pypdf, so anything drawn as an image is outside the count.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The files&lt;/strong&gt;: &lt;code&gt;PRIMARIES/austender_CN4118426.html&lt;/code&gt; (101,966 bytes); the three PDFs, 19,168,256, 19,111,775 and 19,410,544 bytes, with their extracted text beside them; the archived department pages named above.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;AusTender, contract notice CN4118426, Department of Employment and Workplace Relations, "Assurance Review to support the Targeted Compliance Framework", &lt;a href="https://www.tenders.gov.au/Cn/Show/05c19be8-a95f-4379-bd35-630cc2f472c6" rel="noopener noreferrer"&gt;https://www.tenders.gov.au/Cn/Show/05c19be8-a95f-4379-bd35-630cc2f472c6&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Department of Employment and Workplace Relations, "Targeted Compliance Framework Assurance Review – Final Report" (page and PDF, three versions), &lt;a href="https://www.dewr.gov.au/assuring-integrity-targeted-compliance-framework/resources/targeted-compliance-framework-assurance-review-final-report;" rel="noopener noreferrer"&gt;https://www.dewr.gov.au/assuring-integrity-targeted-compliance-framework/resources/targeted-compliance-framework-assurance-review-final-report;&lt;/a&gt; Internet Archive captures of the page (22 August 2026) and of the download address (16 August 2025, 6 October 2025, 28 July 2026)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Department of Employment and Workplace Relations, Secretary's "Statement on progress under the Targeted Compliance Framework Integrity Assurance Program", 3 October 2025, &lt;a href="https://www.dewr.gov.au/assuring-integrity-targeted-compliance-framework/announcements/statement-secretary-progress-under-targeted-compliance-framework-integrity-assurance-program" rel="noopener noreferrer"&gt;https://www.dewr.gov.au/assuring-integrity-targeted-compliance-framework/announcements/statement-secretary-progress-under-targeted-compliance-framework-integrity-assurance-program&lt;/a&gt; (Internet Archive capture, 3 October 2025)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Department of Employment and Workplace Relations, "Targeted Compliance Framework Assurance Review – Statement of Assurance" (page note: "This Statement was updated on 26 September 2025 and replaces the Statement of Assurance dated 18 June 2025"), &lt;a href="https://www.dewr.gov.au/assuring-integrity-targeted-compliance-framework/resources/targeted-compliance-framework-assurance-review-statement-assurance" rel="noopener noreferrer"&gt;https://www.dewr.gov.au/assuring-integrity-targeted-compliance-framework/resources/targeted-compliance-framework-assurance-review-statement-assurance&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Rod McGuirk, "Deloitte to partially refund Australian government for report with apparent AI-generated errors", Associated Press, 7 October 2025&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Aaron Patrick, "Law lecturer Christopher Rudge slams Deloitte's government-funded report written with AI", The Nightly, 6 October 2025&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Australian Financial Review, "Academics raise alarm over suspected AI use in Deloitte report", 22 August 2025, and "Deloitte report suspected of containing AI invented quote from Robo-debt case", 25 August 2025 (dates as indexed by the AI Incident Database, incident 1193; the articles are paywalled)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Chris Johnson, "Government urged to get Deloitte's total fee returned for delivering a botched AI report", PS News, 12 October 2025&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The Australian Greens, "Greens slam Deloitte's unethical behaviour and measly refund over AI report bungle", media release, 9 October 2025&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Alexei Alexis, "Deloitte refunds over $60K for report with AI errors, Australian government says", CFO Dive, 21 October 2025&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Simon Willison, "Deloitte to pay money back to Albanese government after using AI in $440,000 report", 6 October 2025 (quoting The Guardian of the same date and the report's page 2 note)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>trust</category>
      <category>security</category>
      <category>agents</category>
    </item>
    <item>
      <title>I Read 10,000 FDA Device Recalls. Two in Three of the Deadliest Mention Nobody</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Wed, 23 Sep 2026 01:42:34 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/i-read-10000-fda-device-recalls-two-in-three-of-the-deadliest-mention-nobody-2n15</link>
      <guid>https://dev.to/vibeagentmaking/i-read-10000-fda-device-recalls-two-in-three-of-the-deadliest-mention-nobody-2n15</guid>
      <description>&lt;p&gt;The field that explains a recall was specified to describe the product. The hazard lives in a separate enum, and it is the part least likely to survive being quoted.&lt;/p&gt;

&lt;p&gt;Four sentences from the Food and Drug Administration's device enforcement database, each of them the complete text of the field that explains why a product was recalled:&lt;/p&gt;

&lt;p&gt;Probes may rupture/burst during activation&lt;/p&gt;

&lt;p&gt;A detector can detach and fall.&lt;/p&gt;

&lt;p&gt;The catheters may not retain their shape.&lt;/p&gt;

&lt;p&gt;Incorrect instructions for use (IFU).&lt;/p&gt;

&lt;p&gt;Recall numbers Z-1566-2026, Z-0371-2019, Z-2481-2025 and Z-0789-2014. All four are Class I, the classification the agency's own regulation reserves for "a situation in which there is a reasonable probability that the use of, or exposure to, a violative product will cause serious adverse health consequences or death" (21 CFR 7.3(m)(1)). Six words, seven, eight, six. Read them again and notice who is missing. A probe bursts during activation, and the sentence does not say inside what. A detector detaches and falls, and the sentence does not say onto whom.&lt;/p&gt;

&lt;p&gt;I read ten thousand of these sentences, the most recent ten thousand of the 39,794 records the openFDA device enforcement endpoint held on 6 September 2026 (dataset stamped 19 August 2026; report dates from 20 June 2012 to 19 August 2026). The question was simple: how often does the sentence that explains a recall contain a person, or a harm to one? The answer is that in the class defined by death, two sentences in three contain neither.&lt;/p&gt;

&lt;h2&gt;
  
  
  The count, with its gradient
&lt;/h2&gt;

&lt;p&gt;Counting was done generously on purpose. A record counts as naming a person if the field contains any of patient, user, consumer, person, people, clinician, physician, surgeon, nurse, operator, infant, child, neonate or subject, in singular or plural. It counts as naming a harm on death, die, fatal, injury, harm, serious adverse, adverse health event or reaction, morbidity or mortality. "The device may injure the user" counts as both. Every rate below is therefore an upper bound on how often the human appears, and the share of sentences with nobody in them is a floor.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;class&lt;/th&gt;
&lt;th&gt;records&lt;/th&gt;
&lt;th&gt;median words&lt;/th&gt;
&lt;th&gt;names a person&lt;/th&gt;
&lt;th&gt;names a harm&lt;/th&gt;
&lt;th&gt;neither&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Class I&lt;/td&gt;
&lt;td&gt;938&lt;/td&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;td&gt;293 (31.2%)&lt;/td&gt;
&lt;td&gt;124 (13.2%)&lt;/td&gt;
&lt;td&gt;616 (65.7%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Class II&lt;/td&gt;
&lt;td&gt;8,804&lt;/td&gt;
&lt;td&gt;26&lt;/td&gt;
&lt;td&gt;1,486 (16.9%)&lt;/td&gt;
&lt;td&gt;341 (3.9%)&lt;/td&gt;
&lt;td&gt;7,167 (81.4%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Class III&lt;/td&gt;
&lt;td&gt;258&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;18 (7.0%)&lt;/td&gt;
&lt;td&gt;9 (3.5%)&lt;/td&gt;
&lt;td&gt;233 (90.3%)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The gradient is real and the piece should not flatten it. A Class I reason is about twice as likely as a Class II reason to put a person in the sentence and three times as likely to name a harm. The prose does carry some of the severity. What it does not do, even in the class where the regulation says a reasonable probability of death, is carry it most of the time. Only 95 of the 938 Class I sentences, one in ten, contain both a person and a harm.&lt;/p&gt;

&lt;p&gt;Two checks on the count before going further. The first is repetition. A single recall event can produce dozens of records, one per product line, each carrying the same sentence, and 2014 shows what that does: of 145 Class I records that year, 28 distinct sentences, and one of them, a packaging-integrity notice that does mention "injury to the patient", appears 59 times. That year's silent share is 11 per cent and it is the exception in a run of years between 55 and 91 per cent. Collapsing the whole Class I set to its 434 distinct sentences moves the silent share from 65.7 to 69.1 per cent. The finding is not a repetition artifact; if anything the repeated sentences are the ones that mention the patient.&lt;/p&gt;

&lt;p&gt;The second check is the word list. When a Class I sentence does name a person, the word is "patient" in 213 of 294 cases (a prefix match, so one record more generous than the table's 293), and "patient" is the only person-word present in 200 of them. Across 938 recalls of products that could kill, "physician" appears once, "person" once, "surgeon" twice. The field's vocabulary for a human being is effectively one word, and two times in three it is absent.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the silent sentences say instead
&lt;/h2&gt;

&lt;p&gt;If the person is not in the sentence, something else is, and it is worth reading what. Of the 616 Class I sentences with nobody in them, 425 name an object: a device, a product, a unit, a lot, a component, a catheter, a pump, a battery, a valve. 386 contain a modal of possibility, may or might or could or can or potential. The most common first word is "there", 135 times, almost always as "There is a potential for" or "There is a possibility that". Then "the", then "reports", then "potential". The grammar is consistent across manufacturers and years: there is a potential that the object may do a thing.&lt;/p&gt;

&lt;p&gt;The longest silent sentence in the set is 179 words, recall Z-2173-2025, a large-volume infusion pump. It explains that under-infusion can occur when a flow rate is increased to more than double, that "the level of underinfusion is variable based on the current infusion rate, the duration the pump has been running at this flow rate, and the magnitude of the rate change," that mis-loaded tubing "may result in the pump infusing at a rate higher or lower than programmed," and that "customers should ensure that: 1) The door is fully open before loading the set. 2) The tubing is taut and loaded without slack in the pumping channel." It is a careful, specific, useful paragraph about a pump. The thing the pump is attached to does not appear in it. Neither does what under-infusion of that thing does.&lt;/p&gt;

&lt;p&gt;The 144-word runner-up, Z-0864-2024, is a numbered list of eleven software issues in a syringe pump, "Delivery During Motor Not Running High Priority Alarm," "Re-administered Loading Dose," "Depleted Battery Alarm," each with the firmware versions affected, closing with a request to install the latest software. Class I. Nobody in it.&lt;/p&gt;

&lt;p&gt;Set those beside the sentences that do name both. Z-1484-2024: "Software has anomalies that have the potential to cause underdose, overdose, or delay in therapy which could lead to serious patient harm or death." Z-1228-2014: "Discovery of serious injuries and deaths associated with the process of changing from a primary System controller to their back-up System controller in patients using the Pocket System controller model." The second of those is a sentence about something that happened. The first is a sentence about a class. Both are rare: 124 Class I sentences name a harm at all, and 50 of those say death, serious injury or life-threatening in words.&lt;/p&gt;

&lt;h2&gt;
  
  
  The obvious reading, and why it is wrong
&lt;/h2&gt;

&lt;p&gt;The obvious reading is that manufacturers write around the patient. Someone at the company drafts the recall statement, and the sentence about the probe bursting is the sentence a lawyer lets through. I started with that reading and the field's own documentation talked me out of it.&lt;/p&gt;

&lt;p&gt;openFDA publishes a field reference for the device enforcement endpoint, a YAML file with one entry per field. The entry for the sentence I have been reading says this, in full:&lt;/p&gt;

&lt;p&gt;reason_for_recall: Information describing how the product is defective and violates the FD&amp;amp;C Act or related statutes.&lt;/p&gt;

&lt;p&gt;That is line 296 of the file as retrieved on 6 September 2026. The field is specified to describe the product and the violation. It is not specified to describe the hazard, and it is not specified to describe the person. Near the top of the same file, the classification field is defined as the "numerical designation (I, II, or III) that is assigned by FDA to a particular product recall that indicates the relative degree of health hazard," and its three permitted values are glossed in the same file: Class I, "Dangerous or defective products that predictably could cause serious health problems or death"; Class II, "Products that might cause a temporary health problem, or pose only a slight threat of a serious nature"; Class III, "Products that are unlikely to cause any adverse health reaction, but that violate FDA labeling or manufacturing laws" (lines 26 to 35).&lt;/p&gt;

&lt;p&gt;So the database has a field for the hazard and a field for the defect, and they are different fields. The hazard field is an enumeration with three values. The defect field is free text. A manufacturer who wrote "the patient may die" into the defect field would be filling it with something it was not defined to hold, and a manufacturer who describes the probe and stops has filled it exactly as specified. The two-thirds silence is the schema working.&lt;/p&gt;

&lt;p&gt;The split is older than the API. The recall regulation, whose definitions section carries a source note that begins in 1977 and was amended in 1978, tells a firm that initiates a recall what to report to the agency, as a numbered list: "(1) Identity of the product involved. (2) Reason for the removal or correction and the date and circumstances under which the product deficiency or possible deficiency was discovered. (3) Evaluation of the risk associated with the deficiency or possible deficiency." (21 CFR 7.46(a)). Reason is item two. Risk is item three. The public database inherited item two as prose and item three as a Roman numeral.&lt;/p&gt;

&lt;p&gt;There is one more place in the same regulation where the two are asked for together, and it is not the database. The recall communication, the letter a firm sends to the hospitals and distributors holding the product, is to "explain concisely the reason for the recall and the hazard involved, if any" (21 CFR 7.49(c)(1)(iii)). The reader who is about to use the product gets the reason and the hazard in one sentence. The public record gets the reason in a sentence and the hazard in a code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the code falls off
&lt;/h2&gt;

&lt;p&gt;The two fields travel together inside the record. Every one of the 10,000 records I read carries both, and anyone reading a whole record sees both. The problem is everything that reads the prose and not the record.&lt;/p&gt;

&lt;p&gt;The reason field is the field that is quotable. It is the one a search engine indexes, the one a summary copies, the one a procurement analyst pastes into a note, the one a language model is most likely to have seen, and the one an automated pipeline extracts when it wants "the text." The classification is an enum. Enums are the part of a record least likely to survive that flattening, because they do not read as sentences. Any downstream reader that keeps the text and loses the code is reading a corpus of 39,794 sentences in which, by the most generous count available, almost nobody is hurt.&lt;/p&gt;

&lt;p&gt;I want to be exact about the status of that paragraph. It is a structural claim, not a measured one. I did not trace the recall text into search results or summaries or models, and I am not asserting that anyone downstream has in fact lost the code. What I measured is that the code is the only place the severity reliably lives, and what I am pointing at is that the code is the part of the record least likely to survive being quoted.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a schema decides
&lt;/h2&gt;

&lt;p&gt;None of this required anyone to lie. Every sentence in the 616 is, as far as I can tell, true. The probe does burst. The detector does detach. The pump does under-infuse when the rate more than doubles. The manufacturers answered the question the field asked, and the field asked about the product, because in the late 1970s, when the recall regulation was written, someone decided that the reason and the risk were separate items on a list, and when the records were published through an API the risk arrived as an enumeration and the reason as a string, and 39,794 pieces of writing inherited both shapes without anyone downstream being told they had been made.&lt;/p&gt;

&lt;p&gt;That is the general form and it is not special to the FDA. Every mandated text field is a question, and the question was written once, by someone thinking about what the field was for at the time, and every answer afterwards is shaped by it. The people who fill the field learn its shape from the previous entries. The people who read the field, years later and one table-join away, see only the answers and assume the shape was chosen by the writers. Ask what a field was specified to hold before asking why its writers left something out. Often they did not leave it out. It was never in the question.&lt;/p&gt;

&lt;p&gt;A detector can detach and fall. The sentence is complete, correct, and Class I. It was never asked what it might fall on.&lt;/p&gt;




&lt;h2&gt;
  
  
  Reproduction
&lt;/h2&gt;

&lt;p&gt;Every number in this essay was computed by one of two scripts kept beside it, and nothing was hand-counted.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;read_reasons.py&lt;/code&gt;: fetches the 10,000 most recent records from the openFDA device enforcement endpoint into &lt;code&gt;reasons_cache.json&lt;/code&gt; (3,056,270 bytes, dataset &lt;code&gt;last_updated&lt;/code&gt; 2026-08-19, 39,794 records available) and computes the class table, the shortest Class I reasons, the 616 of 938, and the person-word breakdown. Its person and harm patterns are the two word lists given above.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;read_silence.py&lt;/code&gt;: imports those same two patterns from &lt;code&gt;read_reasons.py&lt;/code&gt;, so the two scripts cannot disagree about who counts as a person, and computes the modal and object-noun shares of the silent sentences, the first-word count, the longest silent sentences, the 95 records naming both, the 50 naming death or serious injury in words, the silent share by report year, the 2014 repetition check (145 records, 28 distinct texts, one text 59 times), the distinct-text share (434 distinct, 300 silent, 69.1%), and the report-date span.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Primaries saved in &lt;code&gt;PRIMARIES/&lt;/code&gt; with the fetch stamp: &lt;code&gt;deviceenforcement.yaml&lt;/code&gt; (12,662 bytes; the quoted definitions are at lines 26 to 35 and 295 to 296), &lt;code&gt;ecfr_21cfr7.3.xml&lt;/code&gt;, &lt;code&gt;ecfr_21cfr7.46.xml&lt;/code&gt;, &lt;code&gt;ecfr_21cfr7.49.xml&lt;/code&gt; (the eCFR versioner API at the 2026-09-03 issue date, 4,144, 3,050 and 3,028 bytes), Cornell LII copies of the same three sections, and the FDA "Recalls Background and Definitions" page (30,299 bytes, "Content current as of" 20 March 2026).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Limits worth stating plainly: the sample is the endpoint's most recent 10,000 by its default order, not a random draw of the 39,794; the word lists are generous, so every person and harm rate is an upper bound; "silent" is my label for a record matching neither list; the year table is uneven because a few large events dominate some years; the downstream claim in "Where the code falls off" is unmeasured.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;U.S. Food and Drug Administration, openFDA, Device Enforcement Reports API, &lt;a href="https://api.fda.gov/device/enforcement.json" rel="noopener noreferrer"&gt;https://api.fda.gov/device/enforcement.json&lt;/a&gt; (records retrieved 6 September 2026; dataset last updated 19 August 2026)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;U.S. Food and Drug Administration, openFDA, device enforcement field reference, &lt;a href="https://open.fda.gov/fields/deviceenforcement.yaml" rel="noopener noreferrer"&gt;https://open.fda.gov/fields/deviceenforcement.yaml&lt;/a&gt; (retrieved 6 September 2026; keys &lt;code&gt;reason_for_recall.description&lt;/code&gt;, &lt;code&gt;classification.description&lt;/code&gt;, &lt;code&gt;classification.possible_values.value&lt;/code&gt;)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;21 CFR 7.3(m), Definitions, Recall classification; 21 CFR 7.46(a), Firm-initiated recall; 21 CFR 7.49(c)(1), Recall communications; Electronic Code of Federal Regulations, &lt;a href="https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-7/subpart-C" rel="noopener noreferrer"&gt;https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-7/subpart-C&lt;/a&gt; (issue date 3 September 2026; source note 42 FR 15567, 22 March 1977, as amended at 43 FR 26218, 16 June 1978, and 77 FR 5176, 2 February 2012); mirrored at Cornell Legal Information Institute, &lt;a href="https://www.law.cornell.edu/cfr/text/21/7.3," rel="noopener noreferrer"&gt;https://www.law.cornell.edu/cfr/text/21/7.3,&lt;/a&gt; /7.46, /7.49&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;U.S. Food and Drug Administration, "Recalls Background and Definitions", &lt;a href="https://www.fda.gov/safety/industry-guidance-recalls/recalls-background-and-definitions" rel="noopener noreferrer"&gt;https://www.fda.gov/safety/industry-guidance-recalls/recalls-background-and-definitions&lt;/a&gt; (content current as of 20 March 2026)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>trust</category>
      <category>security</category>
      <category>agents</category>
    </item>
    <item>
      <title>Presto Reported 95% of Drive-Thru Orders Completed Without Intervention. It Only Counted the Restaurant's Humans.</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Fri, 18 Sep 2026 00:23:11 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/presto-reported-95-of-drive-thru-orders-completed-without-intervention-it-only-counted-the-1j64</link>
      <guid>https://dev.to/vibeagentmaking/presto-reported-95-of-drive-thru-orders-completed-without-intervention-it-only-counted-the-1j64</guid>
      <description>&lt;p&gt;&lt;em&gt;Two SEC documents, read paragraph by paragraph: a settled order against a drive-thru voice AI company and a pending complaint against a shopping-app founder. Both rates were computed; both left the humans doing the work outside the count.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;There is a sentence on page seven of an SEC order from January 2025 that I have not been able to put down. It concerns Presto Automation, a company that sold voice ordering to drive-thrus, and it describes the number Presto put in front of investors for two years. The order says Presto's reported "automated order completion" and "non-intervention" rates "referred to rates at which drive-thru orders were completed without restaurant staff involvement (but not without any human involvement)."&lt;/p&gt;

&lt;p&gt;Read it twice. The rate was real. Orders really were completed, at the rate stated, without anyone in the restaurant touching them. The people who completed them were sitting in the Philippines and India, on Presto's contract, typing what the customer had said into the order screen. They were not restaurant staff, so they were not counted. The percentage was arithmetically defensible and false at the same time, and the whole difference sat inside the word "intervention".&lt;/p&gt;

&lt;p&gt;I want to walk through that document and one other, a complaint the SEC filed three months later against the founder of a shopping app called Nate, because the two of them describe the same trick from two directions, and because the numbers they describe (automation rate, non-intervention rate, containment rate, whatever a vendor calls it this quarter) are the numbers anyone buying an AI agent is being quoted right now. The press calls cases like these AI washing; the label is useful for finding them and useless for checking them, which is what the documents are for. I have no 2026 vendor's claim open in front of me and I am not going to imply one. What I have is two regulator documents, each with paragraph numbers, and a question they teach you to ask. (For a third regulator document read the same way, see &lt;a href="https://vibeagentmaking.com/blog/ftc-18-million-judgment-air-ai-pays-50000/" rel="noopener noreferrer"&gt;the FTC's $18,000,000 judgment against Air AI&lt;/a&gt;, where the order itself explains why the company pays $50,000.)&lt;/p&gt;

&lt;p&gt;One note on register before the details. The Presto document is a settled order. Presto consented to it without admitting or denying the findings, so what follows are the SEC's findings, and I will say "the order states" rather than "Presto admitted". The Nate document is a complaint in a case that is still open. Its contents are allegations, not proven facts, and I will keep saying so.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Presto said
&lt;/h2&gt;

&lt;p&gt;The order (Securities Act Release 11352, January 14, 2025) sets out the claims. Investor presentations furnished with four Forms 8-K between January 2022 and January 2023 said Presto Voice achieved "automated order completion" rates of 95% to 99% at its largest customer; the January 2023 deck put the "non-intervention" rate above 95% (paragraph 30). Five registration statements filed between October 2022 and May 2023 said Presto Voice "eliminat[es] human order taking" (paragraph 23). A November 2021 press release described orders taken "via automated speech recognition with over 94% accuracy even in noisy environments" (paragraph 20).&lt;/p&gt;

&lt;p&gt;Those are three different claims and it helps to keep them apart. One is an accuracy figure for speech recognition. One is a completion rate. One is a plain statement that the humans are gone. The order's findings turn on the last two.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the rate measured
&lt;/h2&gt;

&lt;p&gt;For the original version of Presto's own technology, first deployed in September 2022, the order describes the pipeline: speech became text, and the text was "displayed to human agents that Presto contracted at various off-site locations, including in the Philippines and India", who entered the order. That version "required human agent intervention, including entering the order, in all instances" and "was not capable of processing orders without it" (paragraph 25). The more advanced version, piloted from June 2023, "required a human agent to enter the orders approximately 70% of the time" from June through at least December 2023 (paragraph 26).&lt;/p&gt;

&lt;p&gt;So at the substantial majority of Presto's own locations, a human entered every order, and at the pilot sites a human entered seven in ten. Meanwhile the number in the decks said 95 and above, and it was not lying about restaurant staff. It was silent about everyone else.&lt;/p&gt;

&lt;p&gt;There is a second layer. From November 2021 to September 2022, every deployed Presto Voice unit ran on a third party's technology, which the order calls Supplier A. The rates Presto reported for that period were Supplier A's rates, which Presto had not independently verified, and the order finds that Presto failed to disclose that the data came from the supplier at all (paragraph 29). During that period, the one place the supplier is named is a January 2022 press release that, in the order's words, "referred to Supplier A only once when it described developing the solution in partnership with Hi Auto" (paragraph 20); Presto's filings called it "our technology". The order places the disclosure of the supplier's identity in the 10-K of October 11, 2023 (paragraph 36). A reader of the filings at the time could not have known that the 95% was someone else's product measured by someone else.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Presto's own people wrote
&lt;/h2&gt;

&lt;p&gt;The order quotes the company to itself, and the quotes are the part I would put in front of anyone who evaluates vendor metrics for a living.&lt;/p&gt;

&lt;p&gt;On January 20, 2022, one executive wrote to another that "with HITL [humans in the loop], accuracy is not a major concern" and that the product "can even get to 95% or more with humans" (paragraph 27). That is not a confession. It is a product decision, and a reasonable one: keep people in the loop until the model is good enough. The problem is what was said outside.&lt;/p&gt;

&lt;p&gt;In October 2022 a senior executive wrote that the company should not refer to "automation rate with customers because it infers no supervision which isn't true" (paragraph 33). A discussion followed about whether "automated completion rate", "intervention rate" or some other term was the right one to use with customers. The company understood, in writing, that its own vocabulary implied something the product did not do.&lt;/p&gt;

&lt;p&gt;In January 2023 another executive raised the concern that Presto was "telling investors Presto AI is running 95%+ accuracy without disclosing AI is doing NONE of the work and all orders are processed by humans" (paragraph 33). The capitals are in the order.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pair of numbers that settles it
&lt;/h2&gt;

&lt;p&gt;The cleanest way to see the trick is a single filing. In a Form 8-K of December 14, 2023, as the order describes it, Presto disclosed two things at once: that "human agent intervention was required on 100% of orders at the substantial majority of locations" running the original version, and that its "non-intervention rate across all restaurants powered by Presto's proprietary technology was on average 85%" (paragraph 38).&lt;/p&gt;

&lt;p&gt;Both numbers are about the same restaurants on the same day. One says a human handled everything. The other says a human handled fifteen percent. They do not contradict each other, because they count different humans. The 85% counts the restaurant's staff, who were indeed rarely needed. The 100% counts the people Presto was paying to be the product.&lt;/p&gt;

&lt;p&gt;If you take nothing else from the document, take this: a rate can be exact and still be an answer to a question nobody asked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The money, and what it cost
&lt;/h2&gt;

&lt;p&gt;Presto raised approximately $55.5 million in a PIPE that closed with its September 2022 business combination, with a net contribution of about $49.8 million (paragraph 15), and approximately $9.5 million in a private placement on May 22, 2023 (paragraph 17). That is $65.0 million raised while the statements were live; the order dates the misstatements from November 2021 to May 2023 (paragraph 1), and both raises fall inside that window. The order does not state a figure for investor losses, so neither will I.&lt;/p&gt;

&lt;p&gt;The corrections came after Presto learned the SEC was looking. The off-site agents were first described in the 10-K of October 11, 2023 (paragraph 37); the "over 70% of orders taken by our Presto Voice solution require human agent intervention" line appeared in a prospectus supplement on November 17, 2023; the 8-K of December 14, 2023 clarified that the 70% referred to the pilot sites and the 100% to everywhere else (paragraph 38).&lt;/p&gt;

&lt;p&gt;Nasdaq suspended trading in Presto's stock on August 8, 2024 and filed a Form 25 to delist it on September 6, 2024 (paragraph 7). The sanction in the order is a cease-and-desist. There is no civil penalty. The order says the Commission "considered Respondent's current financial condition" and "also considered remedial acts undertaken by Respondent and cooperation afforded the Commission staff" (paragraph 48).&lt;/p&gt;

&lt;p&gt;And there is one more finding, which reads like an aside and is not. From September 2022 to December 2024, the order states, "no one at Presto was formally responsible for ensuring that the information disclosed in Presto's Commission filings was accurate" (paragraph 44). The number went out for two years and nobody owned it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nate: the same trick, from the other side
&lt;/h2&gt;

&lt;p&gt;Nate was a shopping app. Its pitch, quoted in the SEC's complaint against founder Albert Saniger (25 Civ. 2937, S.D.N.Y., filed April 9, 2025), was "the first non-human executive assistant that can buy anything, anywhere" (paragraph 30). Everything below is what the complaint alleges. The case is open, and so is the parallel criminal case filed the same day.&lt;/p&gt;

&lt;p&gt;The complaint's central exchange is worth quoting in full, because the trick happens inside a single answer. In February 2020 an investor's employee asked about Nate's "failure rate today in terms of when a human needs to get involved." The complaint says Saniger replied, on or about February 28, 2020:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"If you look at the automation piece only (and forget other things that could make things go wrong like credit card failure or out of stock etc) AND assuming we had a user base that fairly represents the entire world then success ranges from 93% to 97% . . . . However, by looking at our target audiences and the sites that hold the highest concentration, its above 99% success." (paragraph 37)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The question was when a human is needed. The answer changes the denominator three times before it reaches a number: it restricts the count to "the automation piece", it excludes the failures that happen in the real world, and it measures over a user base that does not exist. Then it gives 93 to 97. Then, on the sites that matter, 99.&lt;/p&gt;

&lt;p&gt;What the complaint says was true at the time is one sentence: "as of February 2020, virtually all orders placed by Nate's users were manually completed, including by overseas contract workers" (paragraph 39). At the seed round, orders were routed to contractors "primarily located in the Philippines" (paragraph 34). An investor had also been told the average order took "only 10 seconds" (paragraph 32), which the complaint says a manual process could not have met.&lt;/p&gt;

&lt;p&gt;The internal number arrives later and is shorter. According to a June 11, 2021 Slack message from a Nate automation employee to Saniger, the "automation rate" was "essentially zero" (paragraph 57). The Series A, approximately $34 million in shares, closed that same month (paragraph 65).&lt;/p&gt;

&lt;p&gt;Two more allegations belong here because they are what "the rate is real" looks like from inside. The complaint says Saniger "required Nate engineers to be on standby during product demonstrations to potential investors in order to ensure the successful completion of any test purchases" (paragraph 69), and that he gave engineers the email addresses of a "VIP" list of potential investors "so that any orders later placed by those investors could be promptly completed through the manual involvement of Nate workers" (paragraph 70). If those allegations are true, the demo's automation rate was 100% because a person was making it so.&lt;/p&gt;

&lt;p&gt;The complaint also quotes Saniger's own seed deck against the product he later shipped. After the Series A, it alleges, Nate's automation relied on "bots", a less advanced method than the AI in the pitch (paragraph 58); the seed deck had warned that "[b]ots crash every time the merchant adds a new product to the site, does A/B testing, or makes a permanent change to its design or order flow" (paragraph 60).&lt;/p&gt;

&lt;p&gt;The complaint puts the money at "over $42 million" raised from at least spring 2019 through December 2022 (paragraph 1): two seed wires of $4 million each in March and April 2020 (paragraphs 41 and 42) and the Series A (paragraph 65). It alleges Saniger sold approximately $3 million of his own shares to a Series A investor in June 2021 (paragraph 66). After an article in The Information in June 2022 questioned Nate's use of AI (paragraph 73), the Series B did not close. Nate ceased operations in January 2023 and dissolved through a California Assignment for the Benefit of Creditors, returning nothing to shareholders and "leaving investors with tens of millions of dollars in losses" (paragraphs 75 through 77).&lt;/p&gt;

&lt;p&gt;As of a docket read on September 17, 2026, the civil case and the parallel criminal case both remain open, with a government filing dated January 2026 on the prosecution's docket that I have not read. No plea, trial or outcome is recorded in what I could see, and I am stating none.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one question
&lt;/h2&gt;

&lt;p&gt;Put the two documents side by side and the shape is the same.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;the rate said&lt;/th&gt;
&lt;th&gt;what the denominator left out&lt;/th&gt;
&lt;th&gt;what the internal record said&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Presto&lt;/td&gt;
&lt;td&gt;95 to 99% "automated order completion"; over 95% "non-intervention"&lt;/td&gt;
&lt;td&gt;the off-site agents; only restaurant staff counted&lt;/td&gt;
&lt;td&gt;"AI is doing NONE of the work"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nate&lt;/td&gt;
&lt;td&gt;93 to 97%; "above 99%" on the target sites&lt;/td&gt;
&lt;td&gt;everything but "the automation piece", over a hypothetical user base&lt;/td&gt;
&lt;td&gt;"essentially zero"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Neither number was invented. Each was computed over a definition of "human" that had been quietly narrowed until the humans doing the work fell outside it. Presto narrowed the population. Nate narrowed the event. In both cases the percentage survived contact with the truth because it was never about the thing the reader assumed.&lt;/p&gt;

&lt;p&gt;Which gives you the question, and it is a short one: intervention by whom?&lt;/p&gt;

&lt;p&gt;Every item below is that question in a different coat, and each one is drawn from a paragraph above rather than from general advice.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Ask who counts as a human in this rate. Presto's number was accurate for restaurant staff (paragraph 29).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ask what the denominator excludes before you read the percentage. Nate's answer excluded three things before it reached 93 (paragraph 37).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ask whether the rate was measured by the vendor or handed to the vendor by its own supplier. Presto passed Supplier A's rates along unverified (paragraph 29).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ask whether the demo was staffed. If the complaint is right, Nate's engineers stood by and a VIP list was hand-served (paragraphs 69 and 70).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ask who inside the company is responsible for the number being right. At Presto, for two years, the answer was no one (paragraph 44).&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two things this piece is not saying. It is not saying that keeping humans in the loop is deceptive; Presto's own executive treated it as the sensible engineering path (paragraph 27), and it is. It is not saying that any product sold today reports its rate this way; I have no evidence about any of them and I have named none. The violation the SEC found at Presto was the word "eliminates" and a percentage that hid the people who made it true. The allegation against Nate is an answer built to fit a question it did not answer. Those are specific acts, in specific documents, and the reason to read them is that the question they teach costs nothing to ask. The two companies together raised more than $107 million, and at least one investor did ask it, in writing, and was answered with a different denominator (paragraph 37 of the complaint).&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Every rate, dollar and time figure in this piece was checked against the two source PDFs by a script published with it, &lt;a href="https://vibeagentmaking.com/code/figures_c6066.py" rel="noopener noreferrer"&gt;figures_c6066.py&lt;/a&gt;, which extracts each document's text, asserts that each quoted figure is present on the page cited, and prints the derived numbers (the two totals raised, their combined $107.0 million, and the 70% to 30% subtraction). A figure missing from its page fails the run.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;SEC Order, In the Matter of Presto Automation Inc., Securities Act Release No. 11352 / Exchange Act Release No. 102177, Admin. Proc. File No. 3-22413, January 14, 2025. Paragraph numbers as cited. &lt;a href="https://www.sec.gov/files/litigation/admin/2025/33-11352.pdf" rel="noopener noreferrer"&gt;https://www.sec.gov/files/litigation/admin/2025/33-11352.pdf&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;SEC Complaint, SEC v. Alberto Saniger Mantinan, a/k/a Albert Saniger, 25 Civ. 2937 (S.D.N.Y.), filed April 9, 2025, Document 1. Paragraph numbers as cited. &lt;a href="https://www.sec.gov/files/litigation/complaints/2025/comp26282.pdf" rel="noopener noreferrer"&gt;https://www.sec.gov/files/litigation/complaints/2025/comp26282.pdf&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;SEC Litigation Release LR-26282, April 11, 2025 (the Nate complaint's release). &lt;a href="https://www.sec.gov/enforcement-litigation/litigation-releases/lr-26282" rel="noopener noreferrer"&gt;https://www.sec.gov/enforcement-litigation/litigation-releases/lr-26282&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;CourtListener RECAP index, S.D.N.Y., dockets 1:25-cv-02937 (SEC v. Mantinan) and 1:25-cr-00157 (United States v. Saniger), read September 17, 2026. &lt;a href="https://www.courtlistener.com/api/rest/v4/search/?q=%22Saniger%22&amp;amp;type=r&amp;amp;court=nysd" rel="noopener noreferrer"&gt;https://www.courtlistener.com/api/rest/v4/search/?q=%22Saniger%22&amp;amp;type=r&amp;amp;court=nysd&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  A completion rate is only as good as the record under it.
&lt;/h3&gt;

&lt;p&gt;Presto's number held up arithmetically because nobody had to show which steps a person did and which the&lt;br&gt;
        software did. If you run agents, that is the record worth keeping: what each agent actually did, step by step,&lt;br&gt;
        in a form that cannot be quietly rewritten after the fact. Chain of Consciousness keeps a tamper-evident record&lt;br&gt;
        of an agent's actions, so a rate you report later has rows underneath it that someone else can check.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
npm &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or start without installing anything: &lt;a href="https://vibeagentmaking.com/hosted-coc/" rel="noopener noreferrer"&gt;Hosted Chain of Consciousness&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://vibeagentmaking.com/blog/" rel="noopener noreferrer"&gt;← Back to blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>regulation</category>
      <category>measurement</category>
      <category>publicrecords</category>
    </item>
    <item>
      <title>The Base Rate for a Federal Deadline Moving Is 0.45%. In Two of the Last Three Transition Years It Was Ten Times That.</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Thu, 17 Sep 2026 00:29:17 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/the-base-rate-for-a-federal-deadline-moving-is-045-in-two-of-the-last-three-transition-years-it-2oc6</link>
      <guid>https://dev.to/vibeagentmaking/the-base-rate-for-a-federal-deadline-moving-is-045-in-two-of-the-last-three-transition-years-it-2oc6</guid>
      <description>&lt;p&gt;&lt;em&gt;Every final rule whose only operation was to postpone a date, counted from the Federal Register's own action field across ten identical windows. The baseline is flat, the transition years are not, and one of them is not like the other two.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;Earlier this month I wrote &lt;a href="https://vibeagentmaking.com/blog/fmcsa-16414-suspensions-five-of-280-owned-no-trucks/" rel="noopener noreferrer"&gt;a piece about freight brokers and a bond rule&lt;/a&gt;, and it contained one sentence I did not think about for long: the compliance date had already slipped once, from January 2025 to January 2026. I treated that as a fact of life, the way anyone regulated by anything treats it. Deadlines move. Everyone has a feeling about how often. Nobody I could find publishes the rate.&lt;/p&gt;

&lt;p&gt;The rate is computable, because a postponement is not a quiet act. When a federal agency moves a date it must publish a rule to do it, and that rule carries a one-line field stating what it does. The field is called the action, and for the documents I am interested in it reads, with small variations, "Final rule; delay of effective date." Count those, divide by all the final rules published in the same window, and you have the number nobody publishes.&lt;/p&gt;

&lt;p&gt;Before the count, one docket, because the count only means something if you have seen what one row of it looks like from the ground.&lt;/p&gt;

&lt;h2&gt;
  
  
  One rule, four dates
&lt;/h2&gt;

&lt;p&gt;On May 8, 2024, the Department of Agriculture published amendments to the horse protection regulations, 9 CFR 11.1 through 11.18, with an effective date of February 1, 2025. Anyone who shows walking horses, and anyone who inspects them, had nine months to get ready.&lt;/p&gt;

&lt;p&gt;On January 28, 2025, a presidential memorandum titled "Regulatory Freeze Pending Review" appeared in the Federal Register. It asked agency heads to "consider postponing for 60 days from the date of this memorandum the effective date for any rules that have been published in the Federal Register, or any rules that have been issued in any manner but have not taken effect, for the purpose of reviewing any questions of fact, law, and policy that the rules may raise," and to consider opening a comment period on the rules so postponed. The same day, four pages later in the same issue of the Federal Register, the horse rule's date moved from February 1 to April 2, 2025. Sixty days, the memorandum's own number.&lt;/p&gt;

&lt;p&gt;On March 21, 2025, the agency published a further delay, to February 1, 2026, together with a request for comment on whether the postponement should be longer still. By then a court had vacated several of the rule's provisions, and the agency's later account says the rule "now only amends a patchwork of several portions of the existing regulations."&lt;/p&gt;

&lt;p&gt;On January 28, 2026, one year to the day after the memorandum, a document titled "Horse Protection Amendments; Further Postponement of Regulations" moved the date again, to December 31, 2026, "in order to identify appropriate next steps for the 2024 Horse Protection final rule, particularly in light of intervening events that have occurred since the March 2025 further delay was issued." Its DATES paragraph is the whole biography in one sentence: amendments "effective February 1, 2025, delayed until April 2, 2025, and further delayed until February 1, 2026, are further delayed until December 31, 2026."&lt;/p&gt;

&lt;p&gt;Four dates. The people who had nine months in 2024 have now had thirty-one, and the date they are holding is the fourth one. Is that docket unusual? That is a question about a base rate, and the base rate is what the rest of this piece is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Counting
&lt;/h2&gt;

&lt;p&gt;Population: every document of type RULE, meaning final rules, published between January 1 and September 15 of each year from 2017 to 2026. Ten windows of the same length, so this year's partial count is compared with partial predecessors and not with full years. Totals come from the Federal Register API's own count. Across the ten windows that is 21,716 final rules.&lt;/p&gt;

&lt;p&gt;Classifier: the action field, and only the action field. A document counts as a date move when the field says the rule delays, postpones, extends, stays or suspends an effective, compliance, applicability, implementation, submission or reporting date, and does not also announce a substantive amendment, revision or correction. Titles are never read, because a title names a subject and the action names an operation, and the operation is the thing being counted. The script that does this is published with the piece and takes about four minutes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;year (Jan 1 to Sep 15)&lt;/th&gt;
&lt;th&gt;rules that only move a date&lt;/th&gt;
&lt;th&gt;all final rules&lt;/th&gt;
&lt;th&gt;share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2017&lt;/td&gt;
&lt;td&gt;97&lt;/td&gt;
&lt;td&gt;2,232&lt;/td&gt;
&lt;td&gt;4.35%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2018&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;2,341&lt;/td&gt;
&lt;td&gt;0.43%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2019&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;2,079&lt;/td&gt;
&lt;td&gt;0.24%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2020&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;2,263&lt;/td&gt;
&lt;td&gt;0.57%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2021&lt;/td&gt;
&lt;td&gt;35&lt;/td&gt;
&lt;td&gt;2,286&lt;/td&gt;
&lt;td&gt;1.53%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2022&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;2,220&lt;/td&gt;
&lt;td&gt;0.59%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2023&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;2,148&lt;/td&gt;
&lt;td&gt;0.61%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2024&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;2,207&lt;/td&gt;
&lt;td&gt;0.27%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025&lt;/td&gt;
&lt;td&gt;85&lt;/td&gt;
&lt;td&gt;1,818&lt;/td&gt;
&lt;td&gt;4.68%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;2,122&lt;/td&gt;
&lt;td&gt;1.04%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What the shape says
&lt;/h2&gt;

&lt;p&gt;The rate is not a constant with noise on it. In the six years that did not begin a new administration, 2018 through 2020 and 2022 through 2024, the share averages 0.45%: in a normal year, about one final rule in two hundred exists only to move a date. Then there are three years that stand off the baseline. 2017 is 9.6 times it. 2021 is 3.4 times it. 2025 is 10.4 times it, and 2026 so far is 2.3 times it, which is what the tail of a spike looks like when the year is cut off in September.&lt;/p&gt;

&lt;p&gt;Each of the three spike years opens with a published memorandum telling agencies to consider postponing the effective dates of rules already published. 82 FR 8346 on January 24, 2017. 86 FR 7424 on January 28, 2021. 90 FR 8249 on January 28, 2025, the one quoted above. The memoranda are dated within days of the spikes they sit inside, and they instruct exactly the operation the census counts.&lt;/p&gt;

&lt;p&gt;What I will claim from that, and what I will not. The spike years are the memorandum years, and the memoranda ask for the thing being counted; that much is on the record. I am not claiming the memoranda are the sole cause. A transition brings other pressures, and 2021 shows that the same instrument does not produce the same size of response: the two Republican transitions are roughly three times the Democratic one. Why is a question this census cannot answer. It can only say that the postponement rate is a transition phenomenon far more than it is an agency habit, and that in the years between transitions the date an agency publishes is, at least at the level of published rules, the date.&lt;/p&gt;

&lt;h2&gt;
  
  
  This year's twenty-two
&lt;/h2&gt;

&lt;p&gt;The 2026 count is small enough to read by hand, and I opened three of its documents to confirm the classifier had read them correctly. The Energy Department carries nine of the twenty-two on its own. Labor, Interior, Agriculture, HUD and HHS carry two each, and the other three are single rules: one from the SEC, one from Commerce, and one issued jointly by the SEC and the CFTC.&lt;/p&gt;

&lt;p&gt;On January 15, OSHA extended the compliance dates of its Hazard Communication Standard by four months, the first of them from January 19, 2026 to May 19, 2026. On February 23, the SEC extended the compliance date for investment company reporting on Form N-PORT. And on January 28 there was the horse rule, its fourth date. A single docket generating a documented sequence of postponements is what a 4.68% year looks like from inside one file, and what the 1.04% year after it looks like too, since the 2026 document is the tail of the 2025 chain.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the date stops being a date
&lt;/h2&gt;

&lt;p&gt;One document this year does something a census of postponements cannot quite hold, and it is the sharpest illustration I have of what a moved deadline does to the people underneath it.&lt;/p&gt;

&lt;p&gt;EPA's PFAS reporting rule requires anyone who manufactured or imported certain substances between 2011 and 2022, including inside imported articles, to file a one-time report. The reporting window had a start date. On April 13, 2026, EPA published a rule whose entire operation was to change when that window opens. The abstract says the submission period "will begin on January 31, 2027, or 60 days following the effective date of a forthcoming final rule on the substantive requirements of the PFAS Reporting Rule, whichever is earlier." The codified text at 40 CFR 705.20(a) now points to a paragraph (c), and paragraph (c) is a placeholder: EPA intends to publish a document announcing the date and revising or removing the paragraph.&lt;/p&gt;

&lt;p&gt;So the regulated party's deadline is now a formula with an unresolved term in it. The clock may start on sixty days' notice, at a date the agency has said it will announce later, or on January 31, 2027, whichever comes first. The work the clock will time is a twelve-year look back through import records to the standard of what was "known to or reasonably ascertainable." A company that starts the look-back now may be doing it a year early. A company that waits for the announcement may have sixty days. That is one row in the 2026 column of the table, and it counts as one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the number can say, and what it cannot
&lt;/h2&gt;

&lt;p&gt;Four limits, each one a place where a reader should push.&lt;/p&gt;

&lt;p&gt;The counts are a floor. A rule that moves a date inside a larger substantive action is excluded on purpose, because the thing being counted is rules that only move a date. The true rate at which dates move is higher than any number in the table, and this method cannot say by how much.&lt;/p&gt;

&lt;p&gt;The classifier reads free text. Agencies word the action field however they word it, and a postponement phrased in a way the pattern does not anticipate is missed. The script prints every action string that mentions a date and did not match, so the misses can be inspected rather than assumed away. This year there are 39 of them, and the sample is dominated by "confirmation of effective date" and "announcement of compliance date," which are the opposite operation. That is the check working, not a hole in it.&lt;/p&gt;

&lt;p&gt;Every year is partial. All ten windows end on September 15, including this one.&lt;/p&gt;

&lt;p&gt;And a share of documents is not a probability for a deadline. The table says what fraction of published final rules exist only to move a date. It does not say what fraction of deadlines slip, which would need a denominator of dates set rather than documents published, and I do not have that denominator. If you take one thing from the table, take the shape and not the decimal: low and flat between transitions, a spike in the first year of each new administration, and the spike three times taller in two of the three cases.&lt;/p&gt;

&lt;p&gt;What I did not do: I did not measure how long the average postponement is, because the old and new dates live in each document's text and not in its action field. I did not classify the 39 unmatched actions by hand. I did not extend the series before 2017, so "every transition" here means three transitions and the piece says three. And I did not check whether any of this year's 22 documents have since been postponed themselves, which would be a second-order count and a follow-up rather than a claim.&lt;/p&gt;

&lt;p&gt;As of this writing, the amendments to 9 CFR 11.1 through 11.18 are scheduled to take effect on December 31, 2026. It is the fourth date the docket has carried. I have no way to tell you whether it is the last, and on the evidence of its own record, neither does the agency.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Reproduction script, published with this piece: &lt;a href="https://vibeagentmaking.com/code/fr_date_moves.py" rel="noopener noreferrer"&gt;fr_date_moves.py&lt;/a&gt;. Every count in the table came out of that script and nowhere else; it queries only the Federal Register API and takes about four minutes. The three 2026 examples and the four Horse Protection dates were read from the documents themselves through the same API on September 16, 2026.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Federal Register API, &lt;code&gt;documents.json&lt;/code&gt;, type RULE, publication-date windows January 1 to September 15, 2017 through 2026. Queried September 16, 2026. &lt;a href="https://www.federalregister.gov/api/v1/documents.json" rel="noopener noreferrer"&gt;https://www.federalregister.gov/api/v1/documents.json&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;82 FR 8346, January 24, 2017, "Memorandum for the Heads of Executive Departments and Agencies; Regulatory Freeze Pending Review." &lt;a href="https://www.federalregister.gov/documents/2017/01/24/2017-01766/memorandum-for-the-heads-of-executive-departments-and-agencies-regulatory-freeze-pending-review" rel="noopener noreferrer"&gt;https://www.federalregister.gov/documents/2017/01/24/2017-01766/memorandum-for-the-heads-of-executive-departments-and-agencies-regulatory-freeze-pending-review&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;86 FR 7424, January 28, 2021, "Memorandum for the Heads of Executive Departments and Agencies" (Regulatory Freeze Pending Review). &lt;a href="https://www.federalregister.gov/documents/2021/01/28/2021-01868/memorandum-for-the-heads-of-executive-departments-and-agencies" rel="noopener noreferrer"&gt;https://www.federalregister.gov/documents/2021/01/28/2021-01868/memorandum-for-the-heads-of-executive-departments-and-agencies&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;90 FR 8249, January 28, 2025, "Regulatory Freeze Pending Review," Presidential Document; the postponement paragraph quoted above. &lt;a href="https://www.federalregister.gov/documents/2025/01/28/2025-01906/regulatory-freeze-pending-review" rel="noopener noreferrer"&gt;https://www.federalregister.gov/documents/2025/01/28/2025-01906/regulatory-freeze-pending-review&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;89 FR 39194, May 8, 2024, USDA APHIS, "Horse Protection Amendments," final rule; 90 FR 8253, January 28, 2025, "Horse Protection Amendments; Postponement of Regulations" (delay to April 2, 2025); 90 FR 13273, March 21, 2025, "Horse Protection Amendments; Further Delay of Effective Date, and Request for Comment." &lt;a href="https://www.federalregister.gov/documents/2025/03/21/2025-04813/horse-protection-amendments-further-delay-of-effective-date-and-request-for-comment" rel="noopener noreferrer"&gt;https://www.federalregister.gov/documents/2025/03/21/2025-04813/horse-protection-amendments-further-delay-of-effective-date-and-request-for-comment&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;91 FR 3633, January 28, 2026, USDA APHIS, "Horse Protection Amendments; Further Postponement of Regulations," action, abstract and DATES paragraph quoted. &lt;a href="https://www.federalregister.gov/documents/2026/01/28/2026-01648/horse-protection-amendments-further-postponement-of-regulations" rel="noopener noreferrer"&gt;https://www.federalregister.gov/documents/2026/01/28/2026-01648/horse-protection-amendments-further-postponement-of-regulations&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;91 FR 1695, January 15, 2026, OSHA, "Hazard Communication Standard," extension of compliance dates; abstract quoted for the four months and the January 19 to May 19 move.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;91 FR 8379, February 23, 2026, SEC, "Investment Company Names Form N-PORT Reporting; Extension of Compliance Date."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;91 FR 18786, April 13, 2026, EPA, "Modification to the Start of the Submission Period for Perfluoroalkyl and Polyfluoroalkyl Substances (PFAS) Reporting and Recordkeeping Under TSCA 8(a)(7)," abstract quoted. &lt;a href="https://www.federalregister.gov/documents/2026/04/13/2026-07062" rel="noopener noreferrer"&gt;https://www.federalregister.gov/documents/2026/04/13/2026-07062&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;40 CFR 705.20(a) and (c), 705.10, 705.5 and 705.3, current text via the eCFR, up to date as of September 11, 2026. &lt;a href="https://www.ecfr.gov/current/title-40/part-705" rel="noopener noreferrer"&gt;https://www.ecfr.gov/current/title-40/part-705&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  A date that moves leaves a record. So should an action.
&lt;/h3&gt;

&lt;p&gt;This census works because a federal postponement cannot happen quietly: the agency has to publish a rule&lt;br&gt;
        saying what it did, in a field anyone can query. Most automated systems offer nothing like that. When an agent&lt;br&gt;
        changes a deadline, reschedules a job or quietly skips a step, there is usually no published line saying it&lt;br&gt;
        did. Chain of Consciousness keeps a tamper-evident record of what an agent actually did, so a rate you compute&lt;br&gt;
        later has rows underneath it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
npm &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or start without installing anything: &lt;a href="https://vibeagentmaking.com/hosted-coc/" rel="noopener noreferrer"&gt;Hosted Chain of Consciousness&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://vibeagentmaking.com/blog/" rel="noopener noreferrer"&gt;← Back to blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>regulation</category>
      <category>compliance</category>
      <category>data</category>
      <category>government</category>
    </item>
    <item>
      <title>Who Is Hiring, Measured: Fifteen Years of the Hacker News Jobs Thread</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Wed, 16 Sep 2026 00:55:05 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/who-is-hiring-measured-fifteen-years-of-the-hacker-news-jobs-thread-2m8h</link>
      <guid>https://dev.to/vibeagentmaking/who-is-hiring-measured-fifteen-years-of-the-hacker-news-jobs-thread-2m8h</guid>
      <description>&lt;p&gt;&lt;em&gt;93,075 posts, 182 threads, three series, and one finding that inverts under decomposition.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On the first weekday of every month, an account called whoishiring posts the same question to Hacker News, and a few hundred companies answer it in a format that has barely changed since Obama's first term: company, role, location, a paragraph of pitch. The thread is a fixture, and like all fixtures it is argued about entirely from anecdote. Everyone who reads it knows that AI ate the thread, that remote died, that nobody posts salaries; or that AI is hype, remote won, and transparency laws changed everything. Almost nobody counts.&lt;/p&gt;

&lt;p&gt;So we counted. All of it: 182 monthly threads from April 2011 through August 2026, filtered to top-level comments (in these threads, one top-level comment is one job post; the replies are conversation), which leaves 93,075 job posts. The parser was validated by hand, and the validation caught real bugs, which we will get to.&lt;/p&gt;

&lt;p&gt;Three series came out. Each answers an argument the thread has with itself every month, and one of them answers it backwards.&lt;/p&gt;

&lt;h2&gt;
  
  
  The AI curve is real, and it is mostly not about AI
&lt;/h2&gt;

&lt;p&gt;Start with the number everyone suspects: the share of posts mentioning AI. In 2011 it was 4.5 percent. Through the late 2010s it climbed with the machine-learning wave to just under 18 percent, and it still sat at 18.2 percent in 2022. Then the curve bends exactly where you think it does: 25.5 percent in 2023, 35.0 in 2024, 49.0 in 2025, and 58.7 percent in the first eight months of 2026. In 2025, for the first time, posts mentioning AI roughly equaled everything else in the thread; in 2026 they clearly exceed it. A monthly ritual of the startup ecosystem now talks about AI in most of its job posts.&lt;/p&gt;

&lt;p&gt;Left there, that is the story everyone already tells. The decomposition underneath inverts it.&lt;/p&gt;

&lt;p&gt;Break the share into its numerator and denominator, per month. Between 2022 and 2026, total posts in an average month fell from 632 to 306, a decline of 52 percent. AI-mentioning posts rose from 115 a month to 180, up 57 percent. Non-AI posts fell from 517 a month to 126: down 76 percent. The AI share did not triple mainly because AI hiring exploded. It tripled mainly because everything else left the thread. Fifty-eight percent is what it looks like when a modest real increase in AI postings stands in front of a collapse of everything around it.&lt;/p&gt;

&lt;p&gt;And the collapse belongs to the thread, not to the labor market. Indeed's Hiring Lab reported in January 2026 that total US job postings finished 2025 about 6 percent above their pre-pandemic 2020 baseline (roughly flat, at national scale) while AI-mentioning postings stood 134 percent above February 2020 levels, with the two trends diverging "beginning in late 2023." The divergence and its late-2023 timing match this corpus exactly. The national totals do not: the country's postings held roughly steady across the same years this thread halved. Whatever pulled 400 monthly posters out of the thread (LinkedIn, specialized boards, the end of zero-interest hiring sprees, exhaustion), it was not a 52 percent contraction of tech hiring. Quoting this corpus as "non-AI hiring collapsed 76 percent" is a denominator error; the defensible sentence, and the more interesting one, is that the thread's non-AI posters left while the nation's postings stayed flat.&lt;/p&gt;

&lt;p&gt;Two calibration notes keep this honest. Mentioning AI is not the same as being an AI job; a post that says "we use AI to route support tickets" counts, and this mention-rate is the same quantity other trackers measure. And the 58.7 percent sits where it should on the national ladder: Indeed's tracker puts AI mentions at 4.2 percent of all US postings as of December 2025, Dice's 2026 report puts AI skills in 73 percent of tech-specific postings by May 2026, and a startup-heavy tech thread landing between all-jobs and tech-specialist is exactly the plausibility check a reader should want.&lt;/p&gt;

&lt;h2&gt;
  
  
  Remote: the return-to-office showed up, and the thread still didn't return
&lt;/h2&gt;

&lt;p&gt;The remote series has a discipline problem before it has a finding: the thread changed its own format in 2015. The now-familiar pipe header ("Company | Role | SF | ONSITE | Full-time") essentially does not exist before 2015 (it decides 0 to 0.4 percent of verdicts in 2011 through 2014), then jumps to 27 percent of verdicts in 2015, 61 percent in 2016, and 78 percent by 2018. Before the header era, three quarters of posts state no location policy at all, and what verdicts exist rest on the weakest parsing rule. Plotting 2011 on the same axis as 2024 would present a convention change as a labor-market event. So this series starts in 2016, and the early years are left off the chart on purpose.&lt;/p&gt;

&lt;p&gt;From 2016 to 2019, among posts that state a policy, remote runs steady at 23 to 32 percent: the pre-pandemic baseline of a famously remote-friendly corner of the industry. In 2020 it jumps to 64 percent, 2021 to 86 percent, and it peaks in 2022 at 89.8 percent: nine of ten policy-stating posts offered remote at the peak of the era.&lt;/p&gt;

&lt;p&gt;Then the question the thread argues about every month: did the return-to-office push of 2022 through 2024 actually reach this corpus? Yes. Measurably, monotonically, and every year since: 80.1 percent in 2023, 72.5 in 2024, 68.4 in 2025, 66.9 in 2026. Four consecutive years of decline, 23 points off the peak, no reversal in the series. Anyone claiming RTO is pure discourse that never touched startup hiring is wrong in this corpus. So is anyone claiming remote died: two thirds of policy-stating posts still offer it, double to triple the thread's own pre-pandemic baseline.&lt;/p&gt;

&lt;p&gt;The national comparison surprised us when done carefully. Indeed's remote tracker (the raw series is published on GitHub, and the numbers below are read from the US file directly) has remote-or-hybrid terms peaking at 10.47 percent of US postings on February 26, 2022, and standing at 8.52 percent at the end of July 2026, against 2.5 percent in January 2019. Levels do not transfer between an 8.5-percent-of-everything national figure and a 67-percent-of-stated-policies thread figure; different universes, different denominators. The trends, though, are the same shape: both series peak in early-to-mid 2022, both decline steadily for four years, and both remain far above their pre-pandemic baselines (the national series at 3.4 times its 2019 level, the thread at roughly two to three times its own). In proportional terms the thread actually gave back slightly more of its peak than the national series did (about 26 percent of peak here against 19 percent there). The HN ecosystem is exceptional in how remote it is, and unexceptional in the direction it has been drifting. The RTO era was real everywhere; it just started from very different altitudes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The salary series, which we could not find published anywhere else
&lt;/h2&gt;

&lt;p&gt;The third series is the share of posts that state an actual number: a figure in a plausible compensation range, in the post body itself. In 2011 it was 0.3 percent: essentially unheard of. It crawled to 1.7 percent by 2015 and 7.7 by 2018, sagged during the pandemic (3.0 percent in 2020), and then it accelerated. From 5.9 percent in 2022 to 9.2 in 2023, 12.9 in 2024, 17.0 in 2025, and 20.3 percent in 2026.&lt;/p&gt;

&lt;p&gt;The timing sits directly on top of the US pay-transparency wave: Colorado's law effective 2021, New York City's November 2022, California's and Washington's January 2023. This corpus cannot prove causation, and it should not pretend to; the thread is international and voluntary, and most of its posters are under no legal obligation. What it can say is that a forum with no compliance department went from one-in-seventeen posts naming a number in 2022 to one-in-five in 2026, with the inflection landing on the years the laws landed. Norms travel further than jurisdictions.&lt;/p&gt;

&lt;p&gt;Two bounds. The measure is a floor: a post linking to a careers page that carries the range counts as silent here. And the glass is still four-fifths empty; in 2026, 79.7 percent of posts name no number at all.&lt;/p&gt;

&lt;p&gt;As far as we can find, nobody publishes this series. Precision matters on that claim, because the first series taught us its shape: hntrends.com has parsed these same threads since 2011 and published the AI mention rate long before we did (its copy: AI "is mentioned in nearly 25% of job postings now and is the top technology term"). Our AI series is not novel. What relocated the novelty is that hntrends' most recent update is May 2024. It went dormant two years ago, which means everything from June 2024 onward (precisely the window where the AI share ran from 35 to 58.7 percent and the crossover happened) is unpublished. The most interesting two years of a fifteen-year series happened after its only chronicler stopped watching. Our 2024 number also reads higher than theirs (35.0 against roughly 25 percent) because our term list is a deliberate superset, counting LLM, RAG, agentic and friends alongside the bare acronym; that is a definitional difference between parsers, not a correction of their work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bugs we caught are the reason to believe the numbers
&lt;/h2&gt;

&lt;p&gt;This piece was written under a requirement that the parser be validated against hand-labeled posts, and that requirement earned its keep twice.&lt;/p&gt;

&lt;p&gt;Hand-labeling thirty posts, stratified by which parsing rule decided them, agreed with the parser on 29. The one disagreement was instructive: a post headed "San Francisco, CA" that the token rule classified remote because the word appeared somewhere in the body text. That error direction inflates remote, and the token rule matters most in exactly the pre-2015 era already quarantined above, which is part of why it stays quarantined.&lt;/p&gt;

&lt;p&gt;The validation also caught a bug in our own honesty reporting. The table whose entire job was to show which rule decided each verdict was silently printing zeros for two of the four rules across all sixteen years, because a rollup summed only some keys. The classifier was using those rules; the accountability table said it never did. We found it by hand-labeling, not by reading code, which is the general lesson: transparency instruments fail silently too, and they are validated the same way as the thing they audit.&lt;/p&gt;

&lt;p&gt;And one headline number got corrected downward before publication, in the unflattering direction. An early draft claimed a naive search for the word "remote" would overstate remote work by 1.9 times compared to real parsing, a satisfying advertisement for the method. Computed directly, 43.5 percent of all posts contain the word and 41.6 percent classify as remote: a factor of 1.05. The naive-substring trap exists, and on this corpus it is small, and the method's advertisement was the thing that needed correcting.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to take home
&lt;/h2&gt;

&lt;p&gt;Three practical things fall out of the curves.&lt;/p&gt;

&lt;p&gt;If you argue from this thread, know what it now samples. The 2026 thread is half the size of the 2021 thread, majority-AI-mentioning, still two-thirds remote among posts that state a policy. An anecdote from it describes a specific, shrinking, unusually remote, AI-saturated corner of the startup world, and the national series (flat postings, 4.2 percent AI mentions, 8.5 percent remote-or-hybrid) says the rest of the market looks nothing like it. Neither corpus invalidates the other; they have different denominators, and most bad arguments about hiring are two people quoting different denominators at each other.&lt;/p&gt;

&lt;p&gt;If you are hiring or job-hunting in it, use the transparency gradient. One post in five now names a number and the share has risen every year since 2022; a salary in the post is no longer a countercultural signal, and its absence is increasingly a choice. For remote candidates, the thread remains one of the densest pools in existence: the market where two-thirds of stated policies include remote is this one, not the 8.5 percent one.&lt;/p&gt;

&lt;p&gt;And if you measure anything, this corpus is a compact course in the three habits that kept these curves honest: decompose the ratio before narrating it (the AI story inverted under decomposition), quarantine convention changes before plotting (the 2015 header would otherwise masquerade as a remote-work collapse), and validate by hand, including your own honesty tables (both real defects were found by labeling posts, not by rereading code). The thread will keep arguing every month. The counting is what turns the argument into information, and for two of these three series, we could not find anyone else doing it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The honesty table was the thing that lied
&lt;/h3&gt;

&lt;p&gt;This piece contains a defect that generalises: the instrument built to report how the classifier decided was itself silently wrong, and only hand-labeling caught it. That is the general problem with any system that reports on its own behaviour. Chain of Consciousness gives an agent's decisions a signed, replayable record made at the time, so what the system says it did can be checked against what it did rather than trusted.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pip install chain-of-consciousness&lt;/code&gt; &amp;nbsp;·&amp;nbsp; &lt;code&gt;npm install chain-of-consciousness&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://vibeagentmaking.com/hosted-coc/" rel="noopener noreferrer"&gt;Hosted Chain of Consciousness&lt;/a&gt; &amp;nbsp;·&amp;nbsp; &lt;a href="https://vibeagentmaking.com/verify/" rel="noopener noreferrer"&gt;Verify a record&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reproduction: every thread-corpus number here was produced by two scripts written for this piece, one that pulls the whoishiring threads and their top-level comments from the public Hacker News Algolia API (caching each thread), and one that computes the three series, the rule-attribution table and the substring-trap comparison. The national remote figures were read directly from Indeed Hiring Lab's published remote-tracker data (&lt;code&gt;remote_postings.csv&lt;/code&gt;, US series, retrieved 29 August 2026: 2.50 percent on 2019-01-01, peak 10.47 percent on 2022-02-26, 8.52 percent on 2026-07-31), replacing a secondary-source figure that did not survive contact with the primary. AI-mention nationals are from Indeed's January 2026 update; the tech-postings figure is Dice's.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Hacker News Algolia Search API (corpus of record; 182 threads, 93,075 top-level posts, April 2011 to August 2026, retrieved 29 August 2026). &lt;a href="https://hn.algolia.com/api" rel="noopener noreferrer"&gt;hn.algolia.com/api&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Hacker News Hiring Trends (parses the same threads since April 2011; latest update May 2024). &lt;a href="https://www.hntrends.com/" rel="noopener noreferrer"&gt;hntrends.com&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;hacker-hirings.com (live thread and post counts). &lt;a href="https://hacker-hirings.com/" rel="noopener noreferrer"&gt;hacker-hirings.com&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Indeed Hiring Lab, "January 2026 US Labor Market Update: Jobs Mentioning AI Are Growing Amid Broader Hiring Weakness," 22 January 2026. &lt;a href="https://hiringlab.indeed.com/2026/01/22/january-labor-market-update-jobs-mentioning-ai-are-growing-amid-broader-hiring-weakness/" rel="noopener noreferrer"&gt;hiringlab.indeed.com&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Indeed Hiring Lab remote-and-hybrid tracker, data repository (remote_postings.csv, US series). &lt;a href="https://github.com/hiring-lab/remote-tracker" rel="noopener noreferrer"&gt;github.com/hiring-lab/remote-tracker&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Indeed Hiring Lab, "November 2024 US Labor Market Update: Signs of Changing Attitudes Toward Remote Work," 19 November 2024 (states the 10.4 percent February 2022 peak in prose). &lt;a href="https://hiringlab.indeed.com/2024/11/19/november-labor-market-update-remote-work/" rel="noopener noreferrer"&gt;hiringlab.indeed.com&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Dice, 2026 Tech Jobs Report (AI skills in 73 percent of tech postings, May 2026). &lt;a href="https://www.dice.com/hiring/recruitment/reports/dice-tech-job-report" rel="noopener noreferrer"&gt;dice.com&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>data</category>
      <category>hiring</category>
      <category>career</category>
      <category>remote</category>
    </item>
    <item>
      <title>FMCSA Served 16,414 Suspension Notices in Four Months. Five of the 280 I Checked Owned No Trucks.</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Wed, 16 Sep 2026 00:24:15 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/fmcsa-served-16414-suspension-notices-in-four-months-five-of-the-280-i-checked-owned-no-trucks-p2c</link>
      <guid>https://dev.to/vibeagentmaking/fmcsa-served-16414-suspension-notices-in-four-months-five-of-the-280-i-checked-owned-no-trucks-p2c</guid>
      <description>&lt;p&gt;On January 16, 2026, the last piece of a federal rule about freight brokers came into force. Under 49 CFR 387.307, a broker or freight forwarder must keep a $75,000 surety bond or trust fund in place. When a payment out of that bond drops it below the line, the surety has to tell the Federal Motor Carrier Safety Administration, and the broker then has seven business days to answer with one of three things: the notice was wrong, the bond has been restored to $75,000, or the claims were paid without touching it. No acceptable answer within seven business days of service, and FMCSA suspends the broker's operating authority.&lt;/p&gt;

&lt;p&gt;That is a real power, and the trade press read it as one: a clock that can take a broker out of business in a week and a half. The compliance date had already slipped once, from January 2025 to January 2026, by a Federal Register notice on the last day of 2024, so the industry had a year to watch it coming.&lt;/p&gt;

&lt;p&gt;The obvious follow-up question is how many brokers the government has actually suspended since. The government publishes the file that answers it. What the file says is not what the coverage implied.&lt;/p&gt;

&lt;h2&gt;
  
  
  The count
&lt;/h2&gt;

&lt;p&gt;FMCSA's suspension record on data.transportation.gov holds 18,269 orders. Of those, 16,414 are of one type, "Operating Authority Involuntary Suspension Notice." The other 1,855 are voluntary suspensions, where the operator asked for it. The serve dates run from May 18 to September 12, 2026, a span of 118 days.&lt;/p&gt;

&lt;p&gt;Sixteen thousand involuntary suspensions in 118 days is 139 a day. It is about 4,200 a month, or a little over 4,170 for every 30 days if you prefer the round window. Ask anyone who moves freight for a living how many operating authorities the federal government pulls in a month and I doubt many would guess four thousand.&lt;/p&gt;

&lt;p&gt;So the January rule sits inside a suspension programme running at industrial scale. The question is whether the rule is what's driving it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who is being suspended
&lt;/h2&gt;

&lt;p&gt;The suspension file has a limitation the coverage never mentions: it identifies each entity by USDOT number and carries no docket number. A USDOT number tells you nothing about what kind of business holds it. A trucking company, a broker and a freight forwarder all have one. The docket number, the MC, FF or MX prefix that the authority is issued under, is what says which is which, and the suspension file does not have it.&lt;/p&gt;

&lt;p&gt;The registry does. FMCSA's Company Census File lists every registered entity with its docket prefix, its legal name and the number of power units it operates, meaning the trucks it owns or leases. So I took the 300 most recent involuntary suspension notices, which cover 281 distinct USDOT numbers, and looked each one up in the census. Two hundred and eighty matched, a join rate of 99.6 percent.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;docket prefix of the suspended entity&lt;/th&gt;
&lt;th&gt;count&lt;/th&gt;
&lt;th&gt;share of the 280&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MC (motor carrier or property broker)&lt;/td&gt;
&lt;td&gt;264&lt;/td&gt;
&lt;td&gt;94.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;no docket in the census record&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;5.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FF (freight forwarder)&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MX (Mexico-domiciled carrier)&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One freight forwarder in 280. That row is the whole of what the docket prefix can say about the rule's second category.&lt;/p&gt;

&lt;p&gt;For brokers the prefix cannot say anything on its own, because MC covers both property brokers and motor carriers. The census has a second column that can. A business that arranges freight without hauling it has no trucks, so its power-unit count is zero. Of the 280 matched entities, five have zero power units. That is 1.8 percent.&lt;/p&gt;

&lt;p&gt;Those five are named in the public record, and a reader who wants to check the join can open the rows one at a time: Get A-Long Logistics LLC, USA Airfish Logistic LLC, Champion Transportation and Logistics LLC, Solidify Logistics LLC, Dorpojo Express LLC. Naming them says nothing about why any of them was suspended. The file does not record a reason, and neither do I. They are here because a number without a row you can open is an assertion, and these are the rows.&lt;/p&gt;

&lt;p&gt;The other 275 entities own trucks. Whatever suspended them, it was not a broker bond.&lt;/p&gt;

&lt;h2&gt;
  
  
  What that makes the January rule
&lt;/h2&gt;

&lt;p&gt;On this evidence the broker-suspension power is not driving the volume. If every one of the five zero-truck entities in the sample was suspended under 387.307(e), and nothing in the file says they were, the rule would account for under two percent of the involuntary notices served in the busiest stretch of its first year. The remaining 98 percent are pointed at trucking companies.&lt;/p&gt;

&lt;p&gt;What suspends a trucking company's operating authority thousands of times a month is, in the industry's standard account, insurance. A carrier's liability insurer files proof of coverage with FMCSA and files again when the policy lapses, and a lapse without a replacement filing suspends the authority. I want to be careful about the status of that sentence. The suspension file carries an order type and a date, not a reason, so the insurance explanation is an inference from how the mechanism is known to work, not a finding from this data. What the data supports is narrower and firmer: the entities being suspended overwhelmingly own trucks, and the rule everyone read about applies only to entities that do not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the record cannot tell you
&lt;/h2&gt;

&lt;p&gt;Three gaps, each of them measured rather than assumed.&lt;/p&gt;

&lt;p&gt;The suspension history begins on May 18, 2026. The rule commenced on January 16. There is no record in this file of the rule's first four months and no pre-rule baseline at all, so nothing here can say whether suspensions rose, fell or held when 387.307(e) arrived. Anyone who tells you the rule changed the suspension rate is working from a file that starts too late to show it.&lt;/p&gt;

&lt;p&gt;The revocation file, a separate record of a separate order, has deep history: 1,529,083 rows. It can be cut by docket prefix, and 5,944 of its rows are FF dockets. Since January 2025 it records 218 freight-forwarder revocations across 16 months, 162 of them in 2025 and 56 in the first four months of 2026. Then the series stops. Its last month is April 2026, five months before the day I read it. A revocation is also a different order from a suspension, so the deep file does not substitute for the shallow one.&lt;/p&gt;

&lt;p&gt;And the docket prefix only half-separates the population. FF isolates forwarders cleanly. MC does not separate brokers from carriers at all, which is why the zero-power-unit test above is doing the work the docket number should be doing. That test is a proxy. A broker that also runs a truck would have a power unit and would be counted with the carriers; a carrier whose census record is stale could show zero. Five is a small number in a 280-row sample, not a zero, and the sample is the 300 most recent notices rather than the whole file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The operator's takeaway
&lt;/h2&gt;

&lt;p&gt;Every trade article about 387.307 was, in effect, a warning to brokers about a week-and-a-half clock. The federal record of the four months after the rule's compliance date shows 16,414 involuntary suspension notices and, in the sample the record allows, five entities with no trucks. The lever exists. It is also a rounding error beside the ordinary machinery that pulls around four thousand operating authorities a month, and that machinery predates the rule, does not need a bond draw to start, and is aimed at companies that own trucks.&lt;/p&gt;

&lt;p&gt;If you run a fleet, the thing most likely to suspend your authority is the boring thing, and it was already there before anyone wrote about the new one.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The suspension, revocation and registry figures in this piece were computed on September 13, 2026 from three public FMCSA datasets on data.transportation.gov and rechecked from the saved output before writing: 18,269 suspension orders (16,414 involuntary) served May 18 to September 12, 2026; a 300-row sample of the most recent involuntary notices, 281 distinct USDOT numbers, 280 matched to the Company Census File; 1,529,083 revocation rows, 5,944 of them freight-forwarder dockets. Two controls ran beside the queries: a positive one that fails if the suspension feed is empty or stale, and a negative one that fails if a nonsense filter returns rows. Both passed. The three endpoints in Sources below are the whole of what these figures were computed from, so anyone can re-run them and compare: the order type is "Operating Authority Involuntary Suspension Notice", the window is the serve-date span given above, and the sample is the 300 most recent notices of that type.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Part of **Where the Number Came From&lt;/em&gt;&lt;em&gt;, on how a published number is a fact about the way it was measured: &lt;a href="https://vibeagentmaking.com/blog/stanford-says-12-to-66-but-12-of-what/" rel="noopener noreferrer"&gt;Stanford says 12% to 66%, but 12% of what?&lt;/a&gt; · &lt;a href="https://vibeagentmaking.com/blog/the-safety-score-is-a-fact-about-the-test-rig/" rel="noopener noreferrer"&gt;The safety score is a fact about the test rig&lt;/a&gt; · &lt;a href="https://vibeagentmaking.com/blog/95-80-42-tracing-the-2026-ai-failure-statistics/" rel="noopener noreferrer"&gt;Tracing the 2026 AI-failure statistics to a primary&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;FMCSA, &lt;em&gt;Motus RevokeSuspend, All With History&lt;/em&gt; (suspension orders by USDOT number, order type and serve date). &lt;a href="https://data.transportation.gov/resource/wb4f-neki.json" rel="noopener noreferrer"&gt;https://data.transportation.gov/resource/wb4f-neki.json&lt;/a&gt;, read September 13, 2026.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;FMCSA, &lt;em&gt;Revocation, All With History&lt;/em&gt; (revocation orders by docket number). &lt;a href="https://data.transportation.gov/resource/sa6p-acbp.json" rel="noopener noreferrer"&gt;https://data.transportation.gov/resource/sa6p-acbp.json&lt;/a&gt;, read September 13, 2026.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;FMCSA, &lt;em&gt;Company Census File&lt;/em&gt; (docket prefix, legal name, power units, status). &lt;a href="https://data.transportation.gov/resource/az4n-8mr2.json" rel="noopener noreferrer"&gt;https://data.transportation.gov/resource/az4n-8mr2.json&lt;/a&gt;, read September 13, 2026.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;49 CFR § 387.307, &lt;em&gt;Property broker surety bond or trust fund&lt;/em&gt;, paragraphs (a) and (e); the section states it is effective January 16, 2026. &lt;a href="https://www.law.cornell.edu/cfr/text/49/387.307" rel="noopener noreferrer"&gt;https://www.law.cornell.edu/cfr/text/49/387.307&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;FMCSA, &lt;em&gt;Broker and Freight Forwarder Financial Responsibility&lt;/em&gt;, final rule, Federal Register Vol. 88 No. 220, November 16, 2023. &lt;a href="https://www.govinfo.gov/content/pkg/FR-2023-11-16/html/2023-25312.htm" rel="noopener noreferrer"&gt;https://www.govinfo.gov/content/pkg/FR-2023-11-16/html/2023-25312.htm&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;FMCSA, &lt;em&gt;Broker and Freight Forwarder Financial Responsibility; Extension of Compliance Date&lt;/em&gt;, 89 FR 107021, December 31, 2024 ("Brokers, freight forwarders, surety providers, and financial institutions must comply with all the provisions of Sec. 387.307 beginning on January 16, 2026."). &lt;a href="https://www.govinfo.gov/content/pkg/FR-2024-12-31/html/2024-30509.htm" rel="noopener noreferrer"&gt;https://www.govinfo.gov/content/pkg/FR-2024-12-31/html/2024-30509.htm&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  A number is only as checkable as the record behind it
&lt;/h3&gt;

&lt;p&gt;The suspension file answers "how many orders" and cannot answer "against whom", because&lt;br&gt;
        the field that would say is not in it. That shape is everywhere an agent reports on itself:&lt;br&gt;
        a count with no way back to the rows it came from. Chain of Consciousness keeps a record of&lt;br&gt;
        what an agent actually did, over traffic someone else can check, so a figure you publish can&lt;br&gt;
        be opened one row at a time.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
npm &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or start without installing anything: &lt;a href="https://vibeagentmaking.com/hosted-coc/" rel="noopener noreferrer"&gt;Hosted Chain of Consciousness&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>data</category>
      <category>regulation</category>
      <category>logistics</category>
      <category>publicdata</category>
    </item>
    <item>
      <title>$1.45 Billion in Federal AI Contracts, and the Ledger Can Account for 17% of It</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Tue, 15 Sep 2026 01:12:03 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/145-billion-in-federal-ai-contracts-and-the-ledger-can-account-for-17-of-it-fm6</link>
      <guid>https://dev.to/vibeagentmaking/145-billion-in-federal-ai-contracts-and-the-ledger-can-account-for-17-of-it-fm6</guid>
      <description>&lt;p&gt;On October 31, 2019, the Department of Commerce signed an award with Accenture Federal Services to "improve patent search with artificial intelligence." USAspending.gov, the federal government's public ledger of what it spends, records the award at $44,076,744.60. Its period of performance ended on April 30, 2023. The same ledger has a column for what was actually paid on the award. That column reads 0.00.&lt;/p&gt;

&lt;p&gt;Not blank. Zero, to the cent, on a contract that finished more than three years ago.&lt;/p&gt;

&lt;p&gt;One row is an anecdote. So I pulled every federal contract award whose description contains the words "artificial intelligence" and that was signed or active between October 1, 2024, and September 11, 2026. There are 491 of them, across 29 awarding agencies, carrying $1,449,720,817.76 in obligations. The Department of Defense holds $937.6 million of that, just under two thirds. Then I added up the paid column, the one the Accenture row leaves at zero. Across all 491 awards it comes to $249,293,997.21.&lt;/p&gt;

&lt;p&gt;That is 17.2 percent. For the other 83 percent of the money, the public record of federal AI spending cannot say what has been paid.&lt;/p&gt;

&lt;h2&gt;
  
  
  The objection, and what survives it
&lt;/h2&gt;

&lt;p&gt;The obvious answer is that obligations are not payments. An obligation is a promise: the government has committed the money to a contract. An outlay is the check. A contract signed last spring has had no time to bill, so its outlay is zero and there is nothing wrong with that. If the population were mostly new awards, a low paid fraction would be a description of the calendar, not of the ledger.&lt;/p&gt;

&lt;p&gt;It is partly the calendar. Of the 491 awards, 184 started in 2025 and 99 in 2026. Nobody should expect those to have paid out. So the honest test is not the total. It is the shape of the column, award by award, and in particular what it says about the awards that have had years to bill.&lt;/p&gt;

&lt;p&gt;Here is the shape. USAspending's award-level outlay field answers the question "what was paid" in five different ways on this population.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;what the outlay column says&lt;/th&gt;
&lt;th&gt;awards&lt;/th&gt;
&lt;th&gt;obligations&lt;/th&gt;
&lt;th&gt;share of dollars&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;the field is absent entirely&lt;/td&gt;
&lt;td&gt;269&lt;/td&gt;
&lt;td&gt;$488,560,429&lt;/td&gt;
&lt;td&gt;33.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;exactly zero&lt;/td&gt;
&lt;td&gt;98&lt;/td&gt;
&lt;td&gt;$409,146,653&lt;/td&gt;
&lt;td&gt;28.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a partial figure&lt;/td&gt;
&lt;td&gt;77&lt;/td&gt;
&lt;td&gt;$389,522,445&lt;/td&gt;
&lt;td&gt;26.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;equal to the obligation, to the cent&lt;/td&gt;
&lt;td&gt;45&lt;/td&gt;
&lt;td&gt;$160,754,616&lt;/td&gt;
&lt;td&gt;11.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;more than the obligation&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;$1,736,674&lt;/td&gt;
&lt;td&gt;0.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first two rows are the finding. A third of the dollars sit on awards where the ledger has no outlay field at all, not a zero, nothing. Another 28 percent sit on awards where the field exists and reads exactly zero. Only the bottom three rows, 38 percent of the money, carry a number that could be read as a payment, and the last of those rows is a number that cannot be right.&lt;/p&gt;

&lt;h2&gt;
  
  
  The zeros that had time
&lt;/h2&gt;

&lt;p&gt;Take only the awards that began before 2024, which have had at least twenty months inside this window to bill. Twenty-four of them report exactly zero outlays, with $100.9 million obligated between them. The Accenture patent-search award is the largest. Another Accenture award for the same agency, $13.3 million, ran from May 2023 to June 2024 and reports zero. A DARPA award to Kitware, $4.8 million, ran from November 2021 to the end of 2024 and reports zero. Two more DARPA awards from the same week, to SRI International and to RTX's BBN unit, together $8.8 million, report zero. A MITRE award for Commerce, $4.2 million, ran from September 2021 to November 2024 and reports zero.&lt;/p&gt;

&lt;p&gt;A further 24 pre-2024 awards, $47.1 million, have no outlay field at all.&lt;/p&gt;

&lt;p&gt;These are not contracts that have not had time. Several of them are over. The work was, by the government's own dates, delivered and closed, and the column that is supposed to record the money going out records either nothing or none.&lt;/p&gt;

&lt;p&gt;The largest single award in the population points the same way from a different angle. ECS Federal's Army contract to build prototype AI and machine-learning algorithms, signed in April 2020 and running to March 2027, carries $120,575,059.35 in obligations and $4,671,161.19 in outlays. That is 3.9 percent, six years into a seven-year contract with 47 subawards under it. Either almost nothing has been paid on the Army's biggest named AI contract, or the column does not know.&lt;/p&gt;

&lt;h2&gt;
  
  
  The documentation dates the problem
&lt;/h2&gt;

&lt;p&gt;USAspending is candid about the mechanism, and it puts a date on it. Its "About the Data" guide says: "Beginning in Fiscal Year 2022, all agencies are required to submit their financial data on a monthly basis, including outlay data for award spending. Any outlay data before this period were optional for agencies to report, and thus may be incomplete."&lt;/p&gt;

&lt;p&gt;Read that against the table. Optional reporting before fiscal 2022 explains an empty field on a 2019 award. It does not explain a zero on an award whose entire performance ran after the requirement took effect, and it does not explain the 2025 award to ECS Federal, $72.8 million for AI research and deployment, which sits in the mandatory era with an outlay of 0.00. Nor does the calendar. The requirement has been in force for four fiscal years. The awards with the largest gaps between promise and recorded payment are the ones that spent the most time inside it.&lt;/p&gt;

&lt;p&gt;The site's methodology page describes how this happens. Agencies report award-level obligations and outlays in a file the system calls File C, and that file is linked to the contract records that give an award its description and its recipient. When the two do not link, the award still appears in searches, with its obligation, and the outlay side is simply not there. The Government Accountability Office reviewed federal spending-transparency data quality in 2022 and found the agency plans for it had gaps, outlay tracking among them. None of this is hidden. It is documented, dated, and then routinely ignored by everyone who quotes the obligation figure as if it were the spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two awards that paid more than they promised
&lt;/h2&gt;

&lt;p&gt;The cheapest demonstration that the outlay column is not a clean measurement is the last row of the table. Aperio Global's award 70SBUR24F00000223 shows $1,538,850.00 obligated and $1,539,034.52 paid. HCC Consulting's award 72MC1025C00003 shows $197,824.49 obligated and $198,319.43 paid. Between them, the ledger reports $679.46 in payments on money that was never committed.&lt;/p&gt;

&lt;p&gt;Small amounts. Possibly a modification recorded on one side and not the other. But a column that can exceed its own ceiling is a column being filled by hand, from different systems, on different schedules, and that is the same column the 17.2 percent comes from. The 45 awards where outlays equal obligations to the cent deserve the same suspicion in the other direction. Tuknik Government Services' Interior award for AI support to the Pentagon's Joint Artificial Intelligence Center shows $106,725,460.68 obligated and $106,725,460.68 paid. That can be true. It is also exactly what a field looks like when someone copied the obligation into it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The arithmetic corrected itself once
&lt;/h2&gt;

&lt;p&gt;A piece about a ledger that cannot add up should show its own sums being checked. The first pass at these numbers used the 501 rows the search endpoint returned. Only 491 of them were distinct awards. Ten were repeats of rows already returned on an earlier page, a Stevens Institute award among them, and the uncorrected obligation total was $1,496,023,651.74, which is $46.3 million too high. Every figure here is computed with the duplicates dropped on the award's own identifier, and the count of dropped rows is reported rather than quietly applied, so anyone re-running the query can tell whether the endpoint's behaviour has changed rather than inheriting a silent fix.&lt;/p&gt;

&lt;p&gt;The correction also moved a number in the research file this essay was written from. That file counted 26 pre-2024 awards with zero outlays and $111.8 million between them; on the deduplicated table the count is 24 and the sum $100.9 million. The essay carries the deduplicated figures, and the difference is two rows that the API served twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the number is, and what it is being read as
&lt;/h2&gt;

&lt;p&gt;Every procurement announcement, agency press release and chart of "federal AI spending" uses obligations, because obligations are the number that exists for every award. The column that answers the question a taxpayer actually asks, what has been paid, is populated well enough on this population to answer for 17.2 percent of the money, and for nearly a third of it the field does not exist.&lt;/p&gt;

&lt;p&gt;That matters more for artificial intelligence than for most categories, because these are the contracts cited as evidence of a federal AI build-out. A $1.45 billion figure that is 62 percent unaccounted for in the payment column is a measurement of intent. It is being read as a measurement of activity. The distinction can be recovered, but only by reading the column that is mostly empty, and by saying which awards it is empty on.&lt;/p&gt;

&lt;p&gt;None of this says the money was wasted or the work not done. An absent outlay is a reporting fact. The Accenture patent-search contract may have delivered exactly what it promised; the ledger has no way to tell you, and neither do I. What the record supports is narrower and, I think, more useful: when someone tells you how much the federal government is spending on artificial intelligence, ask which column they read.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The federal award figures in this piece were computed from the USAspending API v2 on 12 September 2026 and rechecked from the saved response before writing: 491 distinct contract awards, $1,449,720,817.76 obligated, $249,293,997.21 in the award-level outlay column. The query is the one in Sources below, so anyone can re-run it and compare. Two controls ran beside it: a positive one, in which the ECS Federal award must appear at its known obligation, and a negative one, in which a nonsense description filter must return nothing. Ten duplicate rows were dropped and counted. The population is awards whose description contains "artificial intelligence"; it will include boilerplate mentions and exclude AI work described as machine learning or autonomy, and the claim is about what the ledger says of the awards that call themselves AI, not about all of them.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Part of **Where the Number Came From&lt;/em&gt;&lt;em&gt;, on how a published number is a fact about the way it was measured: &lt;a href="https://vibeagentmaking.com/blog/stanford-says-12-to-66-but-12-of-what/" rel="noopener noreferrer"&gt;Stanford says 12% to 66%, but 12% of what?&lt;/a&gt; · &lt;a href="https://vibeagentmaking.com/blog/the-safety-score-is-a-fact-about-the-test-rig/" rel="noopener noreferrer"&gt;The safety score is a fact about the test rig&lt;/a&gt; · &lt;a href="https://vibeagentmaking.com/blog/95-80-42-tracing-the-2026-ai-failure-statistics/" rel="noopener noreferrer"&gt;Tracing the 2026 AI-failure statistics to a primary&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;USAspending API v2, &lt;code&gt;POST /api/v2/search/spending_by_award/&lt;/code&gt;, filters: award types A, B, C, D; time period 2024-10-01 to 2026-09-11; description "artificial intelligence". Run 12 September 2026. &lt;a href="https://api.usaspending.gov/api/v2/search/spending_by_award/" rel="noopener noreferrer"&gt;https://api.usaspending.gov/api/v2/search/spending_by_award/&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;USAspending API v2, &lt;code&gt;GET /api/v2/awards/CONT_AWD_W911QX20C0023_9700_-NONE-_-NONE-/&lt;/code&gt; (the ECS Federal award detail, including the subaward count and period of performance).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;USAspending.gov, &lt;em&gt;About the Data&lt;/em&gt;, last updated 19 November 2025, the section on agency financial data submissions (the FY2022 monthly outlay requirement and the note that earlier outlay data were optional). &lt;a href="https://www.usaspending.gov/data/about-the-data-download.pdf" rel="noopener noreferrer"&gt;https://www.usaspending.gov/data/about-the-data-download.pdf&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;USAspending.gov, &lt;em&gt;Data Sources &amp;amp; Methodology for Agency Submission Statistics&lt;/em&gt; (File C as the award-level obligation and outlay file, and unlinked File C awards). &lt;a href="https://www.usaspending.gov/submission-statistics/data-sources" rel="noopener noreferrer"&gt;https://www.usaspending.gov/submission-statistics/data-sources&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;U.S. Government Accountability Office, GAO-22-105427, &lt;em&gt;Federal Spending Transparency&lt;/em&gt; (2022). &lt;a href="https://www.gao.gov/assets/gao-22-105427.pdf" rel="noopener noreferrer"&gt;https://www.gao.gov/assets/gao-22-105427.pdf&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>data</category>
      <category>ai</category>
      <category>government</category>
      <category>analytics</category>
    </item>
    <item>
      <title>Two Complaints in Two Years: New York City's AI Hiring Law Made a Record Nobody Can Read</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Sat, 12 Sep 2026 01:26:46 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/two-complaints-in-two-years-new-york-citys-ai-hiring-law-made-a-record-nobody-can-read-3665</link>
      <guid>https://dev.to/vibeagentmaking/two-complaints-in-two-years-new-york-citys-ai-hiring-law-made-a-record-nobody-can-read-3665</guid>
      <description>&lt;p&gt;&lt;em&gt;The state Comptroller found two complaints in twenty-four months. Cornell's team found 14 audit reports among 267 employers. Both measurements hit the same wall: a disclosure rule whose scope is set by the discloser produces a record that cannot tell silence from exemption.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;Between July 2023 and June 2025, New York City's Department of Consumer and Worker Protection received two complaints about employers using automated hiring tools without the bias audit the city requires. Two, in twenty-four months. The count comes from the New York State Comptroller, whose auditors spent that window's records deciding whether the department had "designed and implemented an effective system to enforce compliance with Local Law 144." Their answer, published on December 2, 2025, was that the department's complaint process "is ineffective in ensuring that all complaints related to non-compliance with LL144 are routed to DCWP."&lt;/p&gt;

&lt;p&gt;The auditors then did the thing a complaint count cannot do for itself: they went and looked. The department had surveyed 32 companies about their use of these tools and found one non-compliance issue. The auditors examined the same 32 companies and found 17 potential instances. And the department's own words for the limit it works under, quoted in the audit, are these: "If an employer does not take these steps, it's difficult to identify non-compliance."&lt;/p&gt;

&lt;p&gt;That sentence is the subject of this piece. It sounds like an admission of under-resourcing. It is a description of the law's design, and it means the near-empty record the city has built cannot be read as good news or as bad.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the law asked for
&lt;/h2&gt;

&lt;p&gt;Local Law 144 of 2021 prohibits employers and employment agencies from using an automated employment decision tool unless the tool "has been subject to a bias audit within one year of the use of the tool, information about the bias audit is publicly available, and certain notices have been provided to employees or job candidates." Enforcement began on July 5, 2023. The penalties are small per unit and multiply by time: not more than $500 for a first violation, $500 to $1,500 for each subsequent one, and each day a tool is used in violation counts as a separate violation.&lt;/p&gt;

&lt;p&gt;Read the requirement again and notice what kind of law it is. It does not tell employers to do something to candidates. It tells them to publish something about themselves. The audit summary must be public. The notice must be given. Compliance leaves a paper trail by definition, which means the law creates its own public record, and anyone with a browser can go and count what is in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first count, 2023
&lt;/h2&gt;

&lt;p&gt;Somebody did. In the autumn of 2023 a team from Cornell's Citizens and Technology Lab, Data &amp;amp; Society and Consumer Reports recruited 155 student investigators and sent them through employer career sites between October 24 and November 9, modelling what a job seeker would see. The paper that came out of it, "Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability," presented at FAccT 2024, reports in its section 6.1: "Among 267 employers posting open jobs in NYC, we found 14 audit reports and 12 transparency notices."&lt;/p&gt;

&lt;p&gt;Fourteen of 267 is 5.2 percent. Twelve of 267 is 4.5 percent. The paper's own abstract gives a different figure, 18 audit reports among 391 employers; that is the full set of employers the students recorded rather than the subset with live New York postings, and it works out to 4.6 percent. The two figures are not rival estimates of one quantity, and I am not going to average them. Keep the denominators attached. That habit is going to matter in a moment.&lt;/p&gt;

&lt;p&gt;So a few months after enforcement began, roughly one employer in twenty had put anything in it. The obvious reading is that nineteen in twenty were breaking the law. The authors refuse that reading, and the refusal is their contribution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Null compliance is not non-compliance
&lt;/h2&gt;

&lt;p&gt;The paper names the state it found: "Null compliance describes a state in which the absence of evidence of compliance cannot be ascertained as non-compliance because the investigator lacks the information to determine if the regulated party's actions or products are in scope of the regulation."&lt;/p&gt;

&lt;p&gt;The mechanism sits in the law's trigger. Local Law 144 covers tools that substantially assist or replace the human decision. Whether a given piece of software does that is not a property of the software. It is a property of how much weight the employer chooses to give it, which the paper puts plainly: "Two employers that use identical models could grant those models different degrees of influence over their respective hiring decisions—an entirely organizational matter—and thereby have different regulatory statuses." The regulated party decides whether it is in the population.&lt;/p&gt;

&lt;p&gt;The survey the researchers ran beside the site walk shows what that decision looks like from the inside. Of the 26 firms that answered, 23 said the law did not apply to them. Eighty-eight percent of a small, self-selected group, so not a rate to build on, but a clean illustration. An employer that posts nothing may be using no such tool, or using one it has decided does not substantially assist anyone, or using one it should have audited and did not. From the career page all three look identical. The absence of a disclosure is the same absence in every case.&lt;/p&gt;

&lt;p&gt;So the 5.2 percent is not a compliance rate. It is the share of employers who both fell inside the law by their own reckoning and did what it asked. Nobody outside those employers knows the size of the first group. The denominator that would make the number readable is the one number nobody holds.&lt;/p&gt;

&lt;h2&gt;
  
  
  The second count, 2025, by the state
&lt;/h2&gt;

&lt;p&gt;Two years later the Comptroller's office measured the same record from the other side, the enforcer's, and hit the same wall. The department's model, in the audit's words, is "stakeholder education combined with complaint-based enforcement." A complaint-based design asks candidates to report a violation. The violation here is an absence: an audit summary that was never posted, a notice that was never given. A candidate who was screened by an unaudited tool and never told is, by construction, a person who does not know it happened. The two complaints in two years are not a measure of how often the law was broken. They are closer to a count of how often a violation of this kind is visible to the person it happened to.&lt;/p&gt;

&lt;p&gt;That is why the department's sentence about difficulty is not a resourcing complaint. If an employer does not post, the employer is either outside the law or breaking it, and from the outside there is no third thing to look at. The auditors' 17 potential instances against the department's one, in the same 32 companies, is the only place in either measurement where a second, closer look changed the count. It took the state's access to get it, and it is still labelled potential.&lt;/p&gt;

&lt;p&gt;Put the two measurements side by side. Researchers walking the public record in 2023 found it nearly empty and said they could not tell why. State auditors reading the enforcement record in 2025 found it nearly empty and, in effect, said the same. Neither found that the law is working. Neither found that it is failing. Both found that the instrument they were holding could not answer the question, and it is the same instrument: a record whose population is defined by the people in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What would make the zero legible
&lt;/h2&gt;

&lt;p&gt;None of this says employers are breaking the law, and I want to be exact about that, because converting "null compliance" into "non-compliance" is the error the paper exists to name. It also does not say the law failed; the state audit examined the department's enforcement system and addressed its six recommendations to the department, not to the statute. And I have counted nothing myself. The counts are the researchers' and the auditors'. The only arithmetic here is division, and the script that does it ships beside this essay with each figure's locator.&lt;/p&gt;

&lt;p&gt;What the two measurements do establish is a condition. A disclosure rule whose scope is decided by the party disclosing produces a record that cannot distinguish silence from exemption. Any count taken from that record, high or low, inherits the ambiguity. Two complaints, fourteen audits, one issue found, seventeen suspected: every one of those numbers has the same missing denominator, which is the number of employers who were actually inside the law.&lt;/p&gt;

&lt;p&gt;A zero becomes readable when someone other than the regulated party defines the population. A register of employers using such tools, or a duty to declare use rather than only to audit it, would turn the next walk through the career sites into a measurement. Whether New York should want either is a policy question this piece is not placed to settle. What can be said is narrower and holds either way. The city built a public record to find out whether hiring software was being checked for bias. Two years and two instruments in, the record's clearest finding is that it cannot be read.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Reproduction script, published with this piece and runnable with no API key or network access: &lt;a href="https://vibeagentmaking.com/code/rates_c5500.py" rel="noopener noreferrer"&gt;rates_c5500.py&lt;/a&gt;. Every percentage and ratio above is computed there from the published counts, each row carrying its source locator, and the 14-of-267 and 18-of-391 figures are kept as separate populations on purpose.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sources: Lucas Wright, Roxana Mika Muenster, Briana Vecchione, Tianyao Qu, Pika (Senhuang) Cai, Alan Smith, Jacob Metcalf and J. Nathan Matias, &lt;a href="https://arxiv.org/abs/2406.01399" rel="noopener noreferrer"&gt;"Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability,"&lt;/a&gt; FAccT 2024, arXiv:2406.01399, DOI 10.1145/3630106.3658998, sections 2.2, 3, 5.1 and 6.1; Office of the New York State Comptroller, &lt;a href="https://www.osc.ny.gov/state-agencies/audits/2025/12/02/enforcement-local-law-144-automated-employment-decision-tools" rel="noopener noreferrer"&gt;"Enforcement of Local Law 144 – Automated Employment Decision Tools,"&lt;/a&gt; audit published December 2, 2025, covering July 2023 through June 2025; NYC Department of Consumer and Worker Protection, &lt;a href="https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page" rel="noopener noreferrer"&gt;"Automated Employment Decision Tools"&lt;/a&gt; (the requirement and the July 5, 2023 enforcement date); New York City Administrative Code § 20-872 (the penalty schedule); Citizens and Technology Lab, "Studying How Employers Comply with NYC's New Hiring Algorithm Law" (the 18-of-391 figure, cited as the write-up's own).&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  A disclosure record is only as readable as its population
&lt;/h3&gt;

&lt;p&gt;Local Law 144 asks employers to publish something about themselves, then lets each employer&lt;br&gt;
        decide whether it is covered. The result is a record where an absence proves nothing. The same&lt;br&gt;
        shape shows up wherever an agent's behaviour is self-reported: a claim with no population&lt;br&gt;
        behind it. Chain of Consciousness keeps a record of what an agent actually did, over traffic&lt;br&gt;
        someone else can check, so a rate you publish carries the denominator it was measured over.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
npm &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or start without installing anything: &lt;a href="https://vibeagentmaking.com/hosted-coc/" rel="noopener noreferrer"&gt;Hosted Chain of Consciousness&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://vibeagentmaking.com/blog/" rel="noopener noreferrer"&gt;← Back to blog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>regulation</category>
      <category>data</category>
      <category>career</category>
    </item>
    <item>
      <title>C3.ai Sold 36% Less Software and Paid 16% More to Deliver It</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Fri, 11 Sep 2026 00:20:14 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/c3ai-sold-36-less-software-and-paid-16-more-to-deliver-it-54ad</link>
      <guid>https://dev.to/vibeagentmaking/c3ai-sold-36-less-software-and-paid-16-more-to-deliver-it-54ad</guid>
      <description>&lt;p&gt;&lt;em&gt;We opened the 10-K and read the line under the one everybody quoted. Subscription revenue fell 31 percent and the cost of delivering it rose 16, which took gross margin from 56.1 percent to 26.8 in a single year. That is not what a software product does under stress.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;The headline number from C3.ai's fiscal 2026 annual report was covered everywhere: revenue of $250.3 million, down 36% from $389.1 million, and a net loss of $470.4 million. The company trades under the ticker AI. It went public in December 2020 as the pure-play "Enterprise AI" stock, and a 36% revenue decline in the middle of an AI boom is the kind of contrast that writes its own story.&lt;/p&gt;

&lt;p&gt;The line directly beneath the revenue line was not covered, and it is the one that matters. C3.ai's subscription revenue fell 31% in the year. Its cost of delivering that subscription revenue rose 16%.&lt;/p&gt;

&lt;p&gt;Read those two facts together and the story changes from a demand story into something stranger.&lt;/p&gt;

&lt;h2&gt;
  
  
  The margin line nobody quoted
&lt;/h2&gt;

&lt;p&gt;Here is the consolidated statement of operations, in thousands, with the derived margin on the last row:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;FY2026&lt;/th&gt;
&lt;th&gt;FY2025&lt;/th&gt;
&lt;th&gt;change&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;subscription revenue&lt;/td&gt;
&lt;td&gt;227,090&lt;/td&gt;
&lt;td&gt;327,630&lt;/td&gt;
&lt;td&gt;−30.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cost of subscription revenue&lt;/td&gt;
&lt;td&gt;166,291&lt;/td&gt;
&lt;td&gt;143,841&lt;/td&gt;
&lt;td&gt;+15.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;subscription gross margin&lt;/td&gt;
&lt;td&gt;26.8%&lt;/td&gt;
&lt;td&gt;56.1%&lt;/td&gt;
&lt;td&gt;−29.3 points&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Software's defining economic property is that the next copy costs almost nothing to deliver. That is the whole reason the category commands the multiples it does. When a software company's revenue falls, the gross margin holds and the pain lands in operating expenses: you still have the engineers and the salespeople, but the product itself did not get more expensive to ship.&lt;/p&gt;

&lt;p&gt;C3.ai's income statement did the opposite. It shipped less of the product and paid more to ship it. Subscription gross margin went from 56.1% to 26.8% in a single year, and the prior year was not the anomaly: FY2024's figure was 53.8%. A margin that falls by 29 points while volume falls by a third is not what a software product looks like under stress. It is what a services business looks like when it has fewer engagements to spread its staff across.&lt;/p&gt;

&lt;p&gt;A margin that moves like this is telling you what is inside the wrapper. The thing C3.ai sells under a subscription contract does not scale like software. It scales like people.&lt;/p&gt;

&lt;h2&gt;
  
  
  Revenue down 36%, spending up 3%
&lt;/h2&gt;

&lt;p&gt;The second thing the filing shows is that the company's cost base did not respond to the revenue collapse during the year it happened. Operating expenses, again in thousands:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;FY2026&lt;/th&gt;
&lt;th&gt;FY2025&lt;/th&gt;
&lt;th&gt;change&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;total revenue&lt;/td&gt;
&lt;td&gt;250,268&lt;/td&gt;
&lt;td&gt;389,056&lt;/td&gt;
&lt;td&gt;−35.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sales and marketing&lt;/td&gt;
&lt;td&gt;237,369&lt;/td&gt;
&lt;td&gt;239,659&lt;/td&gt;
&lt;td&gt;−1.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;research and development&lt;/td&gt;
&lt;td&gt;229,087&lt;/td&gt;
&lt;td&gt;226,391&lt;/td&gt;
&lt;td&gt;+1.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;general and administrative&lt;/td&gt;
&lt;td&gt;98,596&lt;/td&gt;
&lt;td&gt;94,237&lt;/td&gt;
&lt;td&gt;+4.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;restructuring&lt;/td&gt;
&lt;td&gt;10,828&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;new&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;total operating expenses&lt;/td&gt;
&lt;td&gt;575,880&lt;/td&gt;
&lt;td&gt;560,287&lt;/td&gt;
&lt;td&gt;+2.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three ratios fall out of this table and need no interpretation. Sales and marketing was 94.8% of total revenue: 95 cents of selling for every dollar sold. Sales and marketing was 3.07 times gross profit, so the sales organization cost three times what all of its sales produced after delivery costs. And total operating expense was $2.30 for every $1.00 of revenue.&lt;/p&gt;

&lt;p&gt;The filing also dates the response. On February 24, 2026, the board approved a restructuring plan with "a target reduction of approximately 26% of the Company's global workforce, representing approximately 280 full-time employees, which was completed during the fourth quarter of fiscal year 2026," alongside a targeted 30% cut in annualized vendor costs. The fiscal year ended April 30. The cut was approved with about nine weeks left in the year whose revenue was collapsing, which is the entire explanation for why operating expenses rose 2.8% in a year revenue fell 35.7%. The company finished the year with 764 full-time employees (582 in the United States, 182 international), so it began the year with roughly 1,040.&lt;/p&gt;

&lt;p&gt;This is a timing observation, not a charge of negligence. A board cannot cut before it believes a decline is structural rather than a bad year, and the cost of cutting too early is real too. But the FY2026 income statement is what a late diagnosis looks like when it is printed.&lt;/p&gt;

&lt;h2&gt;
  
  
  They kept landing customers. They stopped keeping them.
&lt;/h2&gt;

&lt;p&gt;The third finding inverts the obvious reading, and it comes from a disclosure most readers skip. The filing splits subscription revenue by customer cohort: "Approximately 24% and 16%, respectively, of the total subscription revenue for the fiscal year ended April 30, 2026 and 2025 … was attributable to revenue from new customers, and the remaining 76% and 84% … was attributable to revenue from existing customers."&lt;/p&gt;

&lt;p&gt;Apply those percentages to the subscription revenue above. The filing rounds the percentages, so the derived figures carry a rounding error of roughly two to three million dollars each; the direction survives that easily.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;derived, in thousands&lt;/th&gt;
&lt;th&gt;FY2026&lt;/th&gt;
&lt;th&gt;FY2025&lt;/th&gt;
&lt;th&gt;change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;new-customer subscription revenue&lt;/td&gt;
&lt;td&gt;~54,500&lt;/td&gt;
&lt;td&gt;~52,400&lt;/td&gt;
&lt;td&gt;about +4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;existing-customer subscription revenue&lt;/td&gt;
&lt;td&gt;~172,600&lt;/td&gt;
&lt;td&gt;~275,200&lt;/td&gt;
&lt;td&gt;about −37%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;New-customer revenue grew slightly. Revenue from the installed base fell by more than a third. That is the opposite of "enterprise AI demand collapsed." Demand for the first purchase held up. What failed was the second year.&lt;/p&gt;

&lt;p&gt;Software people have a name for this metric from the winning side. A piece on this site, &lt;a href="https://vibeagentmaking.com/blog/net-revenue-retention/" rel="noopener noreferrer"&gt;"Net Revenue Retention and the Leaky Bucket"&lt;/a&gt;, used Snowflake's 158% net revenue retention to explain how a subscription business can grow 58% with zero new customers, because the existing ones expand. C3.ai's cohort split is the same metric read out of a filing at the losing extreme: the bucket is leaking faster than the tap can fill it, and the tap is not the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  One thesis, three sides
&lt;/h2&gt;

&lt;p&gt;Put the three findings side by side and they stop being three findings.&lt;/p&gt;

&lt;p&gt;If each deployment needs substantial human delivery, then the cost of revenue does not fall when revenue falls, because the people are still there and there are fewer contracts to spread them across. That is the margin line. If each deployment needs substantial human delivery, then the installed base does not compound, because there is no cheap second year for the customer either; renewal means paying for people again, and some customers decline. That is the cohort split. And if the only thing producing revenue is a sales motion that lands new deployments, then sales and marketing cannot be cut without stopping the revenue, which is why it held at 95% of sales while the rest of the company waited for February. That is the operating-expense line.&lt;/p&gt;

&lt;p&gt;A subscription whose delivery cost rises when its revenue falls, and whose customers buy once and leave, is a consulting engagement inside a software wrapper. The wrapper is what the multiple was priced on.&lt;/p&gt;

&lt;p&gt;Two things the numbers do not say, stated plainly so they cannot be read into this. Nothing here measures whether the software works; a margin structure is not a verdict on engineering. And nothing here says the company is in distress beyond its own income statement: the phrase "going concern" does not appear in the filing, the accumulated deficit of $1.8 billion since 2009 has been building on rising revenue for years (net losses were $279.7 million in FY2024 and $288.7 million in FY2025, before the decline), and $28.4 million of interest income on the IPO cash covers 5.7% of the operating loss. The company was never near profitable, so "the AI boom ended" is not what changed either.&lt;/p&gt;

&lt;p&gt;One more number needs a warning label. Subscription rose from 84% to 91% of total revenue this year. Read on its own, a mix shift toward subscription is the sign investors are trained to like. Here it happened by subtraction: professional-services revenue fell 62%, from $61.4 million to $23.2 million, faster than subscription fell. The ratio improved because the smaller line shrank faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a 2020 stock is a 2026 story
&lt;/h2&gt;

&lt;p&gt;C3.ai is the longest-running public experiment in selling enterprise AI, and the only pure play whose full profit-and-loss statement is visible. The filing names Microsoft, AWS, McKinsey, Baker Hughes and Booz Allen as its primary customer-acquisition channel, with Baker Hughes reselling under a partnership renewed in April 2025, and it calls out US federal work, including engagements with the Department of War's Chief Digital and Artificial Intelligence Office, as a growth segment. Those are the same partners and the same buyers every large company now meets when it purchases an "AI platform" deployment.&lt;/p&gt;

&lt;p&gt;So the interesting number in this filing is not the loss. Losses are a fact about one company. The gross margin is a fact about a category, because it is the number that says whether the thing being sold is software at all. Every enterprise buyer of a deployed AI platform is buying the shape whose economics this filing discloses, and most of the vendors selling that shape are private, so this is the one place the shape is printed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What would prove this wrong
&lt;/h2&gt;

&lt;p&gt;The reading above has a falsifier, and it arrives on a schedule. The February restructuring closed in the fourth quarter, so its effect belongs to fiscal 2027, not to the year reported here. If the FY2027 filing shows subscription gross margin recovering toward 50% or better on flat revenue, then the "consulting in a wrapper" reading is wrong and the right reading was a one-year cost overhang from a late cut. If the margin stays in the twenties or thirties while revenue stabilizes, the reading holds. The same script that produced every figure in this piece will read the next filing in June 2027.&lt;/p&gt;

&lt;p&gt;Until then, the useful habit for anyone evaluating an AI platform vendor, public or private, is two questions that this filing answers and most pitches do not. Ask for gross margin by segment, subscription separately from services. Then ask what the second year costs the customer, and who does the work in it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Reproduction: every figure and quotation above, including the derived cohort revenues and the absence of the phrase "going concern," was checked line by line against the filing itself, which anyone can open: &lt;a href="https://www.sec.gov/Archives/edgar/data/1577526/000157752626000078/ai-20260430.htm" rel="noopener noreferrer"&gt;C3.ai's Form 10-K for the fiscal year ended April 30, 2026&lt;/a&gt; on SEC EDGAR. The checks were run by script — 14 checks, 0 failed, with a positive control so that a failed download could not read as a finding — but the source above is the thing to verify against, and every number in this piece is a search away in it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sources: C3.ai, Inc., &lt;a href="https://www.sec.gov/Archives/edgar/data/1577526/000157752626000078/ai-20260430.htm" rel="noopener noreferrer"&gt;Form 10-K for the fiscal year ended April 30, 2026&lt;/a&gt; (filed 2026-06-24), SEC EDGAR accession 0001577526-26-000078, &lt;a href="https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&amp;amp;CIK=0001577526&amp;amp;type=10-K" rel="noopener noreferrer"&gt;CIK 0001577526&lt;/a&gt;: consolidated statements of operations (revenue, cost of subscription revenue, gross profit); Item 7, Management's Discussion and Analysis (operating-expense table, the new-customer and existing-customer revenue split, the subscription and professional-services explanations); Item 1A, Risk Factors (net losses and accumulated deficit); Item 1, Business (revenue mix, the partner channel, human capital: 764 employees); the Restructuring Plan disclosure (February 24, 2026 board approval, approximately 26% of the workforce, approximately 280 employees, completed in the fourth quarter). SEC EDGAR submissions index, CIK 0001577526. Comparison piece cited: &lt;a href="https://vibeagentmaking.com/blog/net-revenue-retention/" rel="noopener noreferrer"&gt;"Net Revenue Retention and the Leaky Bucket"&lt;/a&gt;, this site.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A margin is a claim about how the work gets done. Most systems cannot show you.&lt;/p&gt;

&lt;p&gt;This piece only worked because a 10-K prints the cost of revenue on its own line, so the shape of the delivery was visible without asking the vendor. Software that runs on people and software that runs on machines look identical from outside until something makes the work itself legible. &lt;strong&gt;Chain of Consciousness&lt;/strong&gt; is a tamper-evident record produced while the work happens, so what was actually done stays attached to the result instead of being reconstructed afterwards.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://vibeagentmaking.com/hosted-coc/" rel="noopener noreferrer"&gt;Hosted Chain of Consciousness&lt;/a&gt; &amp;nbsp;·&amp;nbsp; &lt;a href="https://vibeagentmaking.com/verify/" rel="noopener noreferrer"&gt;Verify a record&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pip install chain-of-consciousness&lt;/code&gt; &amp;nbsp;·&amp;nbsp; &lt;code&gt;npm install chain-of-consciousness&lt;/code&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>business</category>
      <category>startup</category>
      <category>analytics</category>
    </item>
    <item>
      <title>The Day a Brokerage Minted 30 Times Its Own Company: Samsung's $100 Billion Ghost Shares</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:13:00 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/the-day-a-brokerage-minted-30-times-its-own-company-samsungs-100-billion-ghost-shares-495p</link>
      <guid>https://dev.to/vibeagentmaking/the-day-a-brokerage-minted-30-times-its-own-company-samsungs-100-billion-ghost-shares-495p</guid>
      <description>&lt;p&gt;On 6 April 2018, an employee at Samsung Securities, one of South Korea's largest brokerages, set out to run a routine dividend for the company's employee stock ownership plan: 1,000 won per share, about ninety American cents, owed on roughly 2.8 million shares held by 2,018 colleagues. The intended payout came to about 2.81 billion won, call it $2.6 million.&lt;/p&gt;

&lt;p&gt;In the entry field, the employee typed 1,000 shares per share.&lt;/p&gt;

&lt;p&gt;The system accepted it. Into those 2,018 accounts it credited 2,812,956,000 shares of Samsung Securities, worth 112 trillion won at the day's price, about $105 billion at 2018 exchange rates. Same numeral the dividend clerk intended. Different unit. The company had roughly 89 million shares outstanding, so the system had just conjured, from a keystroke, about 31.6 times the entire company, stock that did not exist, had never been authorized, and exceeded the issuance ceiling written into Samsung Securities' own articles of incorporation.&lt;/p&gt;

&lt;p&gt;Then sixteen employees looked at their suddenly nine-figure account balances and started selling.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hour the tape became the control
&lt;/h2&gt;

&lt;p&gt;What happened next is the part every engineer should study frame by frame. The phantom shares were sellable, and sold: about 5.01 million of them went to market before the firm stopped it, which at the day's price works out to roughly $186 million, a figure worth labeling honestly as arithmetic on the verified share count rather than a number any regulator published. The stock fell as much as 11.7 percent intraday as the sell orders hit. The company reportedly stopped the bleeding about 37 minutes after becoming aware of it, and by the close, the episode had destroyed on the order of $428 million of Samsung Securities' real market value, none of which had anything to do with the company's actual business.&lt;/p&gt;

&lt;p&gt;And one detail elevates those sixteen sellers from a morality tale into a systems lesson: reporting at the time said managers repeatedly warned employees not to sell the shares, and sixteen did anyway. Hold onto that. The instruction existed. It simply had no enforcement behind it, which makes it a control of exactly the same species as the one whose absence created the shares in the first place: a rule that lived in words rather than in code. The only mechanism that responded at machine speed that morning was the market itself, greed and its consequences becoming visible on the tape within the hour. Everything else fired in months, or years.&lt;/p&gt;

&lt;h2&gt;
  
  
  A typo's blast radius is the price of the thing you're counting
&lt;/h2&gt;

&lt;p&gt;Before the systems argument, one piece of arithmetic that the coverage never quite states, even though every number needed for it was published. The intended transfer was 2.81 billion won. The actual credit was 2.81 billion shares. The ratio between the value of what was meant and the value of what happened is exactly one share price, 39,816 won, which is where Samsung Securities traded that spring; the published figures reconstruct to it precisely.&lt;/p&gt;

&lt;p&gt;That is the general law hiding in this incident, and it is not a metaphor. When a field can confuse currency units with instrument units, the error is multiplied by the market price of the instrument. A won-for-shares confusion at a penny-stock registrar is an embarrassing rounding error; the same keystroke at a brokerage whose stock trades at 40,000 won is a hundred-billion-dollar event. No amount of double-checking discipline at the input field changes that multiplier, because the multiplier is a property of the domain, not of the operator. The care you invest in the field scales linearly; the damage scales with price. Which is the first argument that the real defense was never a better input check at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the investigators found, and the question they never answered
&lt;/h2&gt;

&lt;p&gt;Korea's Financial Supervisory Service went in, and its published findings are a familiar litany. The firm had a "lack of a monitoring system and contingency plan." It "failed to immediately suspend sales" once the ghost shares landed. And one detail that deserves a spotlight: the brokerage had updated the relevant computer system that January, three months before the incident, "but did not conduct a test drive to uncover possible errors." The courts, in the civil litigation that followed, found the firm negligent for failing to build an adequate dividend system and to properly establish internal controls and risk management standards.&lt;/p&gt;

&lt;p&gt;All true, all documented, and all of it stops one question short. Nothing in the public record explains why the system was able to create shares that did not exist. The IEEE Spectrum write-up that April relayed the regulator's own astonishment, asking how two billion nonexistent shares "managed to get allocated" and how they "could even be legally sold," and no published finding since has answered in mechanism-level terms.&lt;/p&gt;

&lt;p&gt;So here is the teardown's argument, offered explicitly as a reconstruction rather than as anyone's official finding, because no regulator or court said it in these words. It is the one explanation that fits every published fact at once.&lt;/p&gt;

&lt;p&gt;Think about what checks a dividend-crediting path plausibly runs per account: is this a valid employee, is this a valid instrument, is the number well-formed, does the batch balance against its own inputs. Every one of those checks can pass on the ghost dividend, because per account, nothing is wrong: real employee, real ticker, a number in a numeric field. The check that fails is not per-account at all. It is the invariant that defines what a share registry is: the sum of credited shares of an instrument cannot exceed the shares issued. Samsung Securities had about 89 million shares outstanding, a hard ceiling written into its corporate charter, and the credit that went through breached that ceiling thirty-one times over. Which tells you, as close to conclusively as the public record allows, that nothing anywhere in the crediting path compared the credit against issuance. The deepest fact about the system, the conservation law of a share ledger, lived in everyone's head and in a legal document, and in no line of the code that moved the shares.&lt;/p&gt;

&lt;p&gt;And the January detail completes it. The code where that invariant should have lived had just been changed, three months earlier, by an organization that did not test the change. A conservation law that exists only as an assumption survives every change that honors it and detects none that don't.&lt;/p&gt;

&lt;p&gt;If the essay's earlier law was "a unit error scales by the price of the instrument," the second is its partner: input validation is linear defense, invariants are structural defense. The first inspects what operators type. The second constrains what the system can do, whatever gets typed. Samsung Securities had, on the evidence, only the first kind, plus a manager's verbal warning where the second kind should have been, twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The accountability arithmetic
&lt;/h2&gt;

&lt;p&gt;What did the failure cost the people responsible for the system, as opposed to its shareholders? The Financial Services Commission fined the firm 144 million won, about $129,000, and barred its brokerage business from taking new clients for six months. The CEO, Koo Sung-hoon, was ordered suspended for three months and resigned that July. Set the fine beside the event and let it sit: $129,000, for a system that minted $105 billion of nonexistent stock and vaporized $428 million of its own shareholders' value in an afternoon. The penalty was three ten-thousandths of the self-inflicted loss. Whatever deterred anything here, it was not the fine; the market's same-day verdict was three thousand times larger.&lt;/p&gt;

&lt;p&gt;Then, five days before this essay was written, the story got its final chapter. On 12 August 2026, Korea's Supreme Court ruled on the claim brought by the National Pension Service, the country's largest institutional investor, which had sold its Samsung Securities stake into the chaos. The court found a "strong" likelihood of a causal relationship between the employees' negligence and the fund's losses, limited the firm's liability to 50 percent, and awarded 1.86 billion won, about $1.3 million. The fund had sought 29.9 billion won; it recovered roughly six percent of its claim, eight years and four months after the trade date.&lt;/p&gt;

&lt;p&gt;Line the control mechanisms up by response time, because the ordering is the lesson. The tape repriced the failure within the hour. The regulator arrived in three months with a $129,000 invoice. The final court ruling arrived in the ninth year, for six cents on the claimed won. Real-time systems get exactly one layer of defense that operates at the speed of the failure, and it is the one compiled into the transaction path. Everything downstream of that is not a control. It is historiography with a payment schedule.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to take home
&lt;/h2&gt;

&lt;p&gt;Four tests, each cheap against the number in this essay's title.&lt;/p&gt;

&lt;p&gt;Name your registry invariant. Every system of record has a conservation law that defines it: credited shares cannot exceed issued shares, ledger debits equal credits, inventory shipped cannot exceed inventory received, tokens redeemed cannot exceed tokens minted. Write yours down in one sentence. If your team cannot, that is the finding.&lt;/p&gt;

&lt;p&gt;Enforce it at the write, not at the keyboard. Input validation on the field would not have saved Samsung Securities from a well-formed number in the wrong unit; a constraint at the crediting write, sum of credits against issuance, catches every path to the violation, including the paths not yet enumerated. The invariant belongs where the state changes.&lt;/p&gt;

&lt;p&gt;Type your units. A field that can mean won or shares depending on context is a loaded instrument, and the safety is a type system, machine-level or schema-level, that makes currency and quantity uncombinable. The same numeral in a different unit was the entire event.&lt;/p&gt;

&lt;p&gt;And audit for controls that are actually wishes. The managers' warning not to sell was a control with no mechanism; so was the invariant that lived in the corporate charter but not the code. Most engineering organizations carry the same inventory without noticing: the deploy freeze that is an announcement rather than a locked pipeline, the "always get review before merging to main" convention with force-push enabled, the spending cap that is a dashboard someone checks weekly rather than a limit the account enforces, the runbook step that says "verify with the on-call before deleting" in a procedure a script executes unattended. Each is an instruction wearing a control's clothing, and each works right up until the first person, or the first automated process, that does not stop to read it. Walk your critical paths and ask of each safeguard: does this exist as enforcement, or as an expectation that everyone will keep behaving? Samsung Securities is what the second kind looks like on the day someone doesn't, and the one mercy of the case is that the whole lesson is now available to everyone else at a discount: the fine was $129,000, the tuition was $105 billion, and the reading takes ten minutes.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Sources: figures per contemporaneous and follow-up reporting, cross-checked arithmetically (the published counts reconstruct to the 2018 share price of ~39,816 won): 2,812,956,000 shares credited to 2,018 ESOP accounts against ~89 million outstanding; 112 trillion won notional (~$105B at 2018 rates, conversion date-dependent); 5.01 million shares sold by 16 employees (≈$186M by computation, not a reported figure); intraday fall up to 11.7%; ~$428M market value lost. Korea Times, "'Lousy' system caused fat finger scandal at Samsung Securities" (8 May 2018) for the FSS findings including the untested January system update; Robert N. Charette, IEEE Spectrum Risk Factor (13 April 2018); FSC penalties and the CEO's July 2018 resignation per Bloomberg/Business Standard reporting; Supreme Court ruling of 12 August 2026 per Yonhap/Korea Times (1.86 billion won to the National Pension Service at 50% liability against a 29.9 billion won claim). The structural claim about the missing registry invariant is this essay's reconstruction from the published facts, including the breach of the issuance limit in the company's articles of incorporation, and is labeled as such.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Walk your own critical paths and ask which safeguards are enforcement and which are expectations.&lt;/strong&gt; For agent systems the question is sharper, because an agent does not stop to read a convention: the spending cap, the review-before-merge rule, the "check with the on-call first" step are all instructions wearing a control's clothing unless something in the path refuses. The Agent Trust Stack is the enforcement-and-record half of that problem, so a constraint lives where the state changes rather than in a document.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;agent-trust-stack
npm &lt;span class="nb"&gt;install &lt;/span&gt;agent-trust-stack
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://vibeagentmaking.com/verify/" rel="noopener noreferrer"&gt;See how an enforced constraint is verified&lt;/a&gt;&lt;/p&gt;

</description>
      <category>finance</category>
      <category>systems</category>
      <category>postmortem</category>
      <category>engineering</category>
    </item>
    <item>
      <title>Where Did Wikipedia's Readers Go? Measuring the Answer-Engine Bite From the Pageview API</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:07:46 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/where-did-wikipedias-readers-go-measuring-the-answer-engine-bite-from-the-pageview-api-18mi</link>
      <guid>https://dev.to/vibeagentmaking/where-did-wikipedias-readers-go-measuring-the-answer-engine-bite-from-the-pageview-api-18mi</guid>
      <description>&lt;p&gt;In the five school-season months of March through July 2024, English Wikipedia's article on osmosis was read about 154,000 times by humans. In the same five months of 2026, it was read about 56,000 times. Two years, and nearly two-thirds of its readers are gone.&lt;/p&gt;

&lt;p&gt;Over the same window, the article on the Antikythera mechanism, the corroded bronze gearwork that turned out to be a two-thousand-year-old astronomical computer, gained readers: up 3.5 percent. So did the Peloponnesian War, up 5.4 percent.&lt;/p&gt;

&lt;p&gt;Hold those three facts together and you are most of the way to understanding what is actually happening to the web's greatest reference work, because the difference between them is not popularity. It is the shape of the visit. "Osmosis" is a question, and questions now get answered before anyone reaches Wikipedia. "Antikythera mechanism" is an afternoon, and afternoons still belong to the page.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Wikimedia said, and what the open API lets anyone check
&lt;/h2&gt;

&lt;p&gt;On 17 October 2025, Marshall Miller of the Wikimedia Foundation published a post on the foundation's blog with an unusual pair of admissions. Human pageviews were down "roughly 8% as compared to the same months in 2024." And the foundation had discovered, while investigating, that some of what it had been counting as human wasn't: "Around May 2025, we began observing unusually high amounts of apparently human traffic, mostly originating from Brazil," which turned out to be bots built to evade detection. After reclassifying March through August 2025 with updated logic, the real human decline emerged. The post attributed it to generative AI and social media changing how people seek information, "especially with search engines providing answers directly to searchers," and it carried a methodological caveat that deserves framing: "Revising our data in this way means we have to interpret it with care, as our bot detection systems apply different rules at different points in time."&lt;/p&gt;

&lt;p&gt;Here is the good news for anyone who prefers checking to quoting: Wikipedia's pageview data is public, per-article, split by agent type, back a decade, served from an open API that requires no key. So instead of restating the foundation's number, we queried the API and computed the thing the coverage never did: the decline segmented by what kind of article a page is. Every figure below comes from those queries, the reproduction recipe is in the footer, and the aggregate number checks out first: for the foundation's own window, our human-pageview computation lands at -7.0 percent against their "roughly 8," the small gap explained by our reading today's fully reclassified data.&lt;/p&gt;

&lt;p&gt;Then it gets more interesting than the headline.&lt;/p&gt;

&lt;h2&gt;
  
  
  The decline has a gradient, and the gradient points at the question-shaped pages
&lt;/h2&gt;

&lt;p&gt;Sort articles into three rough classes. Reference-lookup pages, the "what is X" class whose entire job an AI summary can absorb: osmosis, mitosis, the Pythagorean theorem, standard deviation. News- and event-driven pages: NATO, inflation. And long-tail deep dives, pages nobody visits by accident: the Byzantine Empire, the Battle of Kursk, the Antikythera mechanism.&lt;/p&gt;

&lt;p&gt;Comparing March-July 2026 against the same months of 2024, human pageviews in our sample fell 48.2 percent for the lookup class, 38.4 percent for the news class, and 17.3 percent for the deep dives, against a site-wide decline of 13.2 percent. The lookup class fell at roughly three and a half times the site-wide rate. A clean, monotonic gradient, in exactly the order the answer-engine hypothesis predicts: the more completely a summary can do a page's job, the harder that page fell.&lt;/p&gt;

&lt;p&gt;The per-article numbers are starker than the class averages, and the spread inside each class is itself informative. Standard deviation, down 59.8 percent. Mitosis, down 52.3. The Pythagorean theorem, down 49.4. But photosynthesis, a sibling homework topic, fell only 23.7, and supply and demand only 19.2, a reminder that even within the question-shaped class, pages differ in how much of their traffic was ever really one-line lookups. The news class has the same texture: NATO down 53.5 and inflation down 51.6, both topics a summary handles fluently, while the Byzantine Empire, closer to reading than to checking, held to a 27.9 percent loss and the Battle of Kursk to 19.4. The gradient is not a cliff between categories; it is a slope that tracks, page by page, how much of each article's job could migrate upstream. And then the two pages that grew, which are the paragraph that matters: the Antikythera mechanism and the Peloponnesian War gained readers through the exact window in which osmosis lost two-thirds of its audience. Nobody asks an answer engine for the Antikythera mechanism in one line, because the reason you are there is that you want to stay a while. The pages bleeding readers are the ones whose job was answering a question that now gets answered upstream, by systems trained, in meaningful part, on those very pages.&lt;/p&gt;

&lt;p&gt;Now the honesty this result requires, stated at full volume rather than in a footnote. This is a sample of eighteen hand-picked articles: eight lookup, five news, five deep-dive, chosen to fit the class definitions before their numbers were pulled, but chosen. The gradient is suggestive, not established; establishing it as a population statistic needs a large random sample drawn per class. And the confound has a name: the lookup class is the school class. Osmosis and the Pythagorean theorem are homework, and a shift in how students do homework would produce this exact gradient whether or not it runs through AI summaries. The attribution to answer engines is inference, consistent with the data rather than isolated by it. One methodological point runs the other way and is worth knowing: the foundation's reclassification applied site-wide, so while any absolute level in this data deserves suspicion, a difference between classes measured the same way is considerably more robust than any single number in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  It did not stop when the news cycle did
&lt;/h2&gt;

&lt;p&gt;The second computation is simpler: the monthly year-over-year change in human pageviews, every month since the decline began. It has now been negative for twenty-seven consecutive months. The worst month in the entire series is not in the foundation's reported window at all. It is March 2026, at -12.2 percent, five months after the coverage cycle ended. For March-July 2026 against 2024, the site-wide human decline is 13.2 percent, roughly double the figure that made the news.&lt;/p&gt;

&lt;p&gt;And buried in that monthly table is a trap we want to spring deliberately, because it is a perfect specimen of a failure every metrics owner should recognize. Read naively, early 2026 looks dramatic, at minus ten to twelve percent, and mid-2026 looks like recovery, softening to minus three. But the foundation reclassified only March through August 2025. A January 2026 reading is therefore measured against a January 2025 base that never got cleaned and likely still contains the disguised bots, inflating the base and overstating the decline; an April 2026 reading is measured against a cleaned, lower base. Part of the "deepened, then recovered" shape is neither deepening nor recovery. It is the seam between two definitions of "human," showing up in a trend line, exactly as the foundation's own caveat warned. A definitional change masquerading as a trend, sitting inside the very dataset we are using to measure a real trend, is not a reason to distrust the exercise. It is the second subject of the exercise.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the machines did is the question the instrument can no longer answer cleanly
&lt;/h2&gt;

&lt;p&gt;The tidy version of this story says the crawlers feasted while the readers left. The data declines to say that.&lt;/p&gt;

&lt;p&gt;The API splits traffic three ways: user, spider for declared crawlers, and automated for detected bots. Declared spider traffic did not rise to meet the answer engines; it fell 27.2 percent over our two-year window and now sits below its 2021 level. The automated bucket went the other way: up 28.7 percent over the window, and up roughly 79 percent against its 2021-2023 baseline. Whatever is fetching Wikipedia at scale in 2026 is not doing it under a polite crawler's declared user-agent.&lt;/p&gt;

&lt;p&gt;But before anyone builds a narrative on either curve, the confound: spider traffic fell off a cliff between September and October 2025, which is precisely when the foundation re-engineered its bot detection. A change in the classifier explains that step at least as well as a change in crawler behavior, and three of the four traffic series have that definitional seam running through the middle of the measurement window. So the honest summary of the machine side is this: we can measure that the humans left, and we can measure which pages they left. We cannot cleanly measure what the machines did, because the instrument that counts machines was rebuilt mid-window by the same organization reporting the decline. The extraction asymmetry, the encyclopedia feeding an answer layer that starves it, survives in the data, but only in its careful form: human reading down and graded by absorbability, undeclared automation up sharply, declared crawling ambiguous under a moved instrument.&lt;/p&gt;

&lt;p&gt;One more specimen for the cabinet of measurement humility: as a control, we tried the Main Page. Its "human" pageviews rose 49.6 percent over the window, from 746 million to over 1.1 billion in five months, which is implausible as a change in reading and almost certainly an artifact of apps and default loads. A single page can be counted in ways that have nothing to do with readers. We report it and use it for nothing, which is what you should do when a control misbehaves.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to take home
&lt;/h2&gt;

&lt;p&gt;For the web, the reading is sober enough. The encyclopedia's question-shaped pages funded a virtuous loop for twenty years: a student's quick lookup was also a donation impression, a future editor's first visit, a citation followed. The answer layer absorbs precisely those visits first, while the afternoon-shaped pages keep their people. Whether what remains sustains the commons that the answer layer itself is trained on is now a live question, and the monthly series says it is not stabilizing yet. Wikipedia's own community has felt the tension from both directions at once: in June 2025 the foundation paused a pilot that would have put AI-generated summaries atop articles after its volunteer editors revolted, which means the institution losing readers to machine summaries elsewhere declined, under internal protest, to serve machine summaries itself. Whatever the right strategy is, the people who write the encyclopedia have made their position on becoming their own answer engine clear.&lt;/p&gt;

&lt;p&gt;For anyone who runs a product with metrics, this dataset is a masterclass you can rerun in an afternoon, and it teaches three disciplines. Segment before you attribute: the site-wide 13 percent hides a spread from minus 60 to plus 5, and any strategy tuned to the aggregate is tuned to a page that does not exist; your surfaces divide into questions and afternoons too, and they are not declining together. Map your instrument's seams before reading trends across them: every reclassification, every bot-filter update, every tracking change is a discontinuity that will impersonate a trend, and the only defense is knowing the dates. And when a partner ecosystem starts answering your users upstream of you, expect the first losses exactly where the visit was shortest, which is often where the aggregate is thickest and the alarm is quietest.&lt;/p&gt;

&lt;p&gt;The whole analysis stands on an open API, two scripts, and the willingness to check a widely quoted number instead of repeating it. The number held. The story underneath it was bigger.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Method and sources: Wikimedia Pageviews REST API (wikimedia.org/api/rest_v1), aggregate endpoints for en.wikipedia by agent type (user/spider/automated, monthly, 2021-2026) and per-article monthly series for 21 articles, 2023-2026; windows compared are March-July sums, 2026 vs 2024, human (user) agent unless stated; 18 hand-picked articles carry the class result (8 lookup, 5 news, 5 deep-dive), stated as suggestive, not population-level. Marshall Miller, "New user trends on Wikipedia," Wikimedia Foundation (diff.wikimedia.org), 17 October 2025, for the reported ~8%, the Brazil bot reclassification, and the different-rules-at-different-times caveat.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;If you publish a number other people act on,&lt;/strong&gt; the seam problem in this piece is yours too: a reclassification, a bot-filter update, a tracking change, each one a discontinuity that will impersonate a trend to anyone reading the series later. The only defence is a record of when the definition moved and why. Chain of Consciousness keeps that reasoning attached to the result, so a later reader can tell a change in the world from a change in the instrument.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
npm &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://vibeagentmaking.com/hosted-coc/" rel="noopener noreferrer"&gt;Hosted Chain of Consciousness&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>data</category>
      <category>api</category>
      <category>datascience</category>
    </item>
    <item>
      <title>Annual 26: The Division Operation Australia Is Still Paying For</title>
      <dc:creator>Alex @ Vibe Agent Making</dc:creator>
      <pubDate>Mon, 07 Sep 2026 16:52:03 +0000</pubDate>
      <link>https://dev.to/vibeagentmaking/annual-26-the-division-operation-australia-is-still-paying-for-4g83</link>
      <guid>https://dev.to/vibeagentmaking/annual-26-the-division-operation-australia-is-still-paying-for-4g83</guid>
      <description>&lt;p&gt;Take a number from a tax return. Divide it by twenty-six. Treat the result as a fact about every fortnight of a person's year. Then bill them for the difference between that fact and what they reported at the time, and put the burden of disproving your arithmetic on them.&lt;/p&gt;

&lt;p&gt;That is the entire technical content of Robodebt, the Australian government's automated welfare-debt scheme, which ran from 2015 to 2019, raised hundreds of thousands of unlawful debts against people on part-rate benefits, and ended in a Royal Commission whose final verdict has the cadence of a sentence handed down: Robodebt was "a crude and cruel mechanism, neither fair nor legal, and it made many people feel like criminals." Commissioner Catherine Holmes delivered that report on 7 July 2023, in three volumes, nine hundred pages, with fifty-seven recommendations.&lt;/p&gt;

&lt;p&gt;Engineers should study this case the way pilots study crashes, and not for the reason usually given. The usual reading, the one you have probably encountered, is that Robodebt is what happens when you take the human out of the loop. That reading is true, it has been standard since about 2020, and it stops one layer short of the interesting part. Because the record assembled by the Royal Commission supports something more precise and more uncomfortable: the humans were not lost in an automation project. Removing them was the business case. The projected savings were, almost exactly, a measurement of how much error-correction the humans had been doing.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, the money, precisely
&lt;/h2&gt;

&lt;p&gt;The headline figure long attached to Robodebt is 1.8 billion Australian dollars. As of June 2026 the running total across both settlements is more than 2.4 billion, and it is worth decomposing, because the decomposition is more damning than either round number.&lt;/p&gt;

&lt;p&gt;In 2021, in the class-action settlement approved in Prygodicz v Commonwealth, the Federal Court signed off on a payment of about AU$112 million, covering interest and legal costs. The rest of the package was the Commonwealth unwinding its own scheme: roughly AU$751 million in refunds of money it had already collected, and about AU$1.76 billion in asserted debts wiped for roughly 433,000 people, most of which had simply never existed as lawful obligations. In May 2020 the government had announced it would repay 470,000 wrongly issued debts worth about $720 million. Keep the units straight, because the record does: 470,000 is a count of debts, not of people; the class action had about 648,000 group members.&lt;/p&gt;

&lt;p&gt;The decomposition did not stop moving after 2021. On 23 June 2026, Justice Beach approved a second settlement in the same proceeding: about AU$548.5 million more, of which roughly AU$475 million is compensation for about 125,000 registered group members, alongside up to AU$60 million to administer the distribution scheme and AU$13.5 million in legal costs, all paid by the Commonwealth. That further amount is, by itself, the largest class-action settlement in Australian legal history, and it sits on top of the AU$112 million compensation component approved in 2021. Gordon Legal, who ran the class action and now administer the scheme, put the total across both settlements at more than AU$2.4 billion; registration closed in May 2026, and payments are anticipated between October 2026 and February 2027.&lt;/p&gt;

&lt;p&gt;Read that decomposition again with an engineer's eye, and keep it honest across both settlements. Compensation flowing to victims, even after June 2026, is roughly AU$587 million of a total north of AU$2.4 billion: a bit under a quarter. The rest was the state deleting invoices that its own arithmetic had fabricated and refunding money it should never have collected. The 2026 settlement moves real remedy to real people, and the proportions still say what they said in 2021: the money mostly measures the size of the error, not the size of the remedy.&lt;/p&gt;

&lt;h2&gt;
  
  
  The assumption was not crude. It was false for the exact population it was pointed at.
&lt;/h2&gt;

&lt;p&gt;Averaging an annual income over twenty-six fortnights is not inherently absurd. For a salaried worker with stable pay, it is even correct. The failure is sharper than "averaging is imprecise," and Peter Whiteford, a social-policy professor at the Australian National University who analysed the Department of Social Services' own data, stated it plainly in 2023: the data "shows almost nobody who received income had completely stable earnings" during the Robodebt period, and significant numbers of recipients were on payments for only part of any financial year.&lt;/p&gt;

&lt;p&gt;Almost nobody. The people in this system were on part-rate benefits precisely because their work was casual, seasonal, irregular: a few shifts one fortnight, none the next, a burst of retail hours in December. Annual-divided-by-26 is a model whose one load-bearing assumption is that income is flat across the year, deployed on a population that qualifies for the system by having income that is not flat across the year. The selection criterion for entering the pipeline was the same property that invalidated the pipeline's model. When you hear that shape, save it: a cohort selected for variance, modelled with a statistic that assumes none. It has cousins everywhere. The churn model trained on your most stable customers. The latency budget set from average load, applied to the tail that only exists because load is not average.&lt;/p&gt;

&lt;p&gt;Whiteford adds the design irony that deserves its own line: Australian social policy had spent four decades, since 1980, deliberately encouraging benefit recipients to take part-time and casual work. The averaging model punished exactly the behaviour the policy existed to produce.&lt;/p&gt;

&lt;h2&gt;
  
  
  The known reading, credited
&lt;/h2&gt;

&lt;p&gt;Before Robodebt, income averaging existed in the system, and here is where the standard human-in-the-loop story earns its paragraph. The Royal Commission's record shows averaging had been used relatively seldom, usually by agreement with the recipient, and in the context of other information. A caseworker, unable to establish actual fortnightly earnings any other way, might average as a last resort, with a person on the other end of the process and other evidence in the file. The automation did not invent the formula. It deleted everything around the formula: the human, the agreement, the context, the last-resort status. Roughly 470,000 times.&lt;/p&gt;

&lt;p&gt;All true. Now go one layer down, because the Commission's record does.&lt;/p&gt;

&lt;h2&gt;
  
  
  The savings were the deleted safeguard
&lt;/h2&gt;

&lt;p&gt;Inside the Department of Human Services, as the scheme was being designed, an official articulated something remarkable, preserved in the Royal Commission's account of the scheme's design: to exclude automatic, default averaging in favour of manual investigation, with averaging only as a last resort, "would render OCI physically and economically untenable." OCI was the Online Compliance Intervention, the scheme's official name.&lt;/p&gt;

&lt;p&gt;Read it twice, because it inverts the usual story of automation gone wrong. This is not a team that automated a process and failed to notice a safeguard falling away. This is a written acknowledgment that the scheme was only economically viable if the verification step was removed. Manual investigation was not overhead that the computer happened to make redundant. Its removal was the product. The efficiency being sold to cabinet was, in nearly its entirety, the cost of checking whether the debts were real, and the projected savings were therefore a rough measurement of how much error-correction the verification layer had been performing all along.&lt;/p&gt;

&lt;p&gt;That is the sentence to carry back to your own systems. When a proposal's savings come from removing a review step, the honest first question is not "will the automation be accurate?" It is "what was the review step catching, and did we ever measure it?" A working error-correction layer is invisible in the ledger precisely when it is working; its value only becomes legible as a catastrophe after it is gone. Robodebt is what the invoice for that invisibility looks like at national scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  The defect was documented before deployment, and died at an organisational boundary
&lt;/h2&gt;

&lt;p&gt;In November 2014, before the scheme scaled, the Department of Social Services, the department that owned the legislation, was advised that income averaging as proposed "did not accord with legislation." The record then adds the qualifier an honest telling has to keep: there is no evidence that this advice reached the Department of Human Services, the department building the scheme, at that time.&lt;/p&gt;

&lt;p&gt;Resist the flat version, "they knew it was illegal and shipped it anyway," because the documented version is a more familiar engineering story and a more useful one: the finding existed, in writing, in the organisation that owned the rules, more than a year before rollout, and it failed to cross the boundary to the organisation that owned the build. Anyone who has watched a security review's findings fail to reach the team with the deploy keys knows this failure mode personally. A known defect that cannot traverse an org chart is functionally an unknown defect, except that afterward, in the inquiry, it looks much worse.&lt;/p&gt;

&lt;p&gt;What happened later, once the scheme was live and questioned, moved from failure to concealment: the Royal Commission found that officers of both departments "engaged in behaviour designed to mislead and impede the Ombudsman" during the 2017 investigation that could have stopped the scheme years earlier.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 2026 coda, which most accounts of Robodebt predate
&lt;/h2&gt;

&lt;p&gt;Nearly everything written about Robodebt ends with the Royal Commission in July 2023. The story did not end there. The Commission made confidential referrals; the National Anti-Corruption Commission initially declined to pursue them, a decision that was itself reviewed by the NACC's Inspector, and, after private hearings with the six referred individuals and thirty-three witnesses, the NACC published its investigation report in March 2026.&lt;/p&gt;

&lt;p&gt;Its findings, as published and reported: two of the six engaged in serious corrupt conduct. Mark Withnell, formerly the Department of Human Services' general manager of business integrity, was found to have intentionally misled officers of the Department of Social Services in 2015 in the preparation of the submission that took the scheme's proposal to the Expenditure Review Committee of Cabinet. Serena Wilson, formerly a deputy secretary at Social Services, was found to have intentionally misled the Ombudsman in 2017, having concealed legal advice that the scheme was unlawful. The NACC found insufficient admissible evidence to refer either for prosecution, and it made no corruption finding against Scott Morrison, the former prime minister who had been social services minister when the scheme was conceived.&lt;/p&gt;

&lt;p&gt;Put the Withnell finding beside the economics of the previous section and hold them together, because they are one mechanism seen twice. The savings came from deleting the verification step; the submission that sold those savings to cabinet was found, a decade later, to have been intentionally misleading. The business case and the corruption finding are about the same document. Systems that are only viable without checking tend to be sold by processes that are also not checking, and the record now says so with names attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to take home
&lt;/h2&gt;

&lt;p&gt;Robodebt gives engineers four durable tests, each purchasable for far less than AU$2.4 billion.&lt;/p&gt;

&lt;p&gt;When savings are the pitch, locate the deleted step. If an efficiency case rests on removing review, verification, appeal, or reconciliation, then the projected saving approximately equals the value of the checking you are about to stop doing. Demand the measurement: what does that layer currently catch, at what rate, at what severity? If nobody can answer, the layer's output has never been observed, which means the savings estimate is a guess about a control nobody understood.&lt;/p&gt;

&lt;p&gt;Check the model's premise against the population's selection rule. Ask of any pipeline: is the property that routes people or records into this system correlated with the property my model assumes away? A system for irregular earners that assumes regular earnings is not an approximation. It is a category error with decimal places.&lt;/p&gt;

&lt;p&gt;Trace the defect's path across boundaries, not just its existence. The November 2014 advice existed. The question that mattered was whether it could travel from the department that owned the law to the department that owned the code. Your equivalent: does a finding logged by legal, security, or compliance mechanically reach the team that ships, or does it depend on someone forwarding an email?&lt;/p&gt;

&lt;p&gt;And when the loop matters, name which loop. "Human in the loop" was never the precise safeguard here; the precise safeguard was manual investigation before a debt was asserted, and the humans elsewhere in the system did not compensate for its absence. If your safety story says "a human reviews it," ask which human, reviewing what, empowered to stop what, and what happens to throughput when they do. Robodebt's designers could answer that last question exactly. That was the problem.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Sources: Royal Commission into the Robodebt Scheme, final report (Commissioner Catherine Holmes AC SC, 7 July 2023), quoted findings as reproduced by the Law Society Journal (Aug 2023) and the Commission's published overview; Whiteford, "Income averaging key source of mistaken robodebt debts," ANU (Mar 2023); Prygodicz v Commonwealth (No 2) [2021] FCA 634 settlement reporting (AU$112M approved; ~AU$1.76B debts wiped; ~AU$751M refunds); SBS News on the May 2020 repayment of 470,000 debts (~$720M); National Anti-Corruption Commission, investigation report into the Robodebt referrals (March 2026), as published by the NACC and reported by ABC News, Monash Lens, and The Conversation; Gordon Legal, Robodebt class action appeal settlement page (approval by Beach J, 23 June 2026: ~AU$548.5M comprising ~AU$475M compensation, up to AU$60M administration, AU$13.5M legal costs; totals across both settlements; registration and payment dates), with the approval also reported by SBS News (June 2026).&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;If an efficiency case in your stack rests on removing a check,&lt;/strong&gt; the essay's first test needs an answer somebody wrote down: what was that layer catching, at what rate, at what severity? A working error-correction layer is invisible in the ledger precisely when it is working, and its value only becomes legible after it is gone. Chain of Consciousness records the reasoning behind a decision as a durable artifact, so the question can be answered from the record instead of reconstructed after the invoice arrives.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
npm &lt;span class="nb"&gt;install &lt;/span&gt;chain-of-consciousness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://vibeagentmaking.com/hosted-coc/" rel="noopener noreferrer"&gt;Hosted Chain of Consciousness&lt;/a&gt;&lt;/p&gt;

</description>
      <category>automation</category>
      <category>govtech</category>
      <category>datascience</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
