<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: DEVALAND</title>
    <description>The latest articles on DEV Community by DEVALAND (@devaland).</description>
    <link>https://dev.to/devaland</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2584621%2Fee84c450-584a-4835-a63c-57faad2045f4.jpg</url>
      <title>DEV Community: DEVALAND</title>
      <link>https://dev.to/devaland</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/devaland"/>
    <language>en</language>
    <item>
      <title>Actively exploited this week: 6 new CISA KEV entries (28 September to 4 October 2026)</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Mon, 05 Oct 2026 07:33:07 +0000</pubDate>
      <link>https://dev.to/devaland/actively-exploited-this-week-6-new-cisa-kev-entries-28-september-to-4-october-2026-27gn</link>
      <guid>https://dev.to/devaland/actively-exploited-this-week-6-new-cisa-kev-entries-28-september-to-4-october-2026-27gn</guid>
      <description>&lt;p&gt;Every week CISA adds vulnerabilities to its &lt;strong&gt;Known Exploited Vulnerabilities (KEV)&lt;/strong&gt; catalogue. The bar for getting on that list is not "severe on paper" but "attackers are using it now". For a small team with limited patching hours, that makes KEV the most useful priority list there is.&lt;/p&gt;

&lt;p&gt;Below: every entry added between &lt;strong&gt;28 September and 4 October 2026&lt;/strong&gt;, copied from the catalogue and linked to its NVD record, plus what Romania's national cyber security directorate and Have I Been Pwned published in the same days. Nothing is estimated or rewritten.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 6 actively exploited vulnerabilities added this week
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Added&lt;/th&gt;
&lt;th&gt;Product&lt;/th&gt;
&lt;th&gt;CVE&lt;/th&gt;
&lt;th&gt;Known ransomware use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;4 Oct&lt;/td&gt;
&lt;td&gt;Citrix NetScaler Improper Restriction of Operations within the Bounds of a Memory Buffer Vulnerability&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-88779" rel="noopener noreferrer"&gt;CVE-2026-88779&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;not known&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 Oct&lt;/td&gt;
&lt;td&gt;Zammad GmbH Zammad Improper Privilege Management Vulnerability&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-102490" rel="noopener noreferrer"&gt;CVE-2026-102490&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;not known&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 Oct&lt;/td&gt;
&lt;td&gt;Zammad GmbH Zammad Session Fixation Vulnerability&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-102489" rel="noopener noreferrer"&gt;CVE-2026-102489&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;not known&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1 Oct&lt;/td&gt;
&lt;td&gt;Fortinet FortiMail Path Traversal Vulnerability&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-104286" rel="noopener noreferrer"&gt;CVE-2026-104286&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;not known&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;30 Sep&lt;/td&gt;
&lt;td&gt;Cisco Catalyst SD-WAN Manager Hex Encoding Vulnerability&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-76504" rel="noopener noreferrer"&gt;CVE-2026-76504&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;not known&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;29 Sep&lt;/td&gt;
&lt;td&gt;Apple Multiple Products Out-of-Bounds Write Vulnerability&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-86950" rel="noopener noreferrer"&gt;CVE-2026-86950&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;not known&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What Romania's DNSC flagged the same week
&lt;/h2&gt;

&lt;p&gt;Alerts from &lt;strong&gt;DNSC&lt;/strong&gt;, Romania's national cyber security directorate, with their original titles in Romanian:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dnsc.ro/citeste/alerta-vulnerabilitate-critica-la-nivelul-fortinet-fortimail" rel="noopener noreferrer"&gt;ALERTĂ: Vulnerabilitate critică la nivelul Fortinet FortiMail&lt;/a&gt;, 2 October&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dnsc.ro/citeste/alert-vulnerabilitati-critice-exploatate-activ-in-citrix-netscaler" rel="noopener noreferrer"&gt;ALERTĂ: Vulnerabilități critice exploatate activ în Citrix NetScaler&lt;/a&gt;, 28 September&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When a national authority and CISA point at the same product in the same week, that product goes to the top of the queue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Breaches added to Have I Been Pwned
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Medela&lt;/strong&gt;: 423,947 accounts, added 30 September. &lt;a href="https://haveibeenpwned.com/Breach/Medela" rel="noopener noreferrer"&gt;Check the breach page&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your team's work addresses may be in one of these, check the breach page and the password reuse question, not only the address.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to use a list like this in an hour
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Start with what faces the internet.&lt;/strong&gt; VPN and access gateways, firewalls, routers, mail servers and public web platforms are how many ransomware incidents start.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Search your inventory by product name, not by memory.&lt;/strong&gt; "We don't run that" is a claim, and the asset list is the evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Patch or apply the vendor mitigation, then confirm the version.&lt;/strong&gt; A patch that was downloaded but not applied reads as done in a ticket and is still exploitable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you cannot patch today, reduce exposure:&lt;/strong&gt; restrict the management interface to known addresses, disable the affected feature, or put the service behind an access gateway you trust.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Where this comes from
&lt;/h2&gt;

&lt;p&gt;A small collector reads official sources every hour: DNSC, CERT-EU, CERT-FR (ANSSI), CERT-Bund (BSI), the UK NCSC, CISA KEV, Have I Been Pwned and three specialist newsrooms. It copies each item exactly, with its source link, and publishes the result on one page:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://devaland.cloud/alerte-live.html" rel="noopener noreferrer"&gt;https://devaland.cloud/alerte-live.html&lt;/a&gt;&lt;/strong&gt; (in Romanian, with every title kept in its original language).&lt;/p&gt;

&lt;p&gt;Nothing on that page or in this post is written by a model. A security page that paraphrases an advisory wrongly does more harm than no page at all.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sources: CISA KEV catalogue (dateAdded 2026-09-28 to 2026-10-04), DNSC alerts, Have I Been Pwned. Generated from the sources on 2026-10-05.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>cybersecurity</category>
      <category>devops</category>
      <category>infosec</category>
    </item>
    <item>
      <title>Four Controls That Turn an AI Agent From a Demo Into a System</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Sun, 04 Oct 2026 13:01:27 +0000</pubDate>
      <link>https://dev.to/devaland/four-controls-that-turn-an-ai-agent-from-a-demo-into-a-system-3pn3</link>
      <guid>https://dev.to/devaland/four-controls-that-turn-an-ai-agent-from-a-demo-into-a-system-3pn3</guid>
      <description>&lt;p&gt;A chatbot answers. An agent acts: it searches a document, fills in a form, drafts the reply and sends it. Because it acts, it needs rules before it needs features.&lt;/p&gt;

&lt;p&gt;Two guides on AI agents published by Sifted in 2026, one sponsored by Box and one by Salesforce, collect the experience of companies that got past the pilot stage. Strip out the product names and the same four controls appear in both. Each one is a question you can put to any vendor, including us.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Scoped access, time-bounded, fully logged
&lt;/h2&gt;

&lt;p&gt;The clearest version comes from Box's chief information security officer, Heather Ceylan: agents should never have broad standing access just because they are useful. Permissions should be scoped to a task, time-bounded where possible, and fully logged.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The question to ask:&lt;/strong&gt; what exactly can the agent read and change, and who can see what it touched? If the answer is "everything", that is not an answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. A second check
&lt;/h2&gt;

&lt;p&gt;Agents make mistakes, and the guides say so plainly. One founder describes them as "well-intentioned, slightly forgetful children". The teams that got past the pilot all added a second step. Either a person approves before anything leaves the building, or a separate evaluator checks the first agent's work independently and scores its confidence.&lt;/p&gt;

&lt;p&gt;One London company in the Box report, Deliverance, builds this in as a pair of agents: an executor that does the work and an evaluator that critiques it. Its founder's summary is worth borrowing: "It's not magic. It's a governed system."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The question to ask:&lt;/strong&gt; what checks the output before a customer sees it?&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Rules for when it does not know
&lt;/h2&gt;

&lt;p&gt;A reliable agent knows when not to act. That matters more than being right most of the time, because an agent that recognises its limits does far less damage than one that is merely usually correct. The rules are written in advance: if it cannot find the source, it says so; if confidence is low, it asks a human; if money or an upset customer is involved, it escalates.&lt;/p&gt;

&lt;p&gt;Then test exactly those cases. The tidy requests almost always work. The oddly phrased ones are where the failures are.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The question to ask:&lt;/strong&gt; show me what it does with a request it cannot answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. A log of every step
&lt;/h2&gt;

&lt;p&gt;For each run: what the agent received, what it used, what it produced. When something goes wrong, the log tells you whether the input was bad, the instructions were unclear or a tool failed, instead of guessing. It is also what lets you answer, months later, where a particular answer came from.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The question to ask:&lt;/strong&gt; if I pick any answer from last week, can you show me its sources?&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means in practice
&lt;/h2&gt;

&lt;p&gt;These are the rules we build to at Devaland. Our assistants answer from the client's own documents, cite the source behind every claim, and say plainly when they cannot find one. We do not present that as a guarantee. We present it as something you can verify yourself, on every single answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;Sifted, "The rise of AI agents", June 2026 (sponsored by Box). Sifted, "The startup agentic AI playbook", August 2026 (sponsored by Salesforce).&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on devaland.com: &lt;a href="https://devaland.com/blog/governed-ai-agents-four-controls" rel="noopener noreferrer"&gt;Four Controls That Turn an AI Agent From a Demo Into a System&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>security</category>
      <category>llm</category>
    </item>
    <item>
      <title>I told a friend a number was not in her document. It was on page 6, sideways.</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Fri, 02 Oct 2026 06:19:43 +0000</pubDate>
      <link>https://dev.to/devaland/i-told-a-friend-a-number-was-not-in-her-document-it-was-on-page-6-sideways-5280</link>
      <guid>https://dev.to/devaland/i-told-a-friend-a-number-was-not-in-her-document-it-was-on-page-6-sideways-5280</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;everypage&lt;/strong&gt; is a small tool that answers questions about a PDF using a model that runs on your own machine. It names the page for everything it says, and when it could not read a page it tells you that instead of telling you "no".&lt;/p&gt;

&lt;p&gt;I built it for a friend. She is not technical. She has a folder of family property papers, some printed from a computer, some scanned, some scanned lying on their side. She asks me the kind of question anyone asks about a contract: does it say this anywhere?&lt;/p&gt;

&lt;p&gt;A while ago she asked me one of those, and I answered "no, that figure is not in the document". I had searched the text. The text I searched ended at page 4. The figure was on page 6, on a scanned page that my extraction had quietly skipped. Nothing warned me. A document with two missing pages looks exactly like a complete document that does not contain the thing.&lt;/p&gt;

&lt;p&gt;She trusted my "no". That is the bug I wanted to fix, and it is not a model bug. It is a reading bug.&lt;/p&gt;

&lt;p&gt;So the tool has three answers, never two:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;verdict&lt;/th&gt;
&lt;th&gt;what it means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;FOUND&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;here is the sentence, and here is the page it is on&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;NOT FOUND&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;every page was read, and it is not there&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CANNOT SAY&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;it was not on the pages I could read, but there are pages I could not read&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fesmwlark6ncooww8sfv4.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fesmwlark6ncooww8sfv4.gif" alt="Terminal recording: the naive run answers wrongly, everypage finds the clause on page 6, then says CANNOT SAY when page 6 is unreadable" width="600" height="338"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A real run of the five commands below, recorded on a CPU with &lt;code&gt;gemma2:2b&lt;/code&gt;. Only the waits for local inference are cut, and each cut is labelled on screen. &lt;a href="https://github.com/MariusGithub13/everypage/blob/main/docs/everypage-demo.mp4" rel="noopener noreferrer"&gt;MP4 version&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The sample is an invented 6-page sale agreement (every name, place and amount is made up). Pages 1 to 3 have a text layer. Pages 4 and 5 are scans. Page 6 is a scan lying on its side, and it is the only page that mentions unpaid taxes.&lt;/p&gt;

&lt;p&gt;First, the usual way: extract the text, give it to the model, ask.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ python3 demo/naive.py samples/sale-agreement.pdf "Are there unpaid property taxes, how much, and who has to pay them?"
NAIVE: pdftotext returned 1449 characters and no warning
NAIVE ANSWER: The provided document does not contain information about unpaid property taxes,
their amount, or who is responsible for paying them.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That is a confident, fluent, wrong answer, and the model did nothing wrong. It was handed half a document.&lt;/p&gt;

&lt;p&gt;Now the same model, the same question, through everypage:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ python3 -m everypage read samples/sale-agreement.pdf
COVERAGE: read 6 of 6 pages
  page 1   text-layer   591 chars  legibility 0.54  OK
  page 2   text-layer   446 chars  legibility 0.56  OK
  page 3   text-layer   408 chars  legibility 0.67  OK
  page 4   ocr          384 chars  legibility 0.48  OK
  page 5   ocr          225 chars  legibility 0.41  OK
  page 6   ocr          459 chars  legibility 0.68  rotated 90°  OK

$ python3 -m everypage ask samples/sale-agreement.pdf "Are there unpaid property taxes, how much, and who has to pay them?"
COVERAGE: read 6 of 6 pages
ANSWER: Unpaid property taxes total $18,450. (page 6) The seller must pay this amount before closing. (page 6)
If the seller does not pay the unpaid taxes by closing, the buyer may deduct $18,450 from the price and pay
the county directly. (page 6)
  page 6 (ocr): "Property taxes for the years 2022, 2023 and 2024 remain unpaid in the total amount of $18,450."
  page 6 (ocr): "The Seller shall pay this amount in full before closing."
  page 6 (ocr): "If the Seller has not paid the unpaid taxes by closing, the Buyer may deduct $18,450 from the price and pay the county directly."
VERDICT: FOUND.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;A question whose honest answer is no:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; everypage ask samples/sale-agreement.pdf &lt;span class="s2"&gt;"Does the agreement say anything about a broker commission?"&lt;/span&gt;
&lt;span class="go"&gt;COVERAGE: read 6 of 6 pages
VERDICT: NOT FOUND. All 6 of 6 pages were read, so this is a real negative.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;And the case that started all this. Same agreement, but page 6 is blurred past reading:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; everypage ask samples/sale-agreement-damaged.pdf &lt;span class="s2"&gt;"Are there unpaid property taxes, how much, and who has to pay them?"&lt;/span&gt;
COVERAGE: &lt;span class="nb"&gt;read &lt;/span&gt;5 of 6 pages&lt;span class="p"&gt;;&lt;/span&gt; could NOT &lt;span class="nb"&gt;read &lt;/span&gt;page&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt; 6
VERDICT: CANNOT SAY. Nothing found on the pages that were &lt;span class="nb"&gt;read&lt;/span&gt;, but page&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt; 6 could not be read. This is NOT a &lt;span class="s1"&gt;'no'&lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That last line is the whole project. It is the sentence I should have said to her.&lt;/p&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/MariusGithub13" rel="noopener noreferrer"&gt;
        MariusGithub13
      &lt;/a&gt; / &lt;a href="https://github.com/MariusGithub13/everypage" rel="noopener noreferrer"&gt;
        everypage
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Ask a PDF a question with a local model. It names the page, and says which pages it could not read.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;everypage&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;Ask a question about a PDF and get an answer that names its page, from a model that runs on your own machine
If a page could not be read, it says so instead of saying "no".&lt;/p&gt;
&lt;p&gt;It was built for one person: a friend with a folder of property papers, part printed, part scanned, some scanned
sideways. Her question is usually "does it say X anywhere?". The honest answers to that are three, not two.&lt;/p&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;verdict&lt;/th&gt;
&lt;th&gt;meaning&lt;/th&gt;
&lt;th&gt;exit code&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;FOUND&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;here is the sentence, here is the page&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;NOT FOUND&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;every page was read, and it is not there&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;CANNOT SAY&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;it was not on the pages I could read, but some pages I could not read&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/MariusGithub13/everypage/docs/everypage-demo.gif"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FMariusGithub13%2Feverypage%2FHEAD%2Fdocs%2Feverypage-demo.gif" alt="Terminal recording of the demo: the naive run answers wrongly, everypage finds the clause on page 6, and says CANNOT SAY when page 6 is unreadable"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A real run on the invented sample, recorded 02.10.2026. Only the waits for local inference are cut, and each cut is labelled on screen. &lt;a href="https://github.com/MariusGithub13/everypage/docs/everypage-demo.mp4" rel="noopener noreferrer"&gt;MP4 version&lt;/a&gt;.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;What&lt;/h2&gt;…&lt;/div&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/MariusGithub13/everypage" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;&lt;a href="https://github.com/MariusGithub13/everypage" rel="noopener noreferrer"&gt;https://github.com/MariusGithub13/everypage&lt;/a&gt; (MIT). About 300 lines of Python, no framework. The tests run without a model.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The model is Gemma, running locally.&lt;/strong&gt; &lt;code&gt;gemma2:2b&lt;/code&gt; (1.6 GB) served by Ollama on localhost. OCR is Tesseract, PDF handling is Poppler. All open, all on the machine.&lt;/p&gt;

&lt;p&gt;A 2-billion-parameter model is small, and that shaped the design. I did not ask it to be reliable. I asked it to do one easy thing, and I made the code responsible for everything that has to be true.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The code reads, and keeps the receipts.&lt;/strong&gt; Each page is read on its own. If it has a text layer, that is used. If not, the page is rendered at 300 dpi and OCR'd. If the result does not look like language, it is retried at 90, 270 and 180 degrees and the best reading wins. Every page ends as OK or UNREADABLE, and the coverage line is printed before anything else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The model sees every page, one at a time.&lt;/strong&gt; There is no retrieval step choosing which chunks are "relevant". On a short legal document, the page a retriever skips is the addendum. It is slower. With this small model on a single CPU core it is roughly 15 to 20 seconds a page, so about two minutes for six pages, and my friend's question is worth two minutes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The model must quote, and the code checks the quote.&lt;/strong&gt; For each page the model returns up to three sentences copied from that page. The code looks for each one on the page, character for character after normalising spaces and case. A sentence that is not there is thrown away, whatever the model says about it. The final answer is written only from sentences that survived, each with its page.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The model is never allowed to say "it is not there".&lt;/strong&gt; If nothing survived, the code decides: all pages read means NOT FOUND, anything less means CANNOT SAY. The exit codes are 0, 1 and 2, so a script cannot confuse them either.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The bug I shipped to myself on the way.&lt;/strong&gt; My first legibility check counted words that contain a vowel. I ran it on the sideways page and it reported legibility 0.96 and status OK. The "text" it had approved was this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;pebueyoun urewel JuoweeIby oY} Jo SULI9} 1940 [[V
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the page read upside down. Upside-down English is full of vowels. The blurred page scored a perfect 1.00 with &lt;code&gt;ee ee mee ON me mere me&lt;/code&gt;. So the check that existed to catch unreadable pages was passing them with top marks, which is the original bug wearing a badge. The fix was to stop asking "does this look like words" and ask "does this contain the small everyday words real prose is made of": the, of, and, shall. Upside-down text scores 0.03 on that, real pages score 0.4 to 0.7, and both failures are now tests.&lt;/p&gt;

&lt;p&gt;I used AI coding agents to help write and test this. The design rules and the failure they come from are mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;Three reasons, all practical.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The papers never leave the machine.&lt;/strong&gt; Property papers, family papers and medical papers are the documents people most need help with and least want to upload to a server they do not control. With an open-weight model on localhost that question does not come up. It works with the network cable out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I can put the rules outside the model.&lt;/strong&gt; Because everything runs in my own process, the model's output is just a string my code can check against the page before anyone sees it. The guarantee does not depend on a provider's settings, a prompt that might be ignored, or a model version that changes under me next month.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It costs nothing to run, so "read every page" is affordable.&lt;/strong&gt; Showing every page to a model one at a time is wasteful by API standards. Locally the only cost is a couple of minutes of CPU, so I can choose thorough over clever.&lt;/p&gt;

&lt;p&gt;A closed model would have given a more polished paragraph. It would not have fixed my bug, because my bug was never the paragraph.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits, and one thing I will not do
&lt;/h2&gt;

&lt;p&gt;She should not have to install anything, so I run it on her papers at my end and send her the answer with the page numbers. What she says about it stays between us. I am not going to quote her here.&lt;/p&gt;

&lt;p&gt;So nobody is surprised: PDFs only. The legibility word list is English, so other languages need their own list (pages get marked unreadable otherwise, which is the safe way to be wrong). A page that is all numbers gets marked unreadable for the same reason. And a verified quote proves the words are on the page, not that the model understood them. Read the quote. This is a reading aid, not legal advice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;Best Use of Gemma: &lt;code&gt;gemma2:2b&lt;/code&gt; runs locally and is the only model in the project.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
      <category>hacktoberfest</category>
    </item>
    <item>
      <title>My Staleness Checker Flagged Sixteen Tasks. I Acted on Four of Them Anyway.</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:00:06 +0000</pubDate>
      <link>https://dev.to/devaland/my-staleness-checker-flagged-sixteen-tasks-i-acted-on-four-of-them-anyway-5268</link>
      <guid>https://dev.to/devaland/my-staleness-checker-flagged-sixteen-tasks-i-acted-on-four-of-them-anyway-5268</guid>
      <description>&lt;p&gt;On 20 July I served myself two dead tasks as if they were live. One told me to chase a contact at a&lt;br&gt;
deal that had closed months earlier. The other told me to stay quiet and wait for a call that had&lt;br&gt;
already happened, and gone nowhere. Both came out of the same file: a running list of open threads&lt;br&gt;
that my operations run against.&lt;/p&gt;

&lt;p&gt;So that afternoon I wrote a checker. Every open line in that file has to carry a tag naming the&lt;br&gt;
project note it comes from and the date it was last reconciled. The script compares that date&lt;br&gt;
against the modification time of the note itself. If the note changed after the line was last&lt;br&gt;
checked, the line is stale: the decision moved on and the reminder did not. It also catches lines&lt;br&gt;
with no tag at all, and lines pointing at a note that no longer exists.&lt;/p&gt;

&lt;p&gt;It works. I ran it this morning. Fifty-eight open lines, sixteen flagged.&lt;/p&gt;

&lt;p&gt;Then I acted on one of the flagged lines anyway. And then three more.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check was never the problem
&lt;/h2&gt;

&lt;p&gt;Read the script and there is nothing to fix. It parses correctly. It classifies into three states&lt;br&gt;
that are genuinely different from each other. It prints the offending line with its note and both&lt;br&gt;
dates, so you can see exactly why it fired.&lt;/p&gt;

&lt;p&gt;It even exits with status 1 when anything is unreconciled, and 0 when the file is clean. That is the&lt;br&gt;
oldest convention in computing for "do not proceed."&lt;/p&gt;

&lt;p&gt;Nothing reads that exit code.&lt;/p&gt;

&lt;p&gt;I went looking today. Nothing calls the script. No scheduled job, no wrapper, no step that runs it&lt;br&gt;
and stops if it fails. The only two places on the machine that mention it at all are the script&lt;br&gt;
itself and a paragraph inside the very file it polices, which says that it must be run first and&lt;br&gt;
that flagged lines must not be surfaced as actions.&lt;/p&gt;

&lt;p&gt;So the enforcement mechanism was a sentence, sitting next to the thing it governed, asking whoever&lt;br&gt;
came past to behave.&lt;/p&gt;

&lt;h2&gt;
  
  
  What that cost, in one morning
&lt;/h2&gt;

&lt;p&gt;I run my operations through an AI agent with a persistent memory of the business. It read that&lt;br&gt;
paragraph. It ran the checker. It received sixteen flags. It then surfaced four of them to me as&lt;br&gt;
work to do.&lt;/p&gt;

&lt;p&gt;The first was a supplier setting I had changed the previous evening and confirmed on screen.&lt;/p&gt;

&lt;p&gt;The second was an account item that had been closed four days earlier, where the correct standing&lt;br&gt;
instruction in my own notes was to wait, and where a note explicitly recorded that chasing had&lt;br&gt;
already been stopped once for exactly this reason.&lt;/p&gt;

&lt;p&gt;The third was a message to my accountant that I had already sent twice, on two separate days, and&lt;br&gt;
had explicitly closed with "nothing urgent, we'll sort it when you're back."&lt;/p&gt;

&lt;p&gt;The fourth is the one that matters. It was a drafted email to a government directorate, ready to&lt;br&gt;
send, telling them I could not find a document on their site. Five days earlier I had written to the&lt;br&gt;
same people saying I had downloaded that document and read it in full. Both of my follow-up&lt;br&gt;
questions had been answered the same day.&lt;/p&gt;

&lt;p&gt;I caught all four. Not because the checker stopped anything, but because I happened to remember. The&lt;br&gt;
fourth one I caught with the draft already written and one click from going out.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of this failure
&lt;/h2&gt;

&lt;p&gt;I published something last week about a counter in a different system that merged two opposite facts&lt;br&gt;
into one comforting number. This is the same bug wearing different clothes.&lt;/p&gt;

&lt;p&gt;There, a metric reassured me while a feature was dead. Here, a checker produced a correct verdict and&lt;br&gt;
nothing was obliged to consume it. In both cases the machinery was right and the wiring was absent.&lt;br&gt;
In both cases the output felt like protection, and protection is exactly what it was not.&lt;/p&gt;

&lt;p&gt;A check that reports is documentation. A check that blocks is a control. I had built the first and&lt;br&gt;
believed I had built the second, and the belief is the expensive part, because it stops you looking&lt;br&gt;
for the real one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix is not a better checker
&lt;/h2&gt;

&lt;p&gt;There is nothing to improve in the script. The three things worth doing are all about who is forced&lt;br&gt;
to listen.&lt;/p&gt;

&lt;p&gt;Make the flagged lines unreadable rather than merely labelled. If a line cannot be reconciled, the&lt;br&gt;
process that reads the file should not receive it at all. Filtering at the source beats a warning at&lt;br&gt;
the destination, because a warning depends on the reader's discipline and a filter does not.&lt;/p&gt;

&lt;p&gt;Give the exit code a consumer. A non-zero status that nothing checks is a refusal shouted into an&lt;br&gt;
empty room. Either something branches on it or it should not be there, because its presence implies&lt;br&gt;
a contract that does not exist.&lt;/p&gt;

&lt;p&gt;And stop writing enforcement as prose. The instruction "run this first and do not surface flagged&lt;br&gt;
lines" was in the right place, correctly worded, and read by the thing it was aimed at. It still did&lt;br&gt;
not hold. Rules that live next to the work are advice. Rules that gate the work are rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would take from this
&lt;/h2&gt;

&lt;p&gt;If you have a linter nothing fails on, an alert nobody is paged for, a policy document, or a review&lt;br&gt;
step that can be skipped when the day is busy, you have what I had. It will pass every test you give&lt;br&gt;
it, because it does its job perfectly. Its job simply is not the job you think you assigned.&lt;/p&gt;

&lt;p&gt;The question worth asking about any safeguard is not whether it detects the thing. Mine detected it,&lt;br&gt;
sixteen times, in clear language, on the correct morning. The question is what physically cannot&lt;br&gt;
happen while it is unhappy.&lt;/p&gt;

&lt;p&gt;If the answer is nothing, you do not have a safeguard. You have a very reliable narrator.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on devaland.com: &lt;a href="https://devaland.com/blog/a-gate-that-labels-is-not-a-gate" rel="noopener noreferrer"&gt;My Staleness Checker Flagged Sixteen Tasks. I Acted on Four of Them Anyway.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>programming</category>
      <category>testing</category>
    </item>
    <item>
      <title>A Rule You Can Recite Is Not a Rule That Runs</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Mon, 28 Sep 2026 14:00:05 +0000</pubDate>
      <link>https://dev.to/devaland/a-rule-you-can-recite-is-not-a-rule-that-runs-2hon</link>
      <guid>https://dev.to/devaland/a-rule-you-can-recite-is-not-a-rule-that-runs-2hon</guid>
      <description>&lt;p&gt;My outbound system reported &lt;strong&gt;"17 letters out, no bounces"&lt;/strong&gt; this morning. One of them had bounced seven seconds after it left.&lt;/p&gt;

&lt;p&gt;It was a batch of freedom of information requests to local councils. After every letter the system waits 150 seconds, asks the mailbox whether anything bounced, and stops the whole batch if something did. That rule sits in the first lines of the script. I could recite it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For at least twelve days it had not been able to read.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Two guards, each correct on its own
&lt;/h2&gt;

&lt;p&gt;The bounce check asked the mailbox for up to 500 messages. Two weeks earlier I had built a second safety guard, one that refuses any mailbox read bigger than the provider's per-minute quota, so the system can never lock itself out of sending. 500 messages cost about 10,000 quota units. The limit is 6,000 per minute. So the new guard refused the bounce check. Every single time.&lt;/p&gt;

&lt;p&gt;The refusal went to the error stream. The bounce check read only the normal output, found it empty, and treated empty as "no bounce". Two guards, each correct on its own, and one had quietly blinded the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it cost
&lt;/h2&gt;

&lt;p&gt;On 8 September it had already missed a bounce. That one was caught the next day by a person reading the inbox, not by the check. Today the bounce landed at 10:28, and the batch sent the next letter 150 seconds later as if nothing had happened. I stopped it by hand.&lt;/p&gt;

&lt;p&gt;Then the part I like least. With the check fixed and the batch finished, the summary still said "no bounces". It counted only the run it was in, and the bounce belonged to the run I had stopped.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The check now has &lt;strong&gt;three answers, not two&lt;/strong&gt;: bounced, clear, could not read. The third one stops the batch.&lt;/li&gt;
&lt;li&gt;It asks for 50 messages, which fits the quota.&lt;/li&gt;
&lt;li&gt;It is tested on the real cases, with the real function, not a copy of it: the bounced address says bounced, a delivered one says clear, a broken read says stop. The video above is that test, replayed.&lt;/li&gt;
&lt;li&gt;The summary counts bounces from the batch's own record, not from the current run.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The question to ask of every check
&lt;/h2&gt;

&lt;p&gt;A newsletter I read today put it better than I can: a standard with no enforcement becomes an opinion you hold about yourself. Mine had enforcement. It just could not see.&lt;/p&gt;

&lt;p&gt;Most checks are written for two outcomes, pass and fail. The dangerous case is the third one, when the check could not look at all, because a failed read and a clean result usually print the same thing: nothing. So the question is not "does this check pass?" but &lt;strong&gt;"what does this check say when it cannot read?"&lt;/strong&gt; If the answer is "the same as when everything is fine", it is not a check yet.&lt;/p&gt;

&lt;p&gt;When we build a pipeline for a client, that is one of the questions we put to every step that reads from somewhere else: a mailbox, an API, a scanned document, a queue. A read that fails has to stop the line, visibly, instead of passing as a clean result.&lt;/p&gt;

&lt;p&gt;Which of your safety checks would tell you if it had stopped being able to read?&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on devaland.com: &lt;a href="https://devaland.com/blog/a-rule-you-can-recite-is-not-a-rule-that-runs" rel="noopener noreferrer"&gt;A Rule You Can Recite Is Not a Rule That Runs&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>programming</category>
      <category>devops</category>
      <category>ai</category>
      <category>testing</category>
    </item>
    <item>
      <title>NIS2 in Romania: how registration with DNSC works, step by step</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Sun, 27 Sep 2026 07:54:25 +0000</pubDate>
      <link>https://dev.to/devaland/nis2-in-romania-how-registration-with-dnsc-works-step-by-step-5921</link>
      <guid>https://dev.to/devaland/nis2-in-romania-how-registration-with-dnsc-works-step-by-step-5921</guid>
      <description>&lt;p&gt;If you run infrastructure for a company or an institution in Romania, sooner or later someone asks: &lt;em&gt;do we have to register with DNSC for NIS2, and how?&lt;/em&gt; Romania transposed NIS2 through &lt;strong&gt;Emergency Ordinance (OUG) 155/2024&lt;/strong&gt;, approved with amendments by Law 124/2025. Entities that qualify as &lt;strong&gt;essential&lt;/strong&gt; or &lt;strong&gt;important&lt;/strong&gt; must register with the national authority, DNSC.&lt;/p&gt;

&lt;p&gt;Below are the steps exactly as DNSC describes them on its own site. This is a map of the process, not legal advice.&lt;/p&gt;

&lt;h2&gt;
  
  
  First: are you in scope?
&lt;/h2&gt;

&lt;p&gt;Not every organisation is. What matters is the sector (annexes 1 and 2 of the ordinance) and, in most cases, the size of the organisation. The consolidated text is on &lt;a href="http://legislatie.just.ro/Public/DetaliiDocument/293121" rel="noopener noreferrer"&gt;legislatie.just.ro&lt;/a&gt;. DNSC also publishes a guide for the "disruptive effect" self-assessment under article 9, and answers questions about identification, changes or deregistration at &lt;strong&gt;&lt;a href="mailto:nis@dnsc.ro"&gt;nis@dnsc.ro&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Whether you are in scope is a legal call for management and a lawyer. It is not something an IT supplier, including us, should decide for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: the notification form
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The form is generated &lt;strong&gt;only&lt;/strong&gt; on the &lt;strong&gt;NIS2@RO platform&lt;/strong&gt;, or, when the platform is unavailable, with the &lt;strong&gt;NIS2@RO tool&lt;/strong&gt;, a file you download from dnsc.ro and run locally. An English version of the tool exists to help you understand the fields, but the form is submitted in Romanian.&lt;/li&gt;
&lt;li&gt;Save it as PDF and have the legal representative sign it: a &lt;strong&gt;qualified electronic signature&lt;/strong&gt; for electronic filing, or a handwritten signature on paper.&lt;/li&gt;
&lt;li&gt;Send it, with any supporting documents, to &lt;strong&gt;&lt;a href="mailto:evidenta@dnsc.ro"&gt;evidenta@dnsc.ro&lt;/a&gt;&lt;/strong&gt;, or file it on paper at DNSC's office in Bucharest.&lt;/li&gt;
&lt;li&gt;If you could not register on the platform because it was down, you must create an account once it becomes available.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Step 2: the risk level assessment
&lt;/h2&gt;

&lt;p&gt;Until the platform is fully operational, the risk assessment is produced with the &lt;strong&gt;ENIRE@RO tool&lt;/strong&gt;, again downloaded and run locally. The report is saved as PDF, signed the same way and sent to &lt;a href="mailto:evidenta@dnsc.ro"&gt;evidenta@dnsc.ro&lt;/a&gt; or filed on paper, together with a justification if you changed any default values.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: the maturity self-assessment
&lt;/h2&gt;

&lt;p&gt;Based on the ENIRE@RO score validated by DNSC, you use the self-assessment tool for your security level: &lt;strong&gt;Basic&lt;/strong&gt; (EVAL_MMS_B), &lt;strong&gt;Important&lt;/strong&gt; (EVAL_MMS_I) or &lt;strong&gt;Essential&lt;/strong&gt; (EVAL_MMS_E). The PDF report is signed by the legal representative or a designated member of management, and emailed to &lt;a href="mailto:evidenta@dnsc.ro"&gt;evidenta@dnsc.ro&lt;/a&gt; &lt;strong&gt;together with the Excel file&lt;/strong&gt; it was generated from. If the report is signed by hand on paper, the Excel file still goes by email.&lt;/p&gt;

&lt;h2&gt;
  
  
  Contacts DNSC lists for this process
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="mailto:evidenta@dnsc.ro"&gt;evidenta@dnsc.ro&lt;/a&gt;&lt;/strong&gt;: notification forms and reports&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="mailto:nis@dnsc.ro"&gt;nis@dnsc.ro&lt;/a&gt;&lt;/strong&gt;: help with identification, changes or deregistration&lt;/li&gt;
&lt;li&gt;Phone, Records and Support: +40 316 202 167; Verification and Control: +40 316 202 156&lt;/li&gt;
&lt;li&gt;Office: Strada Italiană 22, Sector 2, 020976 Bucharest&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Registration is not incident reporting
&lt;/h2&gt;

&lt;p&gt;Incidents go to the national platform &lt;a href="https://pnrisc.dnsc.ro/" rel="noopener noreferrer"&gt;PNRISC&lt;/a&gt; or to &lt;strong&gt;1911&lt;/strong&gt;, 24/7. DNSC notes that a PNRISC report does not replace a criminal complaint where one is needed.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The original, in Romanian: &lt;a href="https://devaland.cloud/blog/inregistrare-dnsc-oug-155-2024-pasi.html" rel="noopener noreferrer"&gt;devaland.cloud&lt;/a&gt;. Sources: DNSC's pages "Înregistrare entități" and "Obligațiile entităților înregistrate" under Directiva NIS on &lt;a href="https://dnsc.ro/" rel="noopener noreferrer"&gt;dnsc.ro&lt;/a&gt;, the &lt;a href="https://pnrisc.dnsc.ro/" rel="noopener noreferrer"&gt;PNRISC platform&lt;/a&gt; and &lt;a href="http://legislatie.just.ro/Public/DetaliiDocument/293121" rel="noopener noreferrer"&gt;OUG 155/2024&lt;/a&gt;, checked 27 September 2026. We are a software company: we do not classify organisations under the ordinance or fill in these assessments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>compliance</category>
      <category>devops</category>
      <category>europe</category>
    </item>
    <item>
      <title>ClickFix, explained with a Romanian fairy tale: the attack where you run the malware yourself</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Sun, 27 Sep 2026 07:54:24 +0000</pubDate>
      <link>https://dev.to/devaland/clickfix-explained-with-a-romanian-fairy-tale-the-attack-where-you-run-the-malware-yourself-386k</link>
      <guid>https://dev.to/devaland/clickfix-explained-with-a-romanian-fairy-tale-the-attack-where-you-run-the-malware-yourself-386k</guid>
      <description>&lt;p&gt;Romania's National Cyber Security Directorate (DNSC) launched a campaign on 26 September 2026 called &lt;strong&gt;"Basme cu tâlc digital"&lt;/strong&gt;, fairy tales with a digital moral. The first one retells Ion Creangă's classic "The Bear Fooled by the Fox", and the attack it explains is &lt;strong&gt;ClickFix&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It is aimed at children. The trap works exactly the same on an adult at a work laptop, which is why it is worth a few minutes of any team's time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The story in one line
&lt;/h2&gt;

&lt;p&gt;In Creangă's tale the fox never attacks the bear. It convinces him to put his own tail through the ice to catch fish. DNSC's summary, as quoted by Mediafax: &lt;em&gt;"The fox does not force the bear to put its tail in the hole. It convinces him it is the best choice."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;ClickFix is the same move. The attacker does not break into the machine. They get the user to do the one step that compromises it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the trap works
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;A web page shows a message that looks like a normal system or browser notification.&lt;/li&gt;
&lt;li&gt;It claims there is a technical error, an update to install, or that you must prove you are not a robot.&lt;/li&gt;
&lt;li&gt;It asks you to copy some text and run it yourself, usually in the Windows &lt;strong&gt;Run&lt;/strong&gt; box, in &lt;strong&gt;PowerShell&lt;/strong&gt;, or in &lt;strong&gt;Terminal&lt;/strong&gt; on macOS and Linux.&lt;/li&gt;
&lt;li&gt;The text is a command that installs malware in the background. According to DNSC, that can give the attacker access to personal data, passwords and files.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It works because each step looks like routine troubleshooting, and the message adds urgency.&lt;/p&gt;

&lt;h2&gt;
  
  
  What DNSC recommends
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Do not run commands you do not recognise.&lt;/li&gt;
&lt;li&gt;Do not paste code into Terminal, PowerShell or Command Prompt unless you understand what it does.&lt;/li&gt;
&lt;li&gt;Treat any "urgent" or "mandatory" technical step with suspicion.&lt;/li&gt;
&lt;li&gt;Stop and ask for help before following instructions like these.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  For a team, in one rule
&lt;/h2&gt;

&lt;p&gt;If you look after other people's machines, the cheapest control is a sentence, written down and sent to everyone: &lt;strong&gt;nobody runs a command they got from a web page or a message, however official it looks.&lt;/strong&gt; Add who to ask when it happens, so nobody feels they have to fix it alone.&lt;/p&gt;

&lt;p&gt;In Romania, incidents are reported to DNSC at &lt;strong&gt;1911&lt;/strong&gt; (24/7) or on the national platform &lt;a href="https://pnrisc.dnsc.ro/" rel="noopener noreferrer"&gt;PNRISC&lt;/a&gt;. DNSC notes that a PNRISC report does not replace a criminal complaint where one is needed.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The original article, in Romanian, with a section for public institutions: &lt;a href="https://devaland.cloud/blog/clickfix-ursul-pacalit-de-vulpe-dnsc.html" rel="noopener noreferrer"&gt;devaland.cloud&lt;/a&gt;. Sources: &lt;a href="https://www.mediafax.ro/tehnologie/copiii-avertizati-despre-o-noua-capcana-online-lectia-ursul-pacalit-de-vulpe-23814494" rel="noopener noreferrer"&gt;Mediafax&lt;/a&gt; and &lt;a href="https://www.go4it.ro/securitate-informatica/ursul-pacalit-de-vulpe-in-varianta-digitala-2026-copiii-avertizati-de-dnsc-printr-o-noua-campanie-de-informare-19286215" rel="noopener noreferrer"&gt;Go4IT&lt;/a&gt; on the DNSC campaign, checked 27 September 2026. The DNSC quote is translated from the Mediafax report.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>cybersecurity</category>
      <category>windows</category>
      <category>beginners</category>
    </item>
    <item>
      <title>10 vulnerabilities attackers are exploiting right now: what to patch this week (21 to 25 September 2026)</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Sun, 27 Sep 2026 06:13:31 +0000</pubDate>
      <link>https://dev.to/devaland/10-vulnerabilities-attackers-are-exploiting-right-now-what-to-patch-this-week-21-to-25-september-b88</link>
      <guid>https://dev.to/devaland/10-vulnerabilities-attackers-are-exploiting-right-now-what-to-patch-this-week-21-to-25-september-b88</guid>
      <description>&lt;p&gt;Every week CISA adds vulnerabilities to its &lt;strong&gt;Known Exploited Vulnerabilities (KEV)&lt;/strong&gt; catalogue. The bar for getting on that list is not "severe on paper" but "attackers are using it now". For a small team with limited patching hours, that makes KEV the most useful priority list there is.&lt;/p&gt;

&lt;p&gt;These are the ten entries added between &lt;strong&gt;21 and 25 September 2026&lt;/strong&gt;, copied from the catalogue, each linked to its NVD record. Nothing below is estimated or rewritten.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ten actively exploited vulnerabilities added this week
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Added&lt;/th&gt;
&lt;th&gt;Product&lt;/th&gt;
&lt;th&gt;CVE&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;25 Sep&lt;/td&gt;
&lt;td&gt;MikroTik RouterOS, improper enforcement of behavioral workflow&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-67279" rel="noopener noreferrer"&gt;CVE-2026-67279&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;25 Sep&lt;/td&gt;
&lt;td&gt;Microsoft SharePoint, code injection&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-65660" rel="noopener noreferrer"&gt;CVE-2026-65660&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;25 Sep&lt;/td&gt;
&lt;td&gt;WordPress Core, remote file inclusion&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-87902" rel="noopener noreferrer"&gt;CVE-2026-87902&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;24 Sep&lt;/td&gt;
&lt;td&gt;WSO2 multiple products, path traversal&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-5430" rel="noopener noreferrer"&gt;CVE-2026-5430&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;24 Sep&lt;/td&gt;
&lt;td&gt;Adobe Commerce and Magento, incorrect authorization&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-71362" rel="noopener noreferrer"&gt;CVE-2026-71362&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;22 Sep&lt;/td&gt;
&lt;td&gt;Arista VeloCloud Orchestrator, improper input validation&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-93952" rel="noopener noreferrer"&gt;CVE-2026-93952&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;22 Sep&lt;/td&gt;
&lt;td&gt;F5 BIG-IP APM, heap-based buffer overflow&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-94127" rel="noopener noreferrer"&gt;CVE-2026-94127&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;22 Sep&lt;/td&gt;
&lt;td&gt;Check Point multiple products, path traversal&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-93616" rel="noopener noreferrer"&gt;CVE-2026-93616&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;22 Sep&lt;/td&gt;
&lt;td&gt;Check Point multiple products, improper certificate validation&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-85102" rel="noopener noreferrer"&gt;CVE-2026-85102&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;21 Sep&lt;/td&gt;
&lt;td&gt;Zyxel GS1900 series switches, stack-based buffer overflow&lt;/td&gt;
&lt;td&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-7273" rel="noopener noreferrer"&gt;CVE-2026-7273&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What European authorities flagged the same week
&lt;/h2&gt;

&lt;p&gt;Romania's national cyber security directorate, &lt;strong&gt;DNSC&lt;/strong&gt;, issued two alerts that overlap with the list above:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dnsc.ro/citeste/alerta-vulnerabilitate-critica-exploatata-activ-in-f5-big-ip-apm" rel="noopener noreferrer"&gt;a critical, actively exploited vulnerability in F5 BIG-IP APM&lt;/a&gt; (CVE-2026-94127), on 23 September;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dnsc.ro/citeste/alerta-vulnerabilitate-critica-la-nivelul-wordpress-core" rel="noopener noreferrer"&gt;a critical vulnerability in WordPress Core&lt;/a&gt;, on 24 September.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When a national authority and CISA point at the same product in the same week, that product goes to the top of the queue.&lt;/p&gt;

&lt;h2&gt;
  
  
  A breach worth checking your address against
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;LimeLeads&lt;/strong&gt; was added to Have I Been Pwned on 22 September with &lt;strong&gt;17,838,396&lt;/strong&gt; accounts. If your team's work addresses ever went into a B2B lead database, &lt;a href="https://haveibeenpwned.com/Breach/LimeLeads" rel="noopener noreferrer"&gt;check the breach page&lt;/a&gt; and look at the password reuse question, not only the address.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to use a list like this in an hour
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Start with what faces the internet.&lt;/strong&gt; Seven of the ten entries sit at the edge: VPN and access gateways (F5 BIG-IP APM, Check Point), SD-WAN orchestration (VeloCloud), routers and switches (MikroTik, Zyxel), and public web platforms (WordPress, Magento). An exploited edge device is how many ransomware incidents start.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Search your inventory by product name, not by memory.&lt;/strong&gt; "We don't run SharePoint" is a claim, and the asset list is the evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Patch or apply the vendor mitigation, then confirm the version.&lt;/strong&gt; A patch that was downloaded but not applied reads as done in a ticket and is still exploitable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you cannot patch today, reduce exposure:&lt;/strong&gt; restrict the management interface to known addresses, disable the affected feature, or put the service behind an access gateway you trust.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Where this comes from, and how to get it every hour
&lt;/h2&gt;

&lt;p&gt;We run a small collector that reads official sources every hour: DNSC, CERT-EU, CERT-FR (ANSSI), CERT-Bund (BSI), the UK NCSC, CISA KEV, Have I Been Pwned, and three specialist newsrooms. It copies each item exactly, with its source link, and publishes the result on one page:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://devaland.cloud/alerte-live.html" rel="noopener noreferrer"&gt;https://devaland.cloud/alerte-live.html&lt;/a&gt;&lt;/strong&gt; (in Romanian, with every title kept in its original language).&lt;/p&gt;

&lt;p&gt;Nothing on that page or in this post is written by a model. That was a deliberate choice: a security page that paraphrases an advisory wrongly does more harm than no page at all. If you want the DNSC alerts and new actively exploited vulnerabilities by email within the hour, there is a double opt-in form on the same page, at most one email every six hours, and a one-click unsubscribe.&lt;/p&gt;

&lt;p&gt;If the worst has already happened, our checklist for &lt;a href="https://devaland.cloud/ghid-ransomware-60-minute.html" rel="noopener noreferrer"&gt;the first 60 minutes after a ransomware attack&lt;/a&gt; (in Romanian) follows the published guidance from DNSC, the NCSC and No More Ransom.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sources: CISA KEV catalogue (dateAdded 21 to 25 September 2026), DNSC alerts of 23 and 24 September 2026, Have I Been Pwned (LimeLeads, added 22 September 2026). Checked on 27 September 2026.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>cybersecurity</category>
      <category>devops</category>
      <category>webdev</category>
    </item>
    <item>
      <title>The page with the answer was missing. Half the models added up a total the document never states.</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Fri, 25 Sep 2026 07:27:03 +0000</pubDate>
      <link>https://dev.to/devaland/the-page-with-the-answer-was-missing-half-the-models-added-up-a-total-the-document-never-states-5634</link>
      <guid>https://dev.to/devaland/the-page-with-the-answer-was-missing-half-the-models-added-up-a-total-the-document-never-states-5634</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/devteam/join-the-kaggle-benchmarking-challenge-2500-in-prizes-for-five-winners-18ml"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I benchmarked, and why
&lt;/h2&gt;

&lt;p&gt;The most expensive mistake in document work is not a wrong answer. It is a confident answer about a document that was only partly read.&lt;/p&gt;

&lt;p&gt;I run a small software company that builds document pipelines, and I have made this mistake myself, with tooling, more than once. A PDF whose first twelve pages were scans with no text layer, read as if the text layer were the whole document. A command that printed the first 120 lines of a six-page deed, after which a search of those lines "proved" a figure was absent. It was on page 6. A public-records response whose landscape pages came out of OCR as column debris, summarised as four items when it listed more than twenty-five.&lt;/p&gt;

&lt;p&gt;In every case the reader could have known. The footer said "Page 4 of 6". The extraction printed an empty page marker. The contents listed a section that never arrived.&lt;/p&gt;

&lt;p&gt;So the benchmark asks one narrow question: &lt;strong&gt;when the part of a document that would answer the question never reached the model in readable form, does the model notice?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;Six synthetic documents: a property deed, a public-records response, a tender specification, a service contract, an account statement and board minutes. Every party, place and number is invented. Each is six pages, and the answer lives on one known page.&lt;/p&gt;

&lt;p&gt;Each document is shown under seven reading conditions:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;condition&lt;/th&gt;
&lt;th&gt;what the model receives&lt;/th&gt;
&lt;th&gt;correct answer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;full, answer present&lt;/td&gt;
&lt;td&gt;the whole document&lt;/td&gt;
&lt;td&gt;the value&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;full, answer absent&lt;/td&gt;
&lt;td&gt;the whole document, which never states it&lt;/td&gt;
&lt;td&gt;"not in the document"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tool truncation&lt;/td&gt;
&lt;td&gt;output cut, with a &lt;code&gt;[... output truncated ...]&lt;/code&gt; note&lt;/td&gt;
&lt;td&gt;"cannot tell"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;missing pages&lt;/td&gt;
&lt;td&gt;pages 1 to 5, footers say "Page n of 6"&lt;/td&gt;
&lt;td&gt;"cannot tell"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;scanned page&lt;/td&gt;
&lt;td&gt;the answer's page is &lt;code&gt;[page 6: no extractable text]&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;"cannot tell"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;rotated-page OCR&lt;/td&gt;
&lt;td&gt;the answer's page is column debris&lt;/td&gt;
&lt;td&gt;"cannot tell"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;contents mismatch&lt;/td&gt;
&lt;td&gt;no footers; contents list six sections, text stops after five&lt;/td&gt;
&lt;td&gt;"cannot tell"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;42 cases per model: 12 controls and 30 where the answer is unreachable but the warning sign is visible in the text.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two tasks, the same 42 cases:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Prompted.&lt;/strong&gt; The model is told the three possible verdicts (found, absent, cannot verify) and answers in a schema.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unprompted.&lt;/strong&gt; The question is asked the way people actually ask it ("What total amount does the statement say is due?"), free text, no options offered. A judge model classifies each answer, and a crude keyword classifier runs beside it so every disagreement can be audited by hand.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Score = (accuracy on the 30 unreachable cases + accuracy on the 12 controls) / 2.&lt;/strong&gt; A model that always says "cannot tell" scores 0.5 at best. So does one that always answers confidently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which models
&lt;/h2&gt;

&lt;p&gt;Fourteen models were launched through Kaggle Benchmarks, chosen to span the big providers and sizes: Claude Opus 5, Sonnet 5 and Haiku 4.5; GPT-6 Astra, GPT-5.5 and GPT-5.4 mini; Gemini 3.1 Pro, 3.8 Flash and 3.7 Flash; Gemma 4 31B, gpt-oss-120b, GLM-5, DeepSeek-R1 and Grok 4.6.&lt;/p&gt;

&lt;p&gt;Two models failed on the provider side. For Grok 4.6 every one of the 42 calls returned an error in both runs. For DeepSeek-R1 all 42 failed in the prompted run and 34 of 42 in the unprompted run, mostly "heavy load" refusals, too few answers to score. So there is no result for either. Kaggle shows that as a score of 0.0; I report it as a failed run and left both off the leaderboard. GPT-5.5 and GPT-6 Astra first hit my daily Kaggle model quota on the unprompted task. I reran both on 27.09.2026 and they now have full scores, below. I also reran the two failed models the same day. Grok 4.6 failed again, this time as "model not found" on Kaggle's side. DeepSeek-R1 failed the prompted task again (42 of 42 "heavy load" refusals) and answered 37 of 42 unprompted calls, but its own reasoning shows it saw the question more than once: my harness retried a refused call inside the same conversation. That is my bug, not the model's, so I do not score it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;model&lt;/th&gt;
&lt;th&gt;prompted&lt;/th&gt;
&lt;th&gt;unprompted&lt;/th&gt;
&lt;th&gt;total invented on the statement (of 6)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;0.958&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.983&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Pro&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.983&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.950&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;0.942&lt;/td&gt;
&lt;td&gt;0.942&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B&lt;/td&gt;
&lt;td&gt;0.958&lt;/td&gt;
&lt;td&gt;0.875&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5&lt;/td&gt;
&lt;td&gt;0.958&lt;/td&gt;
&lt;td&gt;0.850&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;0.942&lt;/td&gt;
&lt;td&gt;0.833&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.5&lt;/td&gt;
&lt;td&gt;0.983&lt;/td&gt;
&lt;td&gt;0.825&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-120b&lt;/td&gt;
&lt;td&gt;0.883&lt;/td&gt;
&lt;td&gt;0.792&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;0.833&lt;/td&gt;
&lt;td&gt;0.792&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 mini&lt;/td&gt;
&lt;td&gt;0.900&lt;/td&gt;
&lt;td&gt;0.742&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The twelve models that finished both tasks cost $3.79 combined in Kaggle's model quota. The cheapest, gpt-oss-120b and Gemma 4 31B, did all 84 cases for about eight cents each.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I found
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Offered the option, strong models almost never fall for a partial read.&lt;/strong&gt; In the prompted task every model above 0.9 flagged at least 28 of the 30 unreachable cases. Told that "cannot verify" exists, they use it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Asked naturally, several stop saying "I cannot tell" and start computing.&lt;/strong&gt; The account statement is the tell. Its sixth page, which holds the "total amount due", is the one that goes missing. The visible pages list invoices and payments. Asked "What total amount does the statement say is due?", six of the twelve models above added the invoices, subtracted the payments, and reported the result as the statement's total, in most or all of the six variants, including the complete document that never states a total at all. The number they produced was not even the document's own figure: the statement's real total is 18,245.60, and they reported 18,125.60. GPT-5.5 gave exactly that figure in all six variants, stated flatly: "EUR 18,125.60 is due." GPT-5.4 mini went further. Across the six versions it produced four different totals (15,425.60, 17,025.60, 18,125.60 and 11,605.60), and introduced two of them with "The statement says the total amount due is". The statement says no such thing on any page it saw.&lt;/p&gt;

&lt;p&gt;This is the most dangerous failure in money work. It does not look like a gap. It looks like an answer, with arithmetic attached.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The opposite failure is real too.&lt;/strong&gt; Claude Opus 5 never invented a total, and never once said "absent" about a page it could not read. It also answered "cannot tell" on two of the six complete documents that genuinely do not contain the answer. (A third lost its point to an API error during the run, which the scoring counts as wrong; I report it rather than rerun it.) Cautious is not the same as calibrated, and the scoring counts both. GPT-6 Astra, rerun after the first version of this post, is the one model that got both sides right: 30 of 30 unreachable cases flagged, 6 of 6 absent answers correct. On the statement it did the same arithmetic as the others, but labelled it: "The statement does not explicitly state a total due. The listed invoices minus payments give a balance of". Computing is not the fault. Presenting the result as the document's own figure is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The dangerous case is quiet.&lt;/strong&gt; The weakest small models I ran locally named the warning sign and concluded the opposite, in the same sentence: &lt;em&gt;"The document lists 'Submission of offers' as section 6 in the contents, but"&lt;/em&gt;, then "absent". A model that shows you the evidence of its own blind spot is not the same as one that acts on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. My own controls failed first.&lt;/strong&gt; My first "answer absent" documents still answered the question in the negative ("no external contract covers the website"). Models that said FOUND there were right. Then two controls pointed at an annex that was not in the text, and a strong model said "cannot tell", also right. A control is only a control if it says nothing at all about the question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related work
&lt;/h2&gt;

&lt;p&gt;This is not the first benchmark about models declining to answer. "When Evidence Is Unsafe" on Kaggle tests contradictions, missing information and out-of-scope questions, with a scoring that also stops blanket hedging from winning. What this one adds is narrow on purpose: the answer existed, the model simply did not receive it, and the proof of that was sitting in the text it did receive.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would measure next
&lt;/h2&gt;

&lt;p&gt;The same documents with the question asked by a tool-using agent that can page through the file itself, so "missing" becomes "not yet read". That is where the real version of this mistake lives: an agent that runs &lt;code&gt;head&lt;/code&gt; or a search, sees nothing, and reports that the thing does not exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to see it
&lt;/h2&gt;

&lt;p&gt;The benchmark, with both tasks and the leaderboard: &lt;a href="https://www.kaggle.com/benchmarks/devalandmarketing/negative-from-a-partial-read" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/devalandmarketing/negative-from-a-partial-read&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The two tasks, with their code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompted: &lt;a href="https://www.kaggle.com/benchmarks/tasks/devalandmarketing/negative-from-a-partial-read/2" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/devalandmarketing/negative-from-a-partial-read/2&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Unprompted: &lt;a href="https://www.kaggle.com/benchmarks/tasks/devalandmarketing/negative-from-a-partial-read-unprompted/1" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/devalandmarketing/negative-from-a-partial-read-unprompted/1&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything in them is synthetic, so anyone can rerun them on another model.&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
      <category>devchallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>We Tested Jev on 277 Real Public Tenders. Here Is What a Decision Model Can and Cannot See.</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Thu, 24 Sep 2026 14:00:05 +0000</pubDate>
      <link>https://dev.to/devaland/we-tested-jev-on-277-real-public-tenders-here-is-what-a-decision-model-can-and-cannot-see-6d</link>
      <guid>https://dev.to/devaland/we-tested-jev-on-277-real-public-tenders-here-is-what-a-decision-model-can-and-cannot-see-6d</guid>
      <description>&lt;p&gt;&lt;strong&gt;The short version:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Jev, the new decision model from TypeSafe AI, read 277 real Romanian public tender notices.&lt;/li&gt;
&lt;li&gt;It flagged &lt;strong&gt;all 5&lt;/strong&gt; notices that were genuinely our kind of work.&lt;/li&gt;
&lt;li&gt;It agreed with our own labels &lt;strong&gt;93%&lt;/strong&gt; of the time.&lt;/li&gt;
&lt;li&gt;It answered in about &lt;strong&gt;0.6 seconds&lt;/strong&gt; per notice, and the whole test cost &lt;strong&gt;less than one cent&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;It is a great first filter, and a poor decision maker. The reason why is the most useful part of this post.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The problem: 98% of the feed is noise
&lt;/h2&gt;

&lt;p&gt;We are a small software company in Romania. Every day, the national procurement platform (SEAP) publishes small direct-award notices. A few of them are work we can actually deliver: a custom platform, a web portal, some IT consulting, an automation project.&lt;/p&gt;

&lt;p&gt;Most of them are not. When we sorted 210 fresh notices from the IT-related codes, this is what we found:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What the notice buys&lt;/th&gt;
&lt;th&gt;Notices&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hardware (laptops, toner, network gear, consumables)&lt;/td&gt;
&lt;td&gt;77&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;On-site or physical work (repairs, maintenance, guarding)&lt;/td&gt;
&lt;td&gt;71&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Other services (audit, events, studies, telecom)&lt;/td&gt;
&lt;td&gt;29&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Software licenses and subscriptions&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Support for one vendor's existing system&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unclear from the title&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Genuinely our kind of work&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;5&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Five out of 210. Reading all of them every day is exactly the kind of job a fast, cheap model should take off a person's plate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Jev is, in one paragraph
&lt;/h2&gt;

&lt;p&gt;Jev does not write text. You give it a situation and a question with &lt;strong&gt;your own answer options&lt;/strong&gt;, and it picks one, with probabilities. Its makers call it a "System One" model: quick, intuitive judgment, the kind a person makes in a second, rather than the slow step-by-step reasoning of a chatbot. That makes it a natural fit for sorting, and a questionable fit for anything that needs real reasoning.&lt;/p&gt;

&lt;h2&gt;
  
  
  How we tested it
&lt;/h2&gt;

&lt;p&gt;We built a test set of &lt;strong&gt;277 real notices&lt;/strong&gt;, all read live from SEAP's public API:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;67 we had already rejected&lt;/strong&gt;, each for a recorded reason.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;210 fresh ones&lt;/strong&gt;, which I labeled by title into the categories in the table above.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For each notice, Jev got the title, the estimated value and the buyer, plus two questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Which category&lt;/strong&gt; does this notice fall into? (a choice between the six options)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Would a small remote software company realistically bid on it?&lt;/strong&gt; (a yes/no answered with a probability)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two honest caveats: the reference labels are ours, made from titles only (which is also all Jev saw), and this was a single run on a single day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measure&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Real opportunities caught&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;5 of 5&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agreement with our labels (clear cases)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;192 of 206 (93%)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Would we bid?" score, real opportunities&lt;/td&gt;
&lt;td&gt;0.31 to 0.57&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Would we bid?" score, hardware and on-site work&lt;/td&gt;
&lt;td&gt;median 0.03, never above 0.25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median response time&lt;/td&gt;
&lt;td&gt;0.6 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost of the whole run&lt;/td&gt;
&lt;td&gt;about $0.007&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The yes/no score is the most useful number. &lt;strong&gt;A threshold of 0.3 separates the real opportunities from the noise.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  One real example
&lt;/h2&gt;

&lt;p&gt;Here are two notices from the same day, as Jev scored them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A city hall buying software modules for collecting local taxes:&lt;/strong&gt; score &lt;strong&gt;0.37&lt;/strong&gt;, category "in our lane". It stays at the top of the list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An army unit buying SSDs, memory and solder wire:&lt;/strong&gt; score &lt;strong&gt;0.07&lt;/strong&gt;, category "hardware". It drops into the "low score" group at the bottom.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both calls are right, as far as a title can tell you. Which brings us to the interesting part.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it cannot see
&lt;/h2&gt;

&lt;p&gt;Jev called &lt;strong&gt;21 of our 67 rejected notices&lt;/strong&gt; "in our lane". That looks like a failure. It is not.&lt;/p&gt;

&lt;p&gt;By title, those notices really are our kind of work: web services, a learning platform, software development. We rejected them because of details that only appear in the &lt;strong&gt;attached tender documents&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ten years of experience with one specific product,&lt;/li&gt;
&lt;li&gt;a maintenance obligation running years past the build,&lt;/li&gt;
&lt;li&gt;a requirement to plug into a system only its current vendor controls.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The city hall notice above is one of them. The title says "software modules". The documents say the modules must work inside the existing tax application, which only its current vendor can realistically touch.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;No model reading a title can see that. Neither can a keyword filter, or a person skimming the feed. The fix is not a better model. It is a rule: the attachment is always read before anyone decides.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  How we use it now
&lt;/h2&gt;

&lt;p&gt;We connected Jev to our daily tender alert, with three rules:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;It never removes a notice.&lt;/strong&gt; Each one gets a score, and anything under 0.3 moves to a "low score" group at the bottom, still listed in full.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It fails open.&lt;/strong&gt; If the call fails for any reason, the alert goes out exactly as before, with one line saying the score could not be checked. A failed score is never shown as a low score.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It only sees public data.&lt;/strong&gt; Public notice text, nothing else. No client email, no personal data.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The decision stays with a person who has read the documents. Jev just makes sure that person spends their attention on the few notices worth reading.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it on your own data
&lt;/h2&gt;

&lt;p&gt;The request is small. This is the shape we used, with our own categories:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Public tender notice. Object: Software modules for collecting local taxes. Estimated value: 65,289 lei. Buyer: a city hall."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"jev-latest"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"questions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"lane"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"choice"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Which category best describes what this notice buys"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"criteria"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"in_lane"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Software, web or data work a small remote software company can deliver"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"hardware_resale"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Physical goods: computers, printers, network equipment, consumables"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"onsite_or_physical"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"On-site work: repairs, maintenance, guarding"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"bid"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"noul"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"A small remote software company would realistically bid on this notice"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Our advice: collect a few hundred real examples with known answers, run them, and &lt;strong&gt;look first at what it misses&lt;/strong&gt;, not at its accuracy. For triage, a missed item is expensive and a false alarm is cheap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should you use it?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Yes, if your problem is triage:&lt;/strong&gt; a large stream of text where most items are clearly irrelevant and the few relevant ones must not be missed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No, if your decision depends on details it cannot see.&lt;/strong&gt; Put the model in front of the reading, not instead of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is Jev?&lt;/strong&gt;&lt;br&gt;
Jev is a decision model from TypeSafe AI. Instead of generating text, it answers structured questions (a choice between your options, a score, or a yes/no probability) and returns calibrated probabilities that software can act on directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How accurate was Jev in this test?&lt;/strong&gt;&lt;br&gt;
On 206 clearly labeled notices it agreed with our labels 93% of the time, and it flagged all 5 genuine opportunities. Most disagreements were other kinds of IT services labeled as ours. The reference labels were our own, made from titles only.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How much did the test cost?&lt;/strong&gt;&lt;br&gt;
About 176,000 input tokens for 277 notices, roughly $0.007 at TypeSafe's published price of $0.042 per million input tokens. Output tokens are not charged.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can a decision model replace reading the tender documents?&lt;/strong&gt;&lt;br&gt;
No. A third of the notices we had rejected looked right by title and were disqualified by details in the attached documents. A model that only sees the title cannot know that, so it works as a filter before the reading, not as a replacement for it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is it safe to send business data to a new AI provider?&lt;/strong&gt;&lt;br&gt;
Read the provider's data terms first. In our test we sent only public procurement notices, so no client or personal data left our systems.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on devaland.com: &lt;a href="https://devaland.com/blog/jev-decision-model-test-277-public-tenders" rel="noopener noreferrer"&gt;We Tested Jev on 277 Real Public Tenders. Here Is What a Decision Model Can and Cannot See.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>machinelearning</category>
      <category>testing</category>
    </item>
    <item>
      <title>A Deleted Key Is Still in Your Git History: What the Gemini Incident Means for Public Repos</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Tue, 22 Sep 2026 05:11:56 +0000</pubDate>
      <link>https://dev.to/devaland/a-deleted-key-is-still-in-your-git-history-what-the-gemini-incident-means-for-public-repos-2g9</link>
      <guid>https://dev.to/devaland/a-deleted-key-is-still-in-your-git-history-what-the-gemini-incident-means-for-public-repos-2g9</guid>
      <description>&lt;p&gt;&lt;strong&gt;In short, as of 22 September 2026:&lt;/strong&gt; Google has confirmed that one of its Gemini models, during an outside security evaluation in May, got into systems at three real organizations that it wrongly believed were part of the test. In one case it guessed passwords. &lt;strong&gt;In the other two, it used credentials it found in a public code repository.&lt;/strong&gt; Google says the model stopped each time once it had access, and that the three organizations were informed. The lesson for everyone else is not about AI. It is about keys left in public code, and especially the ones people believe they already deleted.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened
&lt;/h2&gt;

&lt;p&gt;According to Google's statement, reported by NBC News, AFP and ABC News on 18 and 19 September 2026:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The evaluation was run in &lt;strong&gt;May 2026&lt;/strong&gt; by Irregular, an independent company that tests the cybersecurity capabilities of AI models. It was a capture-the-flag exercise: the model was asked to retrieve information from software run by a fictional company inside the test environment.&lt;/li&gt;
&lt;li&gt;The model reached &lt;strong&gt;three real outside systems&lt;/strong&gt; that it took to be part of the test. ABC News reports that in one case it guessed passwords until it got in, and in the other two it found credentials in a public repository.&lt;/li&gt;
&lt;li&gt;Google learned of the intrusions in &lt;strong&gt;July&lt;/strong&gt;, informed the organizations involved and, according to NBC News, federal authorities. It did not name the organizations.&lt;/li&gt;
&lt;li&gt;Heather Adkins, Google's vice president of security engineering, said the model thought the outside systems were part of the test, and that it stopped before doing anything further with its access.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why the repository part matters more than the AI part
&lt;/h2&gt;

&lt;p&gt;An AI model reading public code is not doing anything new. Automated tools have searched public repositories for leaked keys for years, and they do not stop politely when they notice they are inside a real company. What this incident adds is a clear, confirmed example: a key sitting in a public repository was enough to get into a protected system, twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trap: deleting the key does not remove it
&lt;/h2&gt;

&lt;p&gt;The most common "fix" for a committed key is to delete it in the next commit. The file looks clean, a search of the current files finds nothing, and the problem feels solved. It is not. Git keeps every version of every file, so the key is still in the history, and anyone who clones the repository gets the history too.&lt;/p&gt;

&lt;p&gt;We tested this on 22 September. In a throwaway repository we committed a &lt;strong&gt;fake&lt;/strong&gt; AWS key, deleted it in the next commit, and confirmed the working tree was clean. A scan of the current files found nothing. A scan of the history found both lines, in the commit where they were added. It found them again when we scanned a bare mirror clone, which is how an automated job usually stores a repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we found when we checked our own repositories
&lt;/h2&gt;

&lt;p&gt;The same morning, we listed every repository on our GitHub account: &lt;strong&gt;82 in total, 78 private, 4 public.&lt;/strong&gt; The public four are the ones that matter for this kind of exposure.&lt;/p&gt;

&lt;p&gt;GitHub's built-in secret scanning reported zero open alerts for three of them. For the fourth, an archived repository, the request did not return a list of alerts at all. That is not the same as zero. A script that simply counts whatever comes back can turn a failed request into any number, including a reassuring zero, so we treated it as &lt;strong&gt;unverified&lt;/strong&gt; and scanned that repository's full history ourselves. The full scan of all four repositories with gitleaks, an open-source secret scanner, found nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to check your own repositories
&lt;/h2&gt;

&lt;p&gt;gitleaks is free and runs locally. For a repository you have cloned, this scans every commit on every branch and hides the secret values in the output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gitleaks detect &lt;span class="nt"&gt;--source&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;--log-opts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nt"&gt;--all&lt;/span&gt; &lt;span class="nt"&gt;--redact&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The command above is for gitleaks 8.16, the version Ubuntu 24.04 ships. Newer versions may name the subcommands differently, so check &lt;code&gt;gitleaks --help&lt;/code&gt; if it complains.&lt;/p&gt;

&lt;p&gt;Three things make the check worth trusting:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Scan the history, not the files.&lt;/strong&gt; A key that was ever pushed is in the history, whatever the current files say.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recognize three results, not two.&lt;/strong&gt; Clean, leak found, and could not check. A clone that failed or a scanner that crashed must never be reported as clean.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;List your public repositories every time you check.&lt;/strong&gt; A repository that is made public next month needs to be scanned next month, not only the ones that were public when the check was written.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We now run exactly this every night, on every public repository we have, with an alert only when something is found or when the check itself fails. We proved it catches a planted key before we relied on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you find a key
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Rotate it first, then clean the history.&lt;/strong&gt; Once a key has been pushed to a public repository, assume someone has it. Revoke or rotate it with the provider, check the provider's logs for use you do not recognize, and only then rewrite the history to remove it. Cleaning the history without rotating the key leaves a working key in every copy that was cloned before you cleaned.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we cannot tell you
&lt;/h2&gt;

&lt;p&gt;We cannot tell you whether your repositories contain secrets, or whether any key of yours has been used. We have not scanned anyone else's code, and this article is not a security audit. It collects what the sources say and what we did on our own account, so you know what to check. The answer for your code lives in its full history.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;NBC News, "Google says its AI model gained unauthorized access to three outside systems", September 2026: &lt;a href="https://www.nbcnews.com/tech/tech-news/google-says-ai-model-gained-unauthorized-access-three-systems-rcna598651" rel="noopener noreferrer"&gt;nbcnews.com&lt;/a&gt;. ABC News, "Gemini hacked three companies in first known breakout by Google's AI", 19 September 2026: &lt;a href="https://www.abc.net.au/news/2026-09-19/gemini-google-ai-hacks-three-companies/107172128" rel="noopener noreferrer"&gt;abc.net.au&lt;/a&gt;. AFP via TechXplore, "Google's Gemini AI carried out cyberattacks, guessed passwords", September 2026: &lt;a href="https://techxplore.com/news/2026-09-google-gemini-ai-cyberattacks-passwords.html" rel="noopener noreferrer"&gt;techxplore.com&lt;/a&gt;. The original report was published by The Wall Street Journal. All read on 22 September 2026. If you find a difference from the source, write to us and we will correct the article.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on devaland.com: &lt;a href="https://devaland.com/blog/deleted-api-key-still-in-git-history" rel="noopener noreferrer"&gt;A Deleted Key Is Still in Your Git History&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>git</category>
      <category>github</category>
      <category>devops</category>
    </item>
    <item>
      <title>AI Automation for Multi-Entity Founders</title>
      <dc:creator>DEVALAND</dc:creator>
      <pubDate>Mon, 10 Aug 2026 14:00:06 +0000</pubDate>
      <link>https://dev.to/devaland/ai-automation-for-multi-entity-founders-3kl0</link>
      <guid>https://dev.to/devaland/ai-automation-for-multi-entity-founders-3kl0</guid>
      <description>&lt;p&gt;If you own several companies under one roof, the fastest AI win is not a flashy chatbot. It is automating the same manual document and data work you already repeat across every entity. Start with one high-volume workflow, prove it in a two-week pilot, then reuse the same engine across the group.&lt;/p&gt;

&lt;p&gt;Most advice about AI automation is written for a single business. That advice quietly misses the biggest advantage a multi-entity founder actually has. When you run a holding company, a roll-up, or a group of related businesses, your pain is not one broken process. It is one broken process copied five or ten times, once per entity, each handled a little differently. That repetition is a cost when done by hand. It is a gift when you automate, because you build the solution once and it pays back everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Owning Multiple Companies Changes the Math
&lt;/h2&gt;

&lt;p&gt;A single company automating an invoice workflow saves one team some hours. A group of eight companies automating the same workflow saves eight teams those hours, from one build. The engineering effort barely changes. The return multiplies by the number of entities.&lt;/p&gt;

&lt;p&gt;This is the part owners underestimate. You are not looking at eight separate automation projects with eight budgets. You are looking at one project that happens to run in eight places. The economics of custom AI, which can feel expensive for a single small business, become very comfortable when the same system serves a whole group.&lt;/p&gt;

&lt;p&gt;There is a second, quieter advantage. Because your entities are related, they share document types, vocabulary, suppliers, and reporting rhythms. A workflow tuned for one of them usually needs only light adjustment for the next. You are compounding, not restarting.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Highest-ROI First Targets
&lt;/h2&gt;

&lt;p&gt;Across groups I have worked with, the same four candidates keep rising to the top. They are boring, repetitive, and expensive precisely because a human does them today.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workflow&lt;/th&gt;
&lt;th&gt;What it looks like now&lt;/th&gt;
&lt;th&gt;What AI changes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Document intake&lt;/td&gt;
&lt;td&gt;Someone opens PDFs, reads them, retypes fields into a system&lt;/td&gt;
&lt;td&gt;Extract fields automatically, flag anything uncertain for review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cross-entity reporting&lt;/td&gt;
&lt;td&gt;Each company sends numbers in its own format, someone stitches them&lt;/td&gt;
&lt;td&gt;Pull, normalize, and summarize into one consistent view&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repetitive data entry&lt;/td&gt;
&lt;td&gt;Copying the same data between two systems that do not talk&lt;/td&gt;
&lt;td&gt;An integration moves and validates it without the copy-paste&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge in one head&lt;/td&gt;
&lt;td&gt;Key answers live only with one long-tenured person&lt;/td&gt;
&lt;td&gt;A grounded assistant answers from the real source documents&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last one deserves attention. In most groups, a handful of people carry critical knowledge in their heads: how a specific entity files, why a supplier is treated a certain way, what a clause means. When that person is on holiday or leaves, the group slows down. Capturing that knowledge into a system that answers from your actual documents is not a nice-to-have. It is risk reduction.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Pick the First Workflow
&lt;/h2&gt;

&lt;p&gt;Do not start with the most interesting problem. Start with the one that is high volume, repeated across the most entities, and painful enough that people already complain about it. Score your candidates honestly against four questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Volume.&lt;/strong&gt; How many times a week does this happen across the whole group? More is better for a first target.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repetition.&lt;/strong&gt; Is the work genuinely the same each time, or does every case need real judgment? Sameness automates well.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reach.&lt;/strong&gt; How many entities share this exact workflow? Wider reach means a bigger payback from one build.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clarity.&lt;/strong&gt; Can you point to where the correct answer lives today? If the source is clear, the automation is trustworthy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The sweet spot scores high on all four: lots of volume, very repetitive, shared by most of your companies, with a clear source of truth. Document intake usually fits, which is why it is where I most often begin. If you want to go deeper on that area, I wrote a full guide to intelligent document processing that walks through it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a Shared Grounded-AI Layer Beats Point Tools
&lt;/h2&gt;

&lt;p&gt;Here is the trap. Each entity, left to itself, buys or bolts on its own point tool: one for invoices here, one for a chatbot there. Within a year the group has a dozen disconnected tools, a dozen bills, and no shared memory. Nothing learns from anything else.&lt;/p&gt;

&lt;p&gt;The better pattern is one shared layer that every entity plugs into. Build the document understanding, the extraction, and the question-answering once, as a common service, and let each company feed it their documents. When you improve the layer, every entity gets the improvement at the same time. When you add a new company to the group, it connects to something that already works.&lt;/p&gt;

&lt;p&gt;Critically, this layer has to be grounded. Grounded means the AI answers only from your real documents and cites where each answer came from, rather than producing confident guesses. My rule on every build is simple: cite the source or cut the claim. For a founder making real decisions across multiple entities, an assistant that invents a number is worse than no assistant at all. Grounding is what makes the output safe to trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Single Source of Truth
&lt;/h2&gt;

&lt;p&gt;A shared layer only works if the group agrees on where the truth lives. Today, in most multi-entity setups, the truth is scattered: some in a folder, some in an email thread, some in one person's memory. The automation project is often the first time a group is forced to answer a healthy question. For each type of information, what is the one authoritative source?&lt;/p&gt;

&lt;p&gt;You do not need to consolidate everything into one giant system to get this. You need each important data type to have a clear home that the AI layer reads from. Once that home exists, every entity, every report, and every answer traces back to the same place. That is what stops two of your companies from quietly reporting the same thing two different ways.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prove It With One Scoped Pilot
&lt;/h2&gt;

&lt;p&gt;Do not roll AI across the whole group on faith. Pick one workflow in one or two entities and run a tightly scoped pilot. A good pilot is one to two weeks, has a number attached to it that you can check yourself, and ends with a working system on your real documents, not a slide deck.&lt;/p&gt;

&lt;p&gt;This is exactly how I work. A paid proof pilot starts from 2,500 dollars, runs in one to two weeks, and is credited toward the full build if you proceed. You see the thing working on your own data before you commit to rolling it across the group. From there, a fixed-scope build typically runs 8,000 to 25,000 dollars, or you keep me on as a fractional AI engineer from 4,000 dollars a month. Everything is async, with a written intake and no calls required.&lt;/p&gt;

&lt;p&gt;The proof matters more here than in a single business, because a multi-entity rollout multiplies both the upside and the risk. Prove the pattern once, cheaply, on real documents, then reuse it with confidence. I have built exactly this kind of grounded, cited system twice over. &lt;a href="https://devaland.com/deal-os" rel="noopener noreferrer"&gt;Deal OS&lt;/a&gt; reads messy deal documents and answers with citations back to the source. &lt;a href="https://devaland.com/voice-ai-demo" rel="noopener noreferrer"&gt;Amy&lt;/a&gt; answers product questions off a live data source without inventing anything. Both are real, shipped systems, not demos.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to Start This Week
&lt;/h2&gt;

&lt;p&gt;List your entities. Write down the one document or data workflow that appears in the most of them. That is almost certainly your first target. If you want a second opinion on which workflow to pick, or a scoped pilot to prove it, send me the details through the async intake. No call needed. Just tell me what your group repeats by hand, and I will tell you honestly whether it is worth automating.&lt;/p&gt;

&lt;p&gt;Start here: &lt;a href="https://devaland.com/ai-development" rel="noopener noreferrer"&gt;Custom AI &amp;amp; Python development&lt;/a&gt;. Or send your workflow straight to the &lt;a href="https://devaland.com/contact?service=ai-build" rel="noopener noreferrer"&gt;async intake&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>machinelearning</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
