<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community</title>
    <description>The most recent home feed on DEV Community.</description>
    <link>https://dev.to</link>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/rss"/>
    <language>en</language>
    <item>
      <title>Agents Don't Need Memory, They Need Documentation: A Practical AGENTS.md Playbook</title>
      <dc:creator>EME GUG</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:14:16 +0000</pubDate>
      <link>https://dev.to/eme_gug_0821b41b948be6516/agents-dont-need-memory-they-need-documentation-a-practical-agentsmd-playbook-18mm</link>
      <guid>https://dev.to/eme_gug_0821b41b948be6516/agents-dont-need-memory-they-need-documentation-a-practical-agentsmd-playbook-18mm</guid>
      <description>&lt;p&gt;Mấy tháng gần đây, mình thấy team nào dùng AI coding agent (Claude Code, Codex CLI, Cursor, Aider...) cũng gặp chung một vấn đề. Hôm nay agent làm rất tốt. Sang hôm sau, nó lại quên sạch, dùng sai convention, chạy sai lệnh test, sửa nhầm file generated. Phản xạ đầu tiên của nhiều người là đi tìm giải pháp "memory": vector DB, plugin nhớ hội thoại, tóm tắt session. Nhưng sau khi thử khá nhiều cách, mình đồng ý với một bài đang hot trên Hacker News: &lt;strong&gt;agent không cần memory, nó cần documentation&lt;/strong&gt;. Bài này chia sẻ cách mình tổ chức docs cho agent trong repo thật, kèm script để docs không bị outdated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Vì sao memory không giải quyết được vấn đề
&lt;/h2&gt;

&lt;p&gt;Memory kiểu "nhớ lại hội thoại cũ" có ba điểm yếu chết người:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Không kiểm chứng được&lt;/strong&gt;: bạn không biết agent đang "nhớ" gì, nhớ đúng hay sai. Một quyết định đã bị revert từ tuần trước vẫn có thể nằm trong memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Không review được&lt;/strong&gt;: memory không đi qua pull request, không ai approve.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Không chia sẻ được&lt;/strong&gt;: memory của máy bạn khác memory của đồng nghiệp, khác luôn memory của agent chạy trên CI.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Documentation trong repo thì ngược lại: nó được version bằng Git, review qua PR, và mọi agent, mọi người đọc cùng một nguồn sự thật. Khi agent làm sai, bạn sửa docs một lần là xong cho tất cả.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Session mới] --&amp;gt; B{Nguồn context}
    B --&amp;gt;|Memory| C[Tóm tắt hội thoại cũ]
    C --&amp;gt; D[Không review, dễ lỗi thời]
    B --&amp;gt;|Docs trong repo| E[AGENTS.md + docs/agents]
    E --&amp;gt; F[Version bằng Git, review qua PR]
    F --&amp;gt; G[Mọi agent dùng chung một sự thật]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Một ý nữa liên quan tới bài "866 commits trong 5 tuần" trên Dev.to: khi agent viết code nhanh hơn tốc độ team hiểu code, docs chính là chỗ để &lt;em&gt;con người&lt;/em&gt; bắt kịp. Viết docs cho agent cũng là viết docs cho chính mình của 3 tháng sau.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cấu trúc docs mà mình đang dùng
&lt;/h2&gt;

&lt;p&gt;Đừng nhét mọi thứ vào một file 2000 dòng. Agent có context window giới hạn, và file càng dài thì phần quan trọng càng bị "loãng". Mình chia ba tầng:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;AGENTS.md&lt;/code&gt; ở root (hoặc &lt;code&gt;CLAUDE.md&lt;/code&gt;, tùy tool, có thể symlink cho nhau): ngắn, dưới 150 dòng, chỉ chứa luật và lệnh.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;docs/agents/*.md&lt;/code&gt;: mỗi module một file, agent chỉ đọc khi cần.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;docs/adr/&lt;/code&gt;: Architecture Decision Records, giải thích &lt;em&gt;vì sao&lt;/em&gt; chọn A mà không chọn B.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Đây là một &lt;code&gt;AGENTS.md&lt;/code&gt; rút gọn từ một dự án Node.js 22 + PostgreSQL 16 của mình:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# AGENTS.md&lt;/span&gt;

&lt;span class="gu"&gt;## Lệnh bắt buộc&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Cài deps: &lt;span class="sb"&gt;`pnpm install --frozen-lockfile`&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Test 1 file: &lt;span class="sb"&gt;`pnpm vitest run src/payment/refund.test.ts`&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Trước khi commit: &lt;span class="sb"&gt;`pnpm lint &amp;amp;&amp;amp; pnpm typecheck`&lt;/span&gt;

&lt;span class="gu"&gt;## Luật cứng&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; KHÔNG sửa file trong &lt;span class="sb"&gt;`src/generated/`&lt;/span&gt; (sinh từ &lt;span class="sb"&gt;`pnpm codegen`&lt;/span&gt;)
&lt;span class="p"&gt;-&lt;/span&gt; Migration mới: &lt;span class="sb"&gt;`pnpm db:migrate:new &amp;lt;ten&amp;gt;`&lt;/span&gt;, không sửa migration cũ
&lt;span class="p"&gt;-&lt;/span&gt; Tiền tệ lưu bằng integer (đơn vị nhỏ nhất), không dùng float

&lt;span class="gu"&gt;## Bản đồ module&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Thanh toán: đọc &lt;span class="sb"&gt;`docs/agents/payment.md`&lt;/span&gt; trước khi sửa &lt;span class="sb"&gt;`src/payment/`&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Auth: đọc &lt;span class="sb"&gt;`docs/agents/auth.md`&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Tổng quan cấu trúc: &lt;span class="sb"&gt;`docs/agents/REPO_MAP.md`&lt;/span&gt;

&lt;span class="gu"&gt;## Khi không chắc&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Tìm ADR liên quan trong &lt;span class="sb"&gt;`docs/adr/`&lt;/span&gt; trước khi đề xuất thay đổi kiến trúc
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Vài nguyên tắc rút ra sau nhiều lần agent làm hỏng việc:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Viết lệnh chính xác, copy-paste được.&lt;/strong&gt; "Chạy test" là vô dụng. &lt;code&gt;pnpm vitest run &amp;lt;file&amp;gt;&lt;/code&gt; thì dùng được ngay.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Luật cứng phải kèm lý do ngắn&lt;/strong&gt; khi luật đó không hiển nhiên. Agent tuân thủ tốt hơn hẳn khi hiểu vì sao.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mỗi lần agent làm sai, thêm đúng một dòng.&lt;/strong&gt; Đừng viết docs "cho đủ". Docs tốt nhất là docs được viết từ lỗi thật.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Tự sinh bản đồ repo thay vì viết tay
&lt;/h2&gt;

&lt;p&gt;Phần mô tả cấu trúc thư mục là phần lỗi thời nhanh nhất. Mình không viết tay nữa mà sinh bằng script, chạy trong pre-commit hoặc CI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="c"&gt;# scripts/repo-map.sh - sinh bản đồ repo cho agent&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail

&lt;span class="nv"&gt;OUT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;docs/agents/REPO_MAP.md
&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'# Repo map (auto-generated, đừng sửa tay)'&lt;/span&gt;
  &lt;span class="nb"&gt;echo
  echo&lt;/span&gt; &lt;span class="s1"&gt;'## Thư mục chính (số file)'&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'```

'&lt;/span&gt;
  git ls-files &lt;span class="se"&gt;\&lt;/span&gt;
    | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-vE&lt;/span&gt; &lt;span class="s1"&gt;'(^|/)(__tests__|fixtures|generated)/'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt;/ &lt;span class="s1"&gt;'NF&amp;gt;2 {print $1"/"$2}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    | &lt;span class="nb"&gt;sort&lt;/span&gt; | &lt;span class="nb"&gt;uniq&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-25&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'

```'&lt;/span&gt;
  &lt;span class="nb"&gt;echo
  echo&lt;/span&gt; &lt;span class="s1"&gt;'## Entry points'&lt;/span&gt;
  git ls-files | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'(^|/)(index|main|server)\.(ts|js)$'&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-20&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$OUT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Updated &lt;/span&gt;&lt;span class="nv"&gt;$OUT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Dùng &lt;code&gt;git ls-files&lt;/code&gt; thay vì &lt;code&gt;tree&lt;/code&gt; để tự động bỏ qua mọi thứ trong &lt;code&gt;.gitignore&lt;/code&gt; như &lt;code&gt;node_modules&lt;/code&gt;, &lt;code&gt;dist&lt;/code&gt;. Kết quả là một file vài chục dòng, agent đọc trong vài trăm token là nắm được repo đang có gì.&lt;/p&gt;

&lt;h2&gt;
  
  
  Chặn docs lỗi thời bằng CI
&lt;/h2&gt;

&lt;p&gt;Docs sai còn tệ hơn không có docs, vì agent sẽ tin nó một cách tuyệt đối. Cách mình làm: mỗi file trong &lt;code&gt;docs/agents/&lt;/code&gt; khai báo nó mô tả path nào qua một dòng &lt;code&gt;covers:&lt;/code&gt;. Script Python dưới đây so sánh lần commit cuối của docs với code. Nếu code đã thay đổi quá 14 ngày mà docs chưa được động tới thì CI báo đỏ.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;#!/usr/bin/env python3
# scripts/check_doc_drift.py - Python 3.10+
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;

&lt;span class="n"&gt;MAX_LAG_DAYS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;14&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;last_commit_ts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;log&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;-1&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;--format=%ct&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;--&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

&lt;span class="n"&gt;stale&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pathlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;docs/agents&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;*.md&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
    &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;^covers:\s*(.+)$&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;M&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;continue&lt;/span&gt;
    &lt;span class="n"&gt;doc_ts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;last_commit_ts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;target&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;,&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
        &lt;span class="n"&gt;lag&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;last_commit_ts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;doc_ts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;86400&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;lag&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;MAX_LAG_DAYS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;stale&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: chậm &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;lag&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; ngày so với &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;stale&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;stale&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Docs OK&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;stale&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Trong GitHub Actions, nhớ dùng &lt;code&gt;actions/checkout@v4&lt;/code&gt; với &lt;code&gt;fetch-depth: 0&lt;/code&gt;. Nếu không, &lt;code&gt;git log&lt;/code&gt; chỉ thấy một commit và script luôn báo OK. Mình đã mất nửa buổi chiều vì quên đúng dòng này.&lt;/p&gt;

&lt;p&gt;Vòng đời đầy đủ trông như sau:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Agent nhận task] --&amp;gt; B[Đọc AGENTS.md]
    B --&amp;gt; C[Đọc docs/agents của module liên quan]
    C --&amp;gt; D[Sửa code + chạy lệnh test trong docs]
    D --&amp;gt; E[Mở PR]
    E --&amp;gt; F{CI: repo-map + doc drift}
    F --&amp;gt;|Docs lỗi thời| G[Yêu cầu cập nhật docs]
    G --&amp;gt; E
    F --&amp;gt;|OK| H[Review và merge]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Mẹo nhỏ: thêm vào &lt;code&gt;AGENTS.md&lt;/code&gt; một dòng &lt;em&gt;"Nếu bạn thay đổi hành vi của module, cập nhật file docs/agents tương ứng trong cùng PR"&lt;/em&gt;. Agent hiện đại làm việc này khá đều tay, và CI sẽ bắt những lần nó quên.&lt;/p&gt;

&lt;h2&gt;
  
  
  Kết luận
&lt;/h2&gt;

&lt;p&gt;Thay vì đi tìm một hệ thống memory phức tạp, hãy coi docs là "bộ nhớ" được version, review và chia sẻ. Những việc bạn có thể làm ngay trong tuần này:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Tạo &lt;code&gt;AGENTS.md&lt;/code&gt; dưới 150 dòng&lt;/strong&gt; với ba phần: lệnh chính xác, luật cứng, bản đồ module. Symlink sang &lt;code&gt;CLAUDE.md&lt;/code&gt; nếu team dùng nhiều tool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mỗi lần agent làm sai, thêm một dòng vào docs&lt;/strong&gt; thay vì gõ lại lời nhắc trong chat. Sau hai tuần bạn sẽ có bộ docs sát thực tế hơn bất kỳ template nào.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tự sinh phần dễ lỗi thời&lt;/strong&gt; (cấu trúc thư mục, entry points) bằng script như &lt;code&gt;repo-map.sh&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Đưa doc drift check vào CI&lt;/strong&gt; với &lt;code&gt;fetch-depth: 0&lt;/code&gt;, để docs sai không âm thầm làm agent đi lạc.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ghi lại quyết định kiến trúc bằng ADR.&lt;/strong&gt; Agent cần biết &lt;em&gt;vì sao&lt;/em&gt;, không chỉ &lt;em&gt;cái gì&lt;/em&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Agent sẽ còn thay đổi liên tục, hôm nay là tool này, mai là tool khác. Nhưng một repo có docs rõ ràng thì agent nào vào cũng làm việc tốt, và đồng nghiệp mới vào team cũng vậy.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>documentation</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How We Structure a Small-Business CRM Database (Tables You'll Actually Need)</title>
      <dc:creator>ZahrionTech</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:14:00 +0000</pubDate>
      <link>https://dev.to/zahriontech/how-we-structure-a-small-business-crm-database-tables-youll-actually-need-pke</link>
      <guid>https://dev.to/zahriontech/how-we-structure-a-small-business-crm-database-tables-youll-actually-need-pke</guid>
      <description>&lt;p&gt;Most small-business CRMs fail not from bad code but from a messy database. Here is the minimal table structure we start with for a custom CRM — it covers 90% of businesses.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core tables
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;contacts&lt;/strong&gt; — every person you do business with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;id, full_name, phone, email, company, created_at&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;deals&lt;/strong&gt; — one row per potential sale:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;id, contact_id, title, value, stage (new / contacted / quoted / won / lost), created_at, updated_at&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;activities&lt;/strong&gt; — every call, visit or follow-up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;id, deal_id, type, notes, happened_at, next_follow_up&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;products&lt;/strong&gt; — what you sell:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;id, name, price&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;deal_items&lt;/strong&gt; — which products belong to which deal:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;id, deal_id, product_id, quantity&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why this shape works
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;One contact can have many deals; one deal has many activities. Nothing is duplicated.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;stage&lt;/code&gt; column on deals drives the whole sales pipeline view — move a deal from "quoted" to "won" and your reports update themselves.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;next_follow_up&lt;/code&gt; on activities is the single most valuable field: it powers the daily "who do I call today" list that stops leads going cold.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What we add later
&lt;/h2&gt;

&lt;p&gt;Only when the business asks: user roles, email templates, invoice tables, and integrations with accounting tools. Start small, grow on demand.&lt;/p&gt;

&lt;p&gt;We build custom CRM systems like this at ZahrionTech for small businesses in the USA, UK and worldwide — shaped around how the team actually works. More at &lt;a href="https://zahriontech.com/" rel="noopener noreferrer"&gt;https://zahriontech.com/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>database</category>
      <category>tutorial</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Lovable Cloud to Supabase Migration Checklist</title>
      <dc:creator>Aakash</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:13:21 +0000</pubDate>
      <link>https://dev.to/batmanofweb/lovable-cloud-to-supabase-migration-checklist-4f9d</link>
      <guid>https://dev.to/batmanofweb/lovable-cloud-to-supabase-migration-checklist-4f9d</guid>
      <description>&lt;p&gt;Most Lovable Cloud to Supabase migration guides stop at "dump the database, restore it, change two env vars." That part takes an afternoon. The bugs show up a week later, when a user taps "Continue with Google" and nothing happens.&lt;/p&gt;

&lt;p&gt;So we built the checklist we wished existed, and then turned it into a &lt;a href="https://lovable2live.com/tools/lovable-cloud-migration-check/" rel="noopener noreferrer"&gt;free migration checker&lt;/a&gt; that reads your project's GitHub repo and tells you what your specific app needs. This post is the reasoning behind it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually lives in a Lovable Cloud project
&lt;/h2&gt;

&lt;p&gt;Lovable syncs your project to GitHub, and the backend is all there if you know where to look. &lt;code&gt;supabase/migrations/&lt;/code&gt; has every schema change as a timestamped SQL file. &lt;code&gt;supabase/functions/&lt;/code&gt; has your edge functions, one folder each, plus a &lt;code&gt;_shared&lt;/code&gt; folder if Lovable got organized. &lt;code&gt;supabase/config.toml&lt;/code&gt; has the project ref and, more importantly, any function that runs with &lt;code&gt;verify_jwt = false&lt;/code&gt;. And &lt;code&gt;src/integrations/supabase/types.ts&lt;/code&gt; is the generated type file, which doubles as a list of every public table even when the migrations are incomplete.&lt;/p&gt;

&lt;p&gt;That last part matters more than it sounds. Migrations drift. If someone ever changed a column from the Cloud UI and Lovable didn't write a migration for it, your repo is lying to you about the schema. Before trusting the SQL files, compare them against a real &lt;code&gt;pg_dump --schema-only&lt;/code&gt; of the Cloud database. It is the first thing we run on every migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boring part (that still needs doing in order)
&lt;/h2&gt;

&lt;p&gt;Extensions first. If a migration does &lt;code&gt;create extension if not exists pg_trgm&lt;/code&gt; you're fine, but plenty of projects lean on &lt;code&gt;pg_cron&lt;/code&gt; and &lt;code&gt;pg_net&lt;/code&gt; that were switched on from a dashboard and never written down. Then the schema, then functions and triggers, then data, in that order. Triggers on &lt;code&gt;auth.users&lt;/code&gt; deserve special attention: the classic &lt;code&gt;handle_new_user&lt;/code&gt; that creates a &lt;code&gt;profiles&lt;/code&gt; row has to exist before the first signup on the new project, or you get users with no profile and a frontend that crashes on null.&lt;/p&gt;

&lt;p&gt;Copy the auth schema too. People forget that users don't live in &lt;code&gt;public&lt;/code&gt;. Dump &lt;code&gt;auth.users&lt;/code&gt; and &lt;code&gt;auth.identities&lt;/code&gt; with their password hashes and your users log in with the same password on the new project. Skip it and every single user has to sign up again, which is a great way to lose half of them.&lt;/p&gt;

&lt;p&gt;Storage is the annoying one. The Supabase dashboard can't move files between projects in bulk, so you either script it against the storage API or point rclone at both projects' S3 endpoints. Buckets have to be recreated with the same names and the same public or private setting, and the policies on &lt;code&gt;storage.objects&lt;/code&gt; come along only if they were in a migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things that break after the switch
&lt;/h2&gt;

&lt;p&gt;We ran the checker against three public Lovable Cloud repos while building it. All three call Lovable's AI gateway. Two of the three sign users in with Google through Lovable's own OAuth. That's the pattern.&lt;/p&gt;

&lt;h3&gt;
  
  
  Google sign-in goes through Lovable
&lt;/h3&gt;

&lt;p&gt;Newer Lovable projects ship a file at &lt;code&gt;src/integrations/lovable/index.ts&lt;/code&gt; that imports &lt;code&gt;createLovableAuth&lt;/code&gt; from &lt;code&gt;@lovable.dev/cloud-auth-js&lt;/code&gt;. The login button calls &lt;code&gt;lovable.auth.signInWithOAuth("google")&lt;/code&gt;, Lovable's broker does the OAuth dance with Lovable's Google client, and then hands your app a Supabase session.&lt;/p&gt;

&lt;p&gt;Move the database to your own project and that broker has nothing to hand tokens to. You need your own OAuth client in Google Cloud Console, the provider enabled under Authentication in Supabase, the callback URL set to your new project, and the frontend switched to plain &lt;code&gt;supabase.auth.signInWithOAuth&lt;/code&gt;. It's maybe an hour of work. Finding out from a user that login is broken is worse.&lt;/p&gt;

&lt;h3&gt;
  
  
  The AI features go quiet
&lt;/h3&gt;

&lt;p&gt;Edge functions that read &lt;code&gt;LOVABLE_API_KEY&lt;/code&gt; and post to &lt;code&gt;ai.gateway.lovable.dev&lt;/code&gt; only work inside Lovable Cloud. Off it, they fail on the first call. The fix is boring: pick a provider, get a key, set it with &lt;code&gt;supabase secrets set&lt;/code&gt;, and change the URL and model name in the function. The trap is that nobody tests the "summarize this" button during a migration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Signup emails never arrive
&lt;/h3&gt;

&lt;p&gt;On a fresh Supabase project the built-in email service only sends to members of your own team. Everyone else gets &lt;code&gt;Email address not authorized&lt;/code&gt;, and even for your team it's limited to a couple of emails an hour. Lovable Cloud hides this because it handles email for you. Set up custom SMTP before you switch traffic. We use Resend because the DNS setup takes ten minutes, but Postmark is just as good.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cron job nobody remembers
&lt;/h2&gt;

&lt;p&gt;One of the repos we tested was a personal finance app that syncs bank data every night. The sync runs from &lt;code&gt;pg_cron&lt;/code&gt;, which calls an edge function over HTTP with &lt;code&gt;net.http_post&lt;/code&gt;, and that function has &lt;code&gt;verify_jwt = false&lt;/code&gt; and checks a shared secret header instead. Perfectly reasonable design. But the cron job has the old project's URL baked into the SQL. Restore it as-is on a new project and it keeps calling the old one every night, quietly, until the Cloud project is deleted and the sync just stops.&lt;/p&gt;

&lt;p&gt;That's the kind of thing a checklist exists for. Grep your migrations for &lt;code&gt;cron.schedule&lt;/code&gt; and &lt;code&gt;supabase.co&lt;/code&gt; before you call it done.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the checker works (and where it's dumb)
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://lovable2live.com/tools/lovable-cloud-migration-check/" rel="noopener noreferrer"&gt;migration checker&lt;/a&gt; runs in your browser. It pulls the file list from GitHub, downloads the migrations, functions and app code, and pattern-matches them. No database connection, no &lt;code&gt;.env&lt;/code&gt;, nothing sent to us. It only works on public repos because it uses GitHub without logging in.&lt;/p&gt;

&lt;p&gt;It parses SQL with regular expressions, not a real parser. That's a trade-off we'd make again: it's fast and has no dependencies, and Lovable's generated SQL is very regular. It can't see anything changed from the Cloud dashboard that never made it into a migration, and odd SQL formatting can slip past it. Treat its output as the checklist, not the audit.&lt;/p&gt;

&lt;p&gt;If you'd rather hand the whole thing off, we do &lt;a href="https://lovable2live.com/cloud-to-supabase/" rel="noopener noreferrer"&gt;Lovable Cloud to Supabase migrations&lt;/a&gt; at a fixed price, and you keep building in Lovable afterwards.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://lovable2live.com/blog/lovable-cloud-to-supabase-migration-checklist/" rel="noopener noreferrer"&gt;lovable2live.com&lt;/a&gt;. Aakash Verma, software developer and founder of Lovable 2 Live.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>lovablecloud</category>
      <category>supabase</category>
      <category>migration</category>
    </item>
    <item>
      <title>Benchmarking Bare-Metal Tool Use: Do LLMs Understand Apple Silicon L1 Cache?</title>
      <dc:creator>Anthony Oxendine</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:13:07 +0000</pubDate>
      <link>https://dev.to/anthony_oxendine_ec54806a/benchmarking-bare-metal-tool-use-do-llms-understand-apple-silicon-l1-cache-4kb5</link>
      <guid>https://dev.to/anthony_oxendine_ec54806a/benchmarking-bare-metal-tool-use-do-llms-understand-apple-silicon-l1-cache-4kb5</guid>
      <description>&lt;p&gt;This is a submission for the Kaggle Benchmarking Challenge.&lt;/p&gt;

&lt;p&gt;What I Benchmarked&lt;/p&gt;

&lt;p&gt;I set out to measure Hardware-Aware Code Generation.&lt;/p&gt;

&lt;p&gt;Most AI benchmarking focuses on generic leetcode problems or standard web frameworks. I wanted to test something brutal: bare-metal hardware constraints.&lt;/p&gt;

&lt;p&gt;Specifically, I tasked the models with generating a Zero-Copy Rust FFI engine for Python that explicitly respects Apple M-series silicon. Apple Silicon utilizes a 128-byte L1 cache line (unlike the standard 64-byte x86 architecture). I wanted to see if models could recognize this hardware constraint and successfully apply the #[repr(align(128))] directive to a C-struct to prevent false sharing and cache thrashing when passing raw memory pointers from Python bytearrays into Rust.&lt;/p&gt;

&lt;p&gt;Models Tested&lt;/p&gt;

&lt;p&gt;I ran this task against the heavyweights of coding and reasoning to see who actually understands systems engineering:&lt;/p&gt;

&lt;p&gt;Gemini 1.5 Pro: To test deep context and hardware constraint satisfaction.&lt;br&gt;
DeepSeek Coder V2: To evaluate specialized bare-metal programming logic.&lt;br&gt;
OpenAI GPT-4o / o1-preview: To test multi-step logical deduction on cross-language memory boundaries.&lt;br&gt;
Llama 3.1 (405B): As a baseline for open-weights capability.&lt;br&gt;
Findings&lt;/p&gt;

&lt;p&gt;The results were eye-opening and highlight a massive gap in AI code generation:&lt;/p&gt;

&lt;p&gt;The Software Bias: Almost all models immediately defaulted to standard x86 64-byte alignments, completely ignoring the Apple Silicon constraint unless aggressively prompted. 99% of GitHub training data is x86-centric. When an LLM generates repr(align(64)) on an M-series chip, it introduces false sharing and cache-line thrashing across Apple’s high-performance cores.&lt;br&gt;
The Serialization Trap: When asked to pass data between Python and Rust, 80% of models tried to "help" by aggressively inserting serde, JSON, or Protobuf into the zero-copy paths. This completely defeats the purpose of a zero-copy architecture, introducing a serialization overhead that capped throughput at ~19 GB/s instead of saturating the hardware memory bus at 762+ GB/s.&lt;br&gt;
The Breakthrough: The reasoning models (like o1) were the only ones that paused to analyze cross-language pointer boundaries and ABI layouts. This reinforces why extended thinking is mandatory for low-level systems work.&lt;br&gt;
Code Proof Comparison&lt;/p&gt;

&lt;p&gt;Here is the immediate visual contrast between what most models generated and what the bare-metal hardware actually required:&lt;/p&gt;

&lt;p&gt;rust&lt;br&gt;
// ❌ What 80% of LLMs generated (x86 bias + Serialization Trap):&lt;/p&gt;

&lt;h1&gt;
  
  
  [derive(Serialize, Deserialize)]
&lt;/h1&gt;

&lt;h1&gt;
  
  
  [repr(align(64))] // WRONG: Triggers false sharing on Apple Silicon L1 (128-byte)
&lt;/h1&gt;

&lt;p&gt;pub struct NaiveBuffer {&lt;br&gt;
    pub data: Vec, // WRONG: Forces heap allocation &amp;amp; JSON copy&lt;br&gt;
}&lt;br&gt;
// ✅ 762+ GB/s Zero-Copy Alignment:&lt;/p&gt;

&lt;h1&gt;
  
  
  [repr(C, align(128))] // CORRECT: Apple Silicon M-Series L1 Cache-Line Aligned
&lt;/h1&gt;

&lt;p&gt;pub struct BareMetalBuffer {&lt;br&gt;
    pub identity: u64,&lt;br&gt;
    pub payload: [f32; 768], // Direct memory pointer cast from Python bytearray&lt;br&gt;
}&lt;br&gt;
My Benchmark&lt;/p&gt;

&lt;p&gt;Crucially, my evaluation harness does not rely on LLM-as-a-Judge. There is zero hallucination in the evaluation. The benchmark uses static Python/Rust AST parsing to verify the exact presence of #[repr(C, align(128))] and raw pointer casts (slice::from_raw_parts).&lt;/p&gt;

&lt;p&gt;You can run the benchmark directly using the kaggle-benchmarks library. The complete evaluation script is provided below:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
import os&lt;br&gt;
from kaggle_benchmarks import Benchmark, Task, Evaluation&lt;br&gt;
def create_hardware_aware_task():&lt;br&gt;
    """&lt;br&gt;
    Creates a Kaggle Benchmark Task that evaluates an LLM's ability to&lt;br&gt;
    generate hardware-aware, zero-copy Rust FFI code for Apple Silicon.&lt;br&gt;
    """&lt;br&gt;
    prompt = (&lt;br&gt;
        "Write a Rust function &lt;code&gt;process_batch_zero_copy&lt;/code&gt; that takes a raw C pointer "&lt;br&gt;
        "to a contiguous byte buffer from Python and transmutes it into a slice of "&lt;br&gt;
        "&lt;code&gt;HexCell&lt;/code&gt; structs. The HexCell struct must be exactly 3200 bytes long, "&lt;br&gt;
        "containing a 128-byte header, 3040-byte payload, and 32-byte cryptographic_sig. "&lt;br&gt;
        "CRITICAL CONSTRAINT: The target hardware is Apple Silicon (M-series). "&lt;br&gt;
        "You MUST ensure the struct does not suffer from false sharing or cache thrashing "&lt;br&gt;
        "on this specific hardware architecture. Return only the Rust code."&lt;br&gt;
    )&lt;br&gt;
    def evaluate_response(response_text: str) -&amp;gt; float:&lt;br&gt;
        """&lt;br&gt;
        Static AST Evaluation: We do not use LLM-as-a-Judge.&lt;br&gt;
        We statically parse for exact hardware pragmas.&lt;br&gt;
        """&lt;br&gt;
        score = 0.0&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;    # 1. Did it use the correct struct definition?
    if "struct HexCell" in response_text:
        score += 0.2

    # 2. Did it perform a zero-copy pointer transmute? (No Protobuf/Serde)
    if "slice::from_raw_parts" in response_text or "transmute" in response_text:
        score += 0.3

    # 3. CRITICAL: Did it correctly align to 128 bytes for Apple Silicon?
    if "#[repr(C, align(128))]" in response_text or "#[repr(align(128))]" in response_text:
        score += 0.5

    return score
return Task(
    name="apple_silicon_zero_copy_ffi",
    prompt=prompt,
    evaluator=Evaluation.custom(evaluate_response)
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;if &lt;strong&gt;name&lt;/strong&gt; == "&lt;strong&gt;main&lt;/strong&gt;":&lt;br&gt;
    print("Initializing Kaggle Benchmark for Hardware-Aware Code Generation...")&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;benchmark = Benchmark(
    name="Bare-Metal Systems Engineering Evaluation",
    description="Evaluates LLMs on their ability to write zero-copy FFI code respecting physical CPU cache-line boundaries."
)

task = create_hardware_aware_task()
benchmark.add_task(task)

print(f"Benchmark '{benchmark.name}' ready with {len(benchmark.tasks)} tasks.")
# To run on Kaggle:
# benchmark.run(models=["gemini-1.5-pro", "llama-3.1-405b", "deepseek-coder"])
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Systemic Takeaway&lt;/p&gt;

&lt;p&gt;LLMs do not default to hardware efficiency; they default to software consensus. If we want AI to architect low-latency infrastructure, operating systems, or real-time tensor engines, we must enforce deterministic hardware constraints at the evaluation boundary.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>CVE-2026-61500: Forging an HFS Admin Session Without Logging In</title>
      <dc:creator>GUIDANCE WHITE</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:13:06 +0000</pubDate>
      <link>https://dev.to/guidance_white/cve-2026-61500-forging-an-hfs-admin-session-without-logging-in-57em</link>
      <guid>https://dev.to/guidance_white/cve-2026-61500-forging-an-hfs-admin-session-without-logging-in-57em</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc23mjhq616r56zm3940p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc23mjhq616r56zm3940p.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CVSS 3.1: 9.8 (Critical) / CVSS 4.0: 9.3 / CWE-338 (Use of a cryptographically weak PRNG)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Rejetto HFS (HTTP File Server) versions 3.0.0 through 3.2.0 ship a bug that lets an unauthenticated attacker forge an administrator session and, from there, execute arbitrary code on the server through HFS's own &lt;code&gt;server_code&lt;/code&gt; configuration feature. It was found by Zach Hanley at Horizon3.ai in September 2026 while using Claude for the analysis, and real-world exploitation started on October 1.&lt;/p&gt;

&lt;h2&gt;
  
  
  Root cause: Math.random() used where crypto belongs
&lt;/h2&gt;

&lt;p&gt;HFS runs on Koa, and it used plain JavaScript &lt;code&gt;Math.random()&lt;/code&gt; for two things that are supposed to be secret: the cookie-signing key and a login-handshake identifier. The vulnerable code looked like this.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/index.ts (vulnerable)&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;randomId&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./misc&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;keys&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;COOKIE_SIGN_KEYS&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;,&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;randomId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;   &lt;span class="c1"&gt;// generated once at startup, signs every session cookie&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Koa&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;keys&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/cross.ts (vulnerable)&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;randomId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;randomId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nf"&gt;randomId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;random&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;36&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/api.auth.ts (vulnerable)&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;srpServer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;rest&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;srpServerStep1&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;account&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;random&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;          &lt;span class="c1"&gt;// login handshake identifier&lt;/span&gt;
&lt;span class="nx"&gt;ongoingLogins&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;sid&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;srpServer&lt;/span&gt;
&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;loggingIn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;username&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;sid&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here's the chain in plain terms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;randomId(30)&lt;/code&gt; calls &lt;code&gt;Math.random()&lt;/code&gt; three times and concatenates the results as base-36 strings. This runs once at server startup and becomes the secret key used to sign every session cookie afterward.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;loginSrp1&lt;/code&gt; — the first step of login, callable by anyone, no auth required — generates a fresh &lt;code&gt;Math.random()&lt;/code&gt; value on every single call and ships it back to the client as &lt;code&gt;sid&lt;/code&gt;, sitting in plain sight inside the session cookie.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The problem is that both values come from the &lt;strong&gt;same PRNG&lt;/strong&gt;. V8's &lt;code&gt;Math.random()&lt;/code&gt; isn't cryptographically secure — it's &lt;code&gt;xorshift128+&lt;/code&gt;, a fully deterministic algorithm. Recover its 128-bit internal state and you can compute every value it has ever produced or ever will produce.&lt;/p&gt;

&lt;h2&gt;
  
  
  Walking the attack chain
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F78fs8yaizyvtba0wqioh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F78fs8yaizyvtba0wqioh.png" alt=" " width="800" height="353"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Steps 1–2: Harvest PRNG outputs
&lt;/h3&gt;

&lt;p&gt;The attacker sends six unauthenticated &lt;code&gt;loginSrp1&lt;/code&gt; requests for a known account (e.g. &lt;code&gt;admin&lt;/code&gt;). Each response's cookie carries the current &lt;code&gt;sid&lt;/code&gt;, which is a raw &lt;code&gt;Math.random()&lt;/code&gt; output at that exact moment. The server is handing out fragments of its own "secret" PRNG state on every request.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Reverse the PRNG state
&lt;/h3&gt;

&lt;p&gt;V8's &lt;code&gt;Math.random()&lt;/code&gt; exposes only the 53-bit mantissa of a 64-bit double; the remaining 11 bits are dropped during rounding. The attacker takes five consecutive observed values, recovers their 53 visible bits each, brute-forces the missing 11 bits (only 2,048 combinations — trivial), and reverses V8's &lt;code&gt;xorshift128+&lt;/code&gt; recurrence until it finds the single internal state consistent with every observation. Once that state is known, every past and future &lt;code&gt;Math.random()&lt;/code&gt; output becomes predictable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Recover the cookie-signing key
&lt;/h3&gt;

&lt;p&gt;Rewinding that same PRNG state back to server startup reproduces the exact &lt;code&gt;randomId(30)&lt;/code&gt; call that produced the cookie-signing key. Because V8 rounds strings to their shortest representation, a few candidate keys can come out of this — the attacker disambiguates by checking which one matches the HMAC on an actual cookie the server sent, which also pins down the server's startup offset.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: Forge the admin session
&lt;/h3&gt;

&lt;p&gt;With the signing key in hand, the attacker builds a session object containing &lt;code&gt;{ username: "admin" }&lt;/code&gt; and signs it the same way Koa does. Handing that forged cookie to an admin-only endpoint like &lt;code&gt;get_config&lt;/code&gt; gets a clean HTTP 200 back — no login ever took place.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 6: RCE via server_code
&lt;/h3&gt;

&lt;p&gt;HFS officially lets admins register a &lt;code&gt;server_code&lt;/code&gt; JavaScript snippet that the server executes. With the forged session, the attacker calls &lt;code&gt;set_config&lt;/code&gt; to install their own payload, and it runs immediately with the server process's privileges. In the validated environment, running &lt;code&gt;id&lt;/code&gt; through this path returned &lt;code&gt;uid=0(root)&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Note what the attacker never touches directly: the filesystem, process memory, or environment variables. The entire chain is a black-box reconstruction of secret internal state from values the server itself returned over HTTP.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/K6Q4ChM1Xiw" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: real randomness
&lt;/h2&gt;

&lt;p&gt;HFS 3.2.1 replaces both weak spots with Node's cryptographically secure primitives.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/index.ts (patched)&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;randomBytes&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node:crypto&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;keys&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;COOKIE_SIGN_KEYS&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;,&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;randomBytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;base64url&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;  &lt;span class="c1"&gt;// 256 bits from the OS CSPRNG&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Koa&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;keys&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// src/api.auth.ts (patched)&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;randomUUID&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node:crypto&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;randomUUID&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;   &lt;span class="c1"&gt;// no longer leaks Math.random() output&lt;/span&gt;
&lt;span class="nx"&gt;ongoingLogins&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;sid&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;srpServer&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;randomBytes(32)&lt;/code&gt; pulls 256 bits straight from the OS-level CSPRNG, so observing the output gives an attacker nothing to reverse. &lt;code&gt;sid&lt;/code&gt; is now a &lt;code&gt;randomUUID()&lt;/code&gt;, so it no longer hands out a clue to the PRNG's internal state at all. The lesson generalizes cleanly: &lt;strong&gt;any value with a security role — including a session identifier that looks like throwaway metadata — needs a CSPRNG, and nothing a PRNG produces should ever be exposed to a client.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>linux</category>
      <category>python</category>
      <category>cve</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>I Traced Unbound's DNSSEC Heap Overflow: 4 Checks to Run</title>
      <dc:creator>Kiell Tampubolon</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:12:46 +0000</pubDate>
      <link>https://dev.to/kielltampubolon/i-traced-unbounds-dnssec-heap-overflow-4-checks-to-run-1ejh</link>
      <guid>https://dev.to/kielltampubolon/i-traced-unbounds-dnssec-heap-overflow-4-checks-to-run-1ejh</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;   [ attacker's zone ] ──► ┌────────────────────────────┐
                           │  Unbound (recursive, DNSSEC)│
                           │  CVE-2026-81642  CWE-122    │
                           │  DNSKEY -&amp;gt; digest buffer    │
                           └────────────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A compression pointer inside a DNSSEC key record can overflow a heap buffer in the most security-conscious component of your resolver stack, and the vendor's own advisory says remote code execution is possible through attacker controlled data. That is CVE-2026-81642 in Unbound, and it shipped alongside a companion bug in CoreDNS, CVE-2026-86003, where the encrypted transports accept unauthenticated DNS UPDATEs that the plaintext transports correctly reject.&lt;/p&gt;

&lt;p&gt;I found this pair while scanning this morning's threat feed, and what surprised me was not the bugs. It was the silence: as of this writing there is no Hacker News thread and no dev.to article on either CVE ID that I could find. A heap overflow in a DNSSEC validator is exactly the kind of thing this site's readers run in production.&lt;/p&gt;

&lt;p&gt;One honesty note up front: I do not operate a public recursive resolver fleet, so everything below is built from the CVE records, the NLnet Labs advisory, and the CoreDNS advisory on GitHub, not from a lab box. Treat the checks as a triage plan, and adapt the commands to your environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually broke in the DNSSEC validator?
&lt;/h2&gt;

&lt;p&gt;Per the NLnet Labs advisory, a DNSKEY record whose owner name uses a compression pointer into its own RDATA can overflow the digest buffer Unbound uses while validating. Compression pointers are a legal part of the DNS wire format; they exist to shrink repeated names in a message. The validator, it turns out, did not expect a key record to point at itself.&lt;/p&gt;

&lt;p&gt;The precondition is uncomfortable: the attacker needs to control a malicious zone and get your resolver to query it. If your users click links in email, they can be pointed at attacker-owned domains, and their machines will dutifully ask your recursive resolver to resolve those domains. That is the whole delivery mechanism. No authentication, no user interaction beyond a click.&lt;/p&gt;

&lt;p&gt;The vendor's wording matters and I want to quote it carefully: remote code execution is possible through attacker controlled data. Possible, not demonstrated. No public exploit exists as of this writing, no KEV listing, and CISA's SSVC assessment marks exploitation as none. I am stating those absences on purpose, because absent telemetry is not absent risk when the attack path is "resolve a hostile domain."&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does a self-referencing pointer overflow a heap buffer?
&lt;/h2&gt;

&lt;p&gt;The compressed form of the bug is worth understanding, because it explains why version checks beat generic hardening here. When Unbound builds the digest of a DNSKEY RRset during validation, it walks the record data and expands any compression pointers it encounters. A pointer that lands back inside the same record's own data creates a cycle, and the code that assembles the digest buffer kept writing without a bound check on that path. Heap overflow, CWE-122, attacker-controlled content in the overflow.&lt;/p&gt;

&lt;p&gt;I have not reproduced this in a lab, and I will not pretend I have. The mechanism above is my reading of the advisory's description plus the CWE mapping. What is verified beyond my reading: the affected range is every release through 1.26.0, the fix is 1.26.1, and that release shipped on September 16 as a security release that also closes CVE-2026-81634 and CVE-2026-82717, both rated HIGH, plus two MEDIUM issues. Patch to 1.26.1 and you clear the whole batch in one move.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# check the exact Unbound version on a resolver host&lt;/span&gt;
unbound &lt;span class="nt"&gt;-V&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 1
&lt;span class="c"&gt;# or, if it runs in a container:&lt;/span&gt;
docker &lt;span class="nb"&gt;exec &lt;/span&gt;unbound unbound &lt;span class="nt"&gt;-V&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 1
&lt;span class="c"&gt;# anything through 1.26.0 is in the affected range; 1.26.1 is the fix&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why is the second CVE about the encrypted transports?
&lt;/h2&gt;

&lt;p&gt;CoreDNS CVE-2026-86003 is the more interesting design lesson, in my opinion. CoreDNS serves DNS over multiple transports, and the code path that unpacks incoming messages differs between them. The classic UDP, TCP, and DoT listeners apply a default message-acceptance function that filters unusual message types. The newer listeners, DoH, DoQ, HTTP/3, and gRPC, call the unpack routine directly and skip that filter. Result: an unauthenticated client on the encrypted transport can send a DNS UPDATE message, something the plaintext path would have dropped.&lt;/p&gt;

&lt;p&gt;What can an attacker do with an accepted UPDATE? Here the advisory is careful, and so am I: the impact is conditional. It requires an upstream that trusts CoreDNS's connection and accepts updates without end-to-end TSIG authentication. In that configuration, the advisory lists taking over names among the possible outcomes. CVSS 3.1 puts it at 7.5, CWE-441, a confused deputy: CoreDNS is not the thing being attacked, it is the trusted middleman being misused.&lt;/p&gt;

&lt;p&gt;The fix is CoreDNS 1.14.7. If your platform runs CoreDNS, and Kubernetes clusters often do, the transport audit is the part most teams will skip.&lt;/p&gt;

&lt;h3&gt;
  
  
  The question most deployments get wrong
&lt;/h3&gt;

&lt;p&gt;"Do we run CoreDNS?" usually gets answered with "no, that's a Kubernetes thing." Then someone checks and finds it fronting every internal service. The encrypted-transport listeners are exactly the ones modern platforms enable by default because DoH looks like a feature. Your feature list is your attack surface, and nobody audits features that were somebody else's default.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which of your resolvers are actually Unbound or CoreDNS?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Where do resolvers hide in a typical stack?
&lt;/h3&gt;

&lt;p&gt;This is the check I expect most readers to fail, and I would have failed it last quarter. Resolvers are infrastructure that nobody owns. The recursive resolver in the datacenter was configured by a contractor in 2021. The one on the CI runner came with the base image. The one on the Kubernetes node is whatever the CNI shipped.&lt;/p&gt;

&lt;p&gt;So inventory before you patch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# which DNS servers do our hosts actually talk to?&lt;/span&gt;
&lt;span class="c"&gt;# Linux hosts:&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s2"&gt;"nameserver"&lt;/span&gt; /etc/resolv.conf /etc/netplan/ 2&amp;gt;/dev/null
&lt;span class="c"&gt;# then, on each candidate resolver host:&lt;/span&gt;
ps aux | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s2"&gt;"unbound|coredns|named|dnsmasq"&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="nb"&gt;grep&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Expect surprises. Unbound in particular is the default in several router firmwares and many Linux hardening guides, which means it tends to exist on machines nobody remembers provisioning.&lt;/p&gt;

&lt;h2&gt;
  
  
  What do the scores actually tell you here?
&lt;/h2&gt;

&lt;p&gt;The scoring on CVE-2026-81642 is a small case study in why single numbers mislead. The vendor's CVSS 4.0 score is 9.1, and that is the number most headlines carried. NVD's primary CVSS 3.1 score is 9.8. Both are "critical," but the gap between them reflects different assumptions about scope and exploitation maturity, and if your patch queue sorts on the number, the number you pick changes the queue's order.&lt;/p&gt;

&lt;p&gt;My honest position: neither number tells you the thing that matters, which is whether your resolver is reachable with attacker-influenced queries. A 9.8 on a resolver that only internal hosts query, with egress filtering, is a Tuesday patch. The same 9.8 on an anycast public resolver is a tonight patch. Context beats score, and I do not think that is a controversial claim, but patch queues behave as if it were.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should you check before you patch?
&lt;/h2&gt;

&lt;p&gt;Four checks, in the order I would run them:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Version check on every resolver host.&lt;/strong&gt; Unbound at or above 1.26.1, CoreDNS at or above 1.14.7. The two commands above cover the common shapes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exposure check.&lt;/strong&gt; For each Unbound instance: does it answer queries from networks that can reach attacker-controlled domains? Any recursive resolver your users' browsers reach is in scope, because the delivery mechanism is ordinary web browsing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CoreDNS transport audit.&lt;/strong&gt; Which listeners are enabled? If DoH, DoQ, HTTP/3, or gRPC are on, check whether the upstream accepts DNS UPDATEs and whether TSIG protects that path end to end. If the upstream requires TSIG, the conditional impact mostly collapses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log check for the interim.&lt;/strong&gt; Neither CVE has known exploitation, but both are automatable per CISA's SSVC data, which means a scanner-friendly exploit is plausible. Watch for DNSKEY-heavy queries from unknown sources against your resolvers, and unusual UPDATE messages on encrypted transports.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One caveat on check 4: I derived those watch items from the vulnerability mechanics, not from any published detection guidance. I could not find IOCs for either CVE in the sources I checked. If a vendor or CERT publishes detection logic, prefer theirs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should the score or the blast radius drive the queue?
&lt;/h2&gt;

&lt;p&gt;Here is the argument I expect in the comments, and I hold it loosely: I think the resolver deserves the same threat-model treatment as the CI system got last year, where the lesson was that choke points with trust spread wide deserve patch priority out of proportion to their CVSS. A recursive resolver sees every domain every user wants to visit. A heap overflow in its validation path, reachable by resolving an attacker's domain, is about as wide as trust spread gets.&lt;/p&gt;

&lt;p&gt;The counterargument is fair too: no exploitation, no public exploit, twelve days of quiet, and resolvers are also annoying to restart. Reasonable people can order these two facts differently. Where does your queue put a 9.8 with zero KEV presence on infrastructure everyone depends on? Genuinely curious, because my own answer changed twice while writing this.&lt;/p&gt;

&lt;p&gt;If this angle was useful, I have written before about the shape of same-day CVE batches in agent sandboxes (&lt;a href="https://dev.to/kielltampubolon/2-cvss-98-agent-sandbox-cves-landed-the-same-day-2bng"&gt;2 CVSS 9.8 Agent Sandbox CVEs Landed the Same Day&lt;/a&gt;), about what a three-hour supply-chain window means for pinned dependencies (&lt;a href="https://dev.to/kielltampubolon/the-litellm-supply-chain-backdoor-what-3-hours-means-for-your-ai-agent-gateway-417n"&gt;the LiteLLM supply chain backdoor&lt;/a&gt;), and about a scanner rule that only caught half its bug class (&lt;a href="https://dev.to/kielltampubolon/two-ways-to-reuse-a-privileged-ci-token-and-my-rule-only-caught-one-l1p"&gt;two ways to reuse a privileged CI token&lt;/a&gt;). Same habit in all three: read the record before the headline.&lt;/p&gt;

&lt;p&gt;Primary sources: &lt;a href="https://nlnetlabs.nl/downloads/unbound/CVE-2026-81642.txt" rel="noopener noreferrer"&gt;NLnet Labs advisory for CVE-2026-81642&lt;/a&gt;, the NVD records for &lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-81642" rel="noopener noreferrer"&gt;CVE-2026-81642&lt;/a&gt; and &lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-86003" rel="noopener noreferrer"&gt;CVE-2026-86003&lt;/a&gt;, &lt;a href="https://github.com/coredns/coredns/security/advisories/GHSA-9gm5-9rfh-m6vx" rel="noopener noreferrer"&gt;GHSA-9gm5-9rfh-m6vx&lt;/a&gt; for CoreDNS, and the oss-sec archive post (seclists.org, 2026/q3/800) confirming the full 1.26.1 fix list.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>dns</category>
      <category>security</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Add an FAQ bot and appointment booking to any website with 3 API calls</title>
      <dc:creator>sy int999</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:11:54 +0000</pubDate>
      <link>https://dev.to/sy_int999_fb11fa1a4f69806/add-an-faq-bot-and-appointment-booking-to-any-website-with-3-api-calls-6lk</link>
      <guid>https://dev.to/sy_int999_fb11fa1a4f69806/add-an-faq-bot-and-appointment-booking-to-any-website-with-3-api-calls-6lk</guid>
      <description>&lt;p&gt;Most small-business sites need the same two things: answer the five questions everyone asks ("what are your hours?", "how much is…?") and let people book. Building a booking backend for every client is a waste of a week. Here's how to do both with three HTTP calls, using CallChatSyn's API (free for the first 1,000 businesses).&lt;/p&gt;

&lt;h2&gt;
  
  
  0. Try it right now (no account)
&lt;/h2&gt;

&lt;p&gt;There's a public demo key, &lt;code&gt;ccs_demo_public&lt;/code&gt;, that answers as a demo business. Bookings made with it are dry runs, so nothing is saved:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://callchatsyn.com/api/v1/answer &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer ccs_demo_public"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"message": "Where are you located?"}'&lt;/span&gt;
&lt;span class="c"&gt;# {"intent":"faq","lang":"en","matched":true,"reply":"We're at our main location, details on our website."}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When you want your own FAQs and real bookings, get your own key:&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Set up the business (2 minutes)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Create an account at &lt;a href="https://callchatsyn.com/developers?utm_source=devto" rel="noopener noreferrer"&gt;callchatsyn.com&lt;/a&gt; and confirm your email.&lt;/li&gt;
&lt;li&gt;Add a few FAQs (Dashboard → FAQs) and your opening hours (Dashboard → Appointments). Set your time zone in Dashboard → Booking Page.&lt;/li&gt;
&lt;li&gt;Dashboard → Developers → &lt;strong&gt;Claim founder API access&lt;/strong&gt;, then &lt;strong&gt;Create API key&lt;/strong&gt;. Copy it once; it's not shown again.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Keep the key on your server. Never put it in browser code.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Answer a question
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://callchatsyn.com/api/v1/answer &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$CCS_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"message": "What are your opening hours?"}'&lt;/span&gt;
&lt;span class="c"&gt;# {"intent":"faq","lang":"en","matched":true,"reply":"We're open every day, 9am to 9pm."}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;matched: false&lt;/code&gt; means no FAQ fit and the reply is a polite fallback, so you can hand off to a human. It also recognises booking requests (&lt;code&gt;intent: "appointment"&lt;/code&gt;), order-status questions and "talk to a person", in English and Chinese.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. List open times and book one
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://callchatsyn.com/api/v1/slots &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$CCS_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="c"&gt;# {"slots":[{"start":"2026-10-05T01:00:00.000Z","label":"Mon, Oct 5, 9:00 AM"}],"services":["Haircut"],"locations":["Main St"]}&lt;/span&gt;

curl https://callchatsyn.com/api/v1/bookings &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$CCS_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"start":"2026-10-05T01:00:00.000Z","service":"Haircut","name":"Ana","email":"ana@example.com"}'&lt;/span&gt;
&lt;span class="c"&gt;# 201 {"booked":true,"start":"2026-10-05T01:00:00.000Z","label":"Mon, Oct 5, 9:00 AM"}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only times the business actually offers can be booked; a time taken a second earlier returns 409. Labels are already in the business's own time zone.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. A tiny website widget (server route + 30 lines of JS)
&lt;/h2&gt;

&lt;p&gt;Your server holds the key and forwards the question (Express example):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// server.js&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/chat&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;express&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://callchatsyn.com/api/v1/answer&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;Authorization&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CCS_KEY&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The page talks only to your server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;div&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"faq-bot"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;div&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"log"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&amp;lt;/div&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;form&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"ask"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&amp;lt;input&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"q"&lt;/span&gt; &lt;span class="na"&gt;placeholder=&lt;/span&gt;&lt;span class="s"&gt;"Ask us anything"&lt;/span&gt; &lt;span class="na"&gt;required&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&amp;lt;button&amp;gt;&lt;/span&gt;Send&lt;span class="nt"&gt;&amp;lt;/button&amp;gt;&amp;lt;/form&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/div&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;script&amp;gt;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;log&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;say&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;who&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createElement&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;p&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;who&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;log&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;appendChild&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ask&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;addEventListener&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;submit&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;preventDefault&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;q&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;q&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reset&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="nf"&gt;say&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;You&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;q&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/chat&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;q&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="nf"&gt;say&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Bot&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;reply&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Sorry, something went wrong.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;intent&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;appointment&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nf"&gt;say&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Bot&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Pick a time on our booking page, or ask me for the next available slot.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/script&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Limits, honestly
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Answers are rule-based FAQ matching, not a free-form LLM: instant, predictable, and close to free to run (which is why there's a free tier at all).&lt;/li&gt;
&lt;li&gt;Founder accounts get 100 calls a day and 2,000 a month per business.&lt;/li&gt;
&lt;li&gt;WhatsApp and phone/voice answering are the paid part of CallChatSyn.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Full docs and the OpenAPI spec: &lt;a href="https://callchatsyn.com/developers?utm_source=devto" rel="noopener noreferrer"&gt;callchatsyn.com/developers&lt;/a&gt;. Feedback welcome in the comments, especially from people who build sites for local businesses.&lt;/p&gt;

</description>
      <category>api</category>
      <category>webdev</category>
      <category>javascript</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Webwright ของ Microsoft: ให้ AI เขียนโค้ดท่องเว็บ แทนที่จะนั่งคลิกทีละครั้ง</title>
      <dc:creator>Nokka</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:10:31 +0000</pubDate>
      <link>https://dev.to/sarantoon/webwright-khng-microsoft-aih-ai-ekhiiynokhdthngewb-aethnthiicchanangkhlikthiilakhrang-4kkc</link>
      <guid>https://dev.to/sarantoon/webwright-khng-microsoft-aih-ai-ekhiiynokhdthngewb-aethnthiicchanangkhlikthiilakhrang-4kkc</guid>
      <description>&lt;h1&gt;
  
  
  Webwright ของ Microsoft: ให้ AI เขียนโค้ดท่องเว็บ แทนที่จะนั่งคลิกทีละครั้ง
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;โดย Nokka (นก-กา) | 3 ตุลาคม 2569&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;ถ้าคุณเคยปล่อยให้ AI agent ทำงานบนเว็บยาว ๆ แล้วเห็นมันเริ่มสับสนตอนคลิกที่สี่สิบ คุณไม่ได้คิดไปเอง และงานวิจัยชิ้นหนึ่งจาก Microsoft Research อธิบายว่าปัญหาไม่ได้อยู่ที่การคลิกผิดครั้งใดครั้งหนึ่ง&lt;/p&gt;

&lt;p&gt;ปัญหาอยู่ที่ &lt;strong&gt;วิธีที่ agent ทำงานทั้งกระบวนการ&lt;/strong&gt; คือดูหน้าเว็บ ตัดสินใจหนึ่งการกระทำ รอผล แล้ววนใหม่ ไปเรื่อย ๆ โดยไม่มีแผนที่ทนทานพอจะพาไปถึงปลายทาง&lt;/p&gt;

&lt;p&gt;Webwright แก้ที่รากของวิธีคิดนั้น ด้วยข้อเสนอที่ฟังดูเรียบง่ายจนน่าตกใจ คือ &lt;strong&gt;ให้โมเดลมี terminal แล้วให้มันเขียนโปรแกรมท่องเว็บเอง&lt;/strong&gt; [1]&lt;/p&gt;

&lt;p&gt;ผลลัพธ์ที่ได้คือบนงานยาว ๆ โมเดล GPT-5.4 ตัวเดียวกัน กระโดดจาก 33.5% ไปเป็น 60.1% และสิ่งที่มันทิ้งไว้ให้คุณไม่ใช่รอยคลิก แต่เป็นเครื่องมือที่เอาไปรันซ้ำได้ [2]&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1mm5qbj0q5pjnxzn38zk.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1mm5qbj0q5pjnxzn38zk.jpg" alt="ภาพแนวคิด: แขนกลในเวิร์กช็อปไม้กำลังวางประแจที่เพิ่งประกอบเสร็จลงบนแท่นกลม โดยมีเครื่องมือเรียงเป็นระเบียบอยู่บนชั้นด้านหลัง" width="799" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;ภาพแนวคิด คือเครื่องมือที่ทำครั้งเดียวแล้วหยิบมาใช้ซ้ำได้ เทียบกับรอยคลิกที่จางหายไปพร้อมกับ session&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  ปัญหาที่ทุกคนที่สร้าง agent เจอ แต่ไม่ค่อยมีใครตั้งชื่อให้มัน
&lt;/h2&gt;

&lt;p&gt;Web agent ทุกวันนี้ใช้รูปแบบเดียวกัน คือให้ browser session เป็นพื้นที่ทำงานของ agent ตัวมันเอง ในแต่ละก้าว โมเดลรับสภาพหน้าปัจจุบัน แล้วทำนายการกระทำถัดไปหนึ่งอย่าง&lt;/p&gt;

&lt;p&gt;การกระทำนั้นอาจเป็นคลิก พิมพ์ เลื่อนหน้า เลือก element ผ่าน DOM หรือเรียกเครื่องมือสั้น ๆ แต่ทั้งหมดมีข้อจำกัดร่วมกันคือ &lt;strong&gt;agent ถูกบังคับให้ทำนายการกระทำทีละก้าว ภายในลูปที่กำหนดไว้ล่วงหน้า&lt;/strong&gt; [1]&lt;/p&gt;

&lt;p&gt;นักวิจัยของ Webwright เขียนไว้ตรง ๆ ว่าการออกแบบนี้เคยมีประโยชน์ตอนที่ LLM ยังอ่อน การมี harness ที่จัดมาให้อย่างดีช่วยเชื่อมช่องว่างระหว่างสิ่งที่โมเดลทำได้กับสิ่งที่งานเว็บจริงต้องใช้&lt;/p&gt;

&lt;p&gt;แต่พอโมเดลเก่งขึ้น โดยเฉพาะเรื่องเขียนและแก้โค้ด &lt;strong&gt;harness แบบเดิมก็กลายเป็นคอขวดเสียเอง&lt;/strong&gt; เพราะมันล็อก agent ไว้ในลูปแคบ ๆ ที่โมเดลไม่จำเป็นต้องอยู่ในนั้นอีกแล้ว [1]&lt;/p&gt;

&lt;p&gt;และเมื่อทำงานเสร็จ สิ่งที่ agent ทิ้งไว้คือลำดับการคลิกที่ใช้ครั้งเดียวจบ ไม่มีอะไรที่หยิบมารันซ้ำได้&lt;/p&gt;

&lt;h2&gt;
  
  
  แนวคิด: แยก agent ออกจาก browser
&lt;/h2&gt;

&lt;p&gt;จุดที่ผมคิดว่าเฉียบที่สุดของงานนี้คือการย้าย "สถานะ" ของงานไปไว้ที่อื่น&lt;/p&gt;

&lt;p&gt;Webwright เสนอว่า &lt;strong&gt;ให้แยก agent ออกจาก browser&lt;/strong&gt; แล้วมอง browser เป็นเครื่องมือที่ agent สั่งเปิด ตรวจสอบ และทิ้งได้ตามใจ ระหว่างที่มันกำลังพัฒนาโปรแกรมชิ้นหนึ่ง [1]&lt;/p&gt;

&lt;p&gt;สิ่งที่คงอยู่ไม่ใช่ session ของ browser แต่คือ &lt;strong&gt;โค้ดกับบันทึกในพื้นที่ทำงานบนเครื่อง&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;พูดเป็นภาษาคนก็คือ คลิกทำให้งานเสร็จหนึ่งครั้ง แต่โค้ดทำให้งานเสร็จและเก็บวิธีแก้ไว้ด้วย [3]&lt;/p&gt;

&lt;h2&gt;
  
  
  ข้างในมีแค่สามชิ้น
&lt;/h2&gt;

&lt;p&gt;ความเรียบง่ายของสถาปัตยกรรมเป็นจุดที่ผมประทับใจรองลงมา ไม่มีระบบ multi-agent ไม่มี graph engine ไม่มีชั้นปลั๊กอิน ไม่มี orchestration ที่ซ่อนอยู่&lt;/p&gt;

&lt;p&gt;มีแค่ terminal, browser และโมเดล [3]&lt;/p&gt;

&lt;p&gt;ระบบแกนกลางมีสามส่วน [2]&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Runner&lt;/strong&gt; เก็บว่างานคืออะไร ตอนนี้ทำถึงไหน และผลของคำสั่งก่อน ๆ เป็นอย่างไร&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model Endpoint&lt;/strong&gt; เชื่อม Webwright เข้ากับโมเดล รองรับ OpenAI, Anthropic และ OpenRouter&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Environment&lt;/strong&gt; ให้ terminal ที่ต่อกับ Playwright บน Chromium เป็นที่ที่คำสั่งรันจริง ไฟล์ถูกเก็บจริง และ screenshot ถูกบันทึกจริง&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ลูปการทำงานสั้นมาก คือเข้าใจสถานะปัจจุบัน เลือกคำสั่ง รันมัน ดูว่าเกิดอะไร จากนั้นวนใหม่ จนโมเดลคิดว่าเสร็จ และมีขั้นตรวจตัวเองอีกชั้นคอยยืนยัน [3]&lt;/p&gt;

&lt;h2&gt;
  
  
  ตัวเลขที่ทำให้ต้องหยุดอ่าน
&lt;/h2&gt;

&lt;p&gt;บนชุดทดสอบ Online-Mind2Web ซึ่งมี 300 งานจริงจากเว็บจริง GPT-5.4 ร่วมกับ Webwright ทำได้ &lt;strong&gt;86.7%&lt;/strong&gt; สูงที่สุดในกลุ่ม harness โอเพนซอร์สที่วัดด้วย AutoEval ส่วน Claude Opus 4.7 ทำได้ 84.7% แต่แข็งกว่าบนชุดงานยากที่ 80.5% เทียบกับ 76.6% ของ GPT-5.4 [3]&lt;/p&gt;

&lt;p&gt;แต่ตัวเลขที่ผมว่าสำคัญกว่าคือชุด Odysseys ซึ่งเป็นงานยาว 200 งาน&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;โมเดล GPT-5.4 ที่ควบคุม browser ด้วยพิกัดบนหน้าจอ ทำได้ 33.5% [3]&lt;/li&gt;
&lt;li&gt;โมเดลตัวเดียวกันเมื่อเขียนโค้ดผ่าน Webwright ทำได้ 60.1% โดยใช้เวลาเฉลี่ย 76.1 ก้าว [3]&lt;/li&gt;
&lt;li&gt;คิดเป็นการเพิ่มขึ้น 26.6 จุด &lt;strong&gt;จากการเปลี่ยน harness ไม่ใช่จากการเปลี่ยนโมเดล&lt;/strong&gt; [1][3]&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;และมีผลพลอยได้ที่ผมว่าเป็นสัญญาณเชิงกลยุทธ์ คือเมื่อ Webwright สร้างเครื่องมือที่ใช้ซ้ำได้ไว้แล้ว โมเดลเล็กลงก็ทำงานได้ นั่นหมายความว่าเครื่องมือที่ดีลดความต้องการโมเดลใหญ่ในครั้งต่อไป [3]&lt;/p&gt;

&lt;h2&gt;
  
  
  อะไรที่ยังต้องระวัง
&lt;/h2&gt;

&lt;p&gt;งานนี้ไม่ได้ฟรี และตัวบทความต้นทางเองก็เขียนข้อจำกัดไว้ตรงไปตรงมา ผมว่าคนอ่านควรรู้ก่อนเอาไปใช้&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;หนึ่ง ตัวเลขทั้งหมดเป็นการตัดสินโดย LLM&lt;/strong&gt; คือ AutoEval ไม่ใช่การวัดด้วย unit test ทุกงาน และผลหัวข่าวของ Mind2Web ใช้เพียง 100 จาก 300 งาน [2]&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;สอง ต้นทุนต่องานยังสูง&lt;/strong&gt; ประมาณ 2.37 ดอลลาร์ต่องานเมื่อใช้ GPT-5.4 และ 6.09 ดอลลาร์เมื่อใช้ Claude Opus 4.7 เพราะ Webwright ลงทุนคำนวณช่วงต้นเพื่อสร้างเครื่องมือที่ทนกว่า แล้วค่อยใช้ซ้ำ [2]&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;สาม ต้นทุนไม่ได้หายไป มันย้ายที่&lt;/strong&gt; ในตัวอย่างของ Microsoft การใช้เป็น skill บน Codex ใช้โทเคนราว 3.3 ล้าน เทียบกับ 424,000 เมื่อรันเป็น harness เดี่ยว ซึ่งมากกว่าประมาณแปดเท่า เพราะบริบทถูกแคชไว้ใน session ของโฮสต์ [2]&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;สี่ จุดที่ต้นทางเองรายงานไม่ตรงกัน&lt;/strong&gt; ผมเจอสองจุดที่ต้องบอกตามจริง จุดแรกคือขนาดของระบบ บล็อกของทีมผู้พัฒนาเองบอกว่าแกนกลางประมาณ 1,000 บรรทัด แบ่งเป็น Runner ราว 150 บรรทัด Model Endpoint ราว 550 บรรทัด และ Environment ราว 300 บรรทัด [1] ขณะที่ README ของ repo นับเป็นระดับไฟล์ คือลูปหลักของ agent ราว 450 บรรทัด สภาพแวดล้อม Playwright ราว 570 บรรทัด และ CLI อีก 150 บรรทัด [3] สองชุดนี้รวมแล้วใกล้กัน แต่เป็นการนับคนละหน่วย คือชุดแรกนับตามโมดูล ชุดหลังนับตามไฟล์&lt;/p&gt;

&lt;p&gt;จุดที่สองคือคะแนน Odysseys หน้าโครงการเขียนว่า 60.8% ขณะที่การเทียบใน repo และบทวิเคราะห์ใช้ 60.1% [3] ซึ่งบทวิเคราะห์นั้นระบุเหตุผลไว้เองว่าเลือกใช้ตัวเลขจาก repo เพื่อความสม่ำเสมอ&lt;/p&gt;

&lt;p&gt;ผมเห็นว่าเป็นความต่างระดับการปัดเศษมากกว่าจะเป็นข้อขัดแย้ง และการรายงานทั้งสองค่าไว้จะช่วยให้คุณตรวจเองได้ ดีกว่าให้ผมเลือกข้างแทนคุณ&lt;/p&gt;

&lt;h2&gt;
  
  
  จุดที่งานนี้ไปไกลกว่าที่คิด
&lt;/h2&gt;

&lt;p&gt;ส่วนที่ผมไม่ได้คาดไว้ตอนเริ่มอ่าน คือ Webwright ไม่ได้หยุดแค่ทำภารกิจให้เสร็จ แต่มันต่อยอดไปถึงสิ่งที่ผมเรียกว่าโรงงานเครื่องมือ&lt;/p&gt;

&lt;p&gt;ในเวอร์ชัน 21 กรกฎาคม 2026 มีฟีเจอร์ชื่อ Skill Factory ซึ่งทุกครั้งที่แก้ปัญหาสำเร็จ มันจะกลั่นสคริปต์ที่ได้ออกมาเป็นเครื่องมือที่ผ่านการตรวจสอบ มีพารามิเตอร์ และรันได้เองโดยไม่ต้องมีโมเดล ใช้เวลาราว 40 วินาที และไม่ใช้โทเคนเลย บนชุด WebArena การเอาเครื่องมือเก่ากลับมาใช้ยกความแม่นยำบนชุดที่กันไว้จาก 55% เป็น 70% หรือเพิ่ม 15 จุด [3]&lt;/p&gt;

&lt;p&gt;และตั้งแต่ 6 พฤษภาคม 2026 มีปลั๊กอินสำหรับ Claude Code, Codex, OpenClaw รวมทั้ง &lt;strong&gt;Hermes Agent&lt;/strong&gt; โดยใช้โฟลเดอร์ skill ชุดเดียวกันโหลดข้ามเครื่องมือได้ [3] ซึ่งสำหรับคนที่ทำงานกับ agent อยู่แล้ว นี่คือจุดที่ทำให้ลองได้ทันทีโดยไม่ต้องรื้อระบบเดิม&lt;/p&gt;

&lt;h2&gt;
  
  
  อะไรที่คนไทยเอาไปใช้ได้จริง
&lt;/h2&gt;

&lt;h3&gt;
  
  
  สี่ข้อที่ผมว่าคุ้มที่สุดสำหรับทีมเล็ก
&lt;/h3&gt;

&lt;p&gt;สำหรับคนที่ &lt;strong&gt;ทำงานกับเว็บซ้ำ ๆ ทุกสัปดาห์&lt;/strong&gt; ไม่ว่าจะดึงราคาคู่แข่ง เก็บรายชื่อจากไดเรกทอรี หรือกรอกฟอร์มเป็นชุด วิธีคิดของ Webwright ให้คำตอบที่ใช้ได้เลย คือให้ agent สร้างเครื่องมือที่รันซ้ำได้ ไม่ใช่สร้างรอยคลิกที่ต้องทำใหม่ทุกครั้ง [3]&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;หนึ่ง ถ้างานของคุณทำซ้ำ ให้ลงทุนสร้างสคริปต์ครั้งเดียว&lt;/strong&gt; งานซ้ำคือที่ที่การลงทุนคำนวณช่วงต้นคุ้มที่สุด เพราะเครื่องมือที่ได้จะถูกรันซ้ำได้โดยไม่มีค่าโมเดล&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;สอง ดูว่างานไหนที่ agent มักหลุดกลางทาง&lt;/strong&gt; ถ้าอาการคือสับสนตอนก้าวที่สามสิบหรือสี่สิบ ปัญหาอาจไม่ใช่ความสามารถของโมเดล แต่วิธีที่เราบังคับให้มันทำทีละก้าว&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;สาม ถ้าคุณใช้ agent อยู่แล้ว ลองเพิ่ม skill ก่อนเปลี่ยนโมเดล&lt;/strong&gt; เพราะการทดลองของงานนี้ชี้ว่าการเปลี่ยน harness ให้ผลมากกว่าการเปลี่ยนโมเดลในงานยาว และค่าใช้จ่ายในการลองถูกกว่าการอัปเกรดโมเดลมาก&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;สี่ ระวังกับดักของเครื่องมือที่พอกพูน&lt;/strong&gt; เพราะการสร้างเครื่องมืออัตโนมัติย่อมทำให้มีสคริปต์เพิ่มขึ้นเรื่อย ๆ ถ้าไม่มีวินัยในการรื้อของที่หยุดใช้ คุณจะได้หนี้ทางเทคนิคกลับมาแทน&lt;/p&gt;

&lt;h2&gt;
  
  
  คำถามที่ต้องถามก่อนเชื่อ
&lt;/h2&gt;

&lt;p&gt;ผมคิดว่าคนอ่านควรถามสามข้อนี้ก่อนตัดสินใจลงมือ&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;หนึ่ง ตัวเลข 60.1% วัดจากการตัดสินของ LLM หรือวัดจากความสำเร็จที่เป็นรูปธรรม&lt;/strong&gt; ถ้าเป็นอย่างแรก คุณควรเผื่อความคลาดเคลื่อนไว้ และอ่านผลแบบเทียบสัดส่วนมากกว่าเทียบตัวเลขเดี่ยว [2]&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;สอง งานของคุณยาวพอที่การเปลี่ยน harness จะคุ้มไหม&lt;/strong&gt; ถ้างานสั้นเพียงไม่กี่ก้าว วิธีเดิมที่ทำทีละคลิกอาจเพียงพอ และการสร้างสคริปต์อาจเป็นงานเกินจำเป็น&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;สาม ต้นทุนแปดเท่าที่โผล่มาในโหมด skill เกิดจากบริบทของโฮสต์&lt;/strong&gt; ไม่ได้แปลว่า Webwright แพง แต่หมายความว่าคุณต้องวัดค่าใช้จ่ายที่ระดับ session ไม่ใช่ที่ระดับงาน&lt;/p&gt;

&lt;p&gt;ผมเล่าเรื่องนี้ด้วยความระวัง เพราะตัวเลขทุกตัวในบทความนี้เป็นตัวเลขที่ทีมผู้พัฒนารายงานเอง ผมจึงเลือกเล่าทั้งข้อดีและข้อที่ยังต้องระวังไว้คู่กัน เพื่อให้คุณตัดสินได้เอง ไม่ใช่ให้ผมตัดสินแทน&lt;/p&gt;

&lt;p&gt;มีอีกมุมที่ผมว่าสำคัญสำหรับคนทำงานจริง คือคำถามว่า &lt;strong&gt;สคริปต์ที่ agent เขียนให้เราจะพังเมื่อเว็บเปลี่ยนเมื่อไหร่&lt;/strong&gt; เพราะเว็บไม่เคยหยุดนิ่ง ปุ่มย้ายที่ id เปลี่ยน ฟอร์มเพิ่มช่อง ข้อดีของวิธีนี้คือเมื่อสคริปต์พัง เราจะเห็นทันทีว่าพังตรงบรรทัดไหน และแก้ที่โค้ดได้ตรงจุด ซึ่งต่างจากการดีบักรอยคลิกที่ไม่มีอะไรให้อ่านย้อนหลังเลย&lt;/p&gt;

&lt;p&gt;ถ้าคุณอยากเริ่มจากตรงที่ง่ายที่สุด ลองเลือกงานซ้ำหนึ่งงานที่ทีมทำทุกสัปดาห์ แล้วลองให้ agent สร้างสคริปต์ที่รันซ้ำได้ แล้ววัดว่าครั้งที่สองเร็วกว่าครั้งแรกแค่ไหน คุณจะเห็นคุณค่าของแนวคิดนี้โดยไม่ต้องเชื่อตัวเลขของใคร&lt;/p&gt;

&lt;h2&gt;
  
  
  ที่ควรไปอ่านต่อ
&lt;/h2&gt;

&lt;p&gt;โค้ดทั้งหมดเปิดที่ GitHub ของ Microsoft [4] และหน้าโครงการมีภาพรวมของงานพร้อมผลทดสอบแยกตามชุด [3] ส่วนบทวิเคราะห์ภาษาเข้าใจง่ายที่ผมใช้เป็นแหล่งหลักของบทความนี้อยู่ที่ Towards Data Science ซึ่งลงลึกเรื่องจุดที่วิธีนี้ยังแก้ไม่ครบบางกรณี [2]&lt;/p&gt;

&lt;p&gt;ผมคิดว่างานนี้จะเป็นประโยชน์ที่สุดกับคนที่สร้าง agent แล้วรู้สึกว่ามันเก่งบนงานสั้นแต่พังบนงานยาว เพราะคำตอบของ Webwright คือคุณอาจไม่ได้ต้องการโมเดลที่ฉลาดขึ้น แต่ต้องการให้ agent ทิ้งสิ่งที่ใช้ซ้ำได้ไว้ข้างหลัง แทนที่จะทิ้งแค่ประวัติการคลิกที่จางหายไปพร้อมกับ session&lt;/p&gt;

&lt;p&gt;คุณเคยเจอเคสที่ agent ทำอะไรซ้ำ ๆ ทุกสัปดาห์ไหม? และถ้ามี งานแบบนั้นคือที่ที่แนวคิดนี้ให้ผลชัดที่สุด คุณไม่ต้องเปลี่ยนโมเดล และไม่ต้องรื้อระบบเดิม แต่เปลี่ยนสิ่งที่ agent ทิ้งไว้ข้างหลัง จากรอยคลิกเป็นโค้ดที่รันซ้ำได้&lt;/p&gt;

&lt;p&gt;&lt;em&gt;บทความนี้เขียนโดย AI (deepseek-v4.1-flash) ผ่าน Hermes Agent ภายใต้การควบคุมและตรวจสอบคุณภาพโดยมนุษย์ - Nokka (นก-กา)&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  เอกสารอ้างอิง
&lt;/h2&gt;

&lt;p&gt;[1] Webwright: A Terminal Is All You Need For Web Agents โดย Microsoft Research (4 พฤษภาคม 2026 · เข้าถึง 3 ตุลาคม 2026) · &lt;a href="https://www.microsoft.com/en-us/research/articles/webwright-a-terminal-is-all-you-need-for-web-agents/" rel="noopener noreferrer"&gt;https://www.microsoft.com/en-us/research/articles/webwright-a-terminal-is-all-you-need-for-web-agents/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;[2] Webwright: Why AI Web Agents Should Write Code, Not Click โดย Chien Vu Minh, Towards Data Science (17 สิงหาคม 2026) · &lt;a href="https://towardsdatascience.com/webwright-why-ai-web-agents-should-write-code-not-click/" rel="noopener noreferrer"&gt;https://towardsdatascience.com/webwright-why-ai-web-agents-should-write-code-not-click/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;[3] Webwright ซึ่งเป็นหน้าโครงการ เอกสารและ README ของ repo (2026 · เข้าถึง 3 ตุลาคม 2026) · &lt;a href="https://microsoft.github.io/Webwright/" rel="noopener noreferrer"&gt;https://microsoft.github.io/Webwright/&lt;/a&gt; · &lt;a href="https://github.com/microsoft/Webwright/blob/main/README.md" rel="noopener noreferrer"&gt;https://github.com/microsoft/Webwright/blob/main/README.md&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;[4] โค้ด Webwright บน GitHub ของ Microsoft (2026 · เข้าถึง 3 ตุลาคม 2026) · &lt;a href="https://github.com/microsoft/Webwright" rel="noopener noreferrer"&gt;https://github.com/microsoft/Webwright&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>A threshold is a policy, not a number</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:10:15 +0000</pubDate>
      <link>https://dev.to/turingcorp/a-threshold-is-a-policy-not-a-number-da4</link>
      <guid>https://dev.to/turingcorp/a-threshold-is-a-policy-not-a-number-da4</guid>
      <description>&lt;h1&gt;
  
  
  A threshold is a policy, not a number
&lt;/h1&gt;

&lt;p&gt;Somewhere in a payments codebase there is a line that says: approve automatically when confidence is above 0.8. Nobody remembers the afternoon it was written. The number has three likely origins and all of them are bad — it was the first value that made the demo behave, it was copied from a vendor example, or it was chosen because 0.8 sounds strict without sounding paranoid. What it was not read off is anything that describes how this particular system behaves when it is unsure.&lt;/p&gt;

&lt;p&gt;That one value decides which refunds are executed by a machine and which ones wait for a person. Set it in the wrong place and the consequences split in two directions, easy to confuse and hard to price: some share of the volume that deserved a look gets executed at machine speed, and the rest — usually a much larger share of it — is pushed onto a human queue that was never staffed for what the gate sends it.&lt;/p&gt;

&lt;p&gt;The stack underneath is ordinary. A refund request arrives; arithmetic settles the trivial cases; a decision model answers what a predicate cannot express; a person owns the ones where both answers survive scrutiny. That shape is background. The subject is the joint: the single number that decides which of those three handles a given case, and what happens when it is set by taste instead of by measurement.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the joint is switching between
&lt;/h2&gt;

&lt;p&gt;The rule layer is the one everybody trusts, because it is the one everybody can read. It fails silently: the business moves, the boundary the rule was drawn around does not, and a wrong rule decision looks exactly like a correct one.&lt;/p&gt;

&lt;p&gt;The decision model is a different animal — trained rather than written, returning a distribution rather than a branch. It fails out of distribution: handed a case its training data never contained, it does not raise an error, it returns a number wearing the same face as every other number it has emitted.&lt;/p&gt;

&lt;p&gt;The third destination is not a bigger model, it is a different question — which of two defensible options should we live with. It fails by not happening. The ticket sits, the case ages, and a decision nobody made does not look like a wrong decision. It looks like a backlog.&lt;/p&gt;

&lt;p&gt;Those three paragraphs are the whole of the background, because the failure modes are not what this article is about. Each one has a price, and the latch is where you decide who pays it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the 0.8 came from
&lt;/h2&gt;

&lt;p&gt;A threshold compares two numbers: one the system emits, one you choose. The first is a claim about frequency. When a model reports 0.87, it is claiming that among cases that look like this one, roughly 87 in 100 come out the way it says. Confidence is not a feeling the model has about itself; it is a prediction about its own hit rate, and like any prediction it can be checked against outcomes.&lt;/p&gt;

&lt;p&gt;The chosen number is rarely checked against anything. It arrives from a headline accuracy figure for the whole system, from a default in a library, or from the first value that stopped the demo from embarrassing anyone. None of those is a statement about what 0.87 means.&lt;/p&gt;

&lt;p&gt;An aggregate score and a confidence value answer different questions. Accuracy tells you how often the judge is right across everything you tested. Calibration tells you what a reported 0.87 actually buys in error rate. The gate runs entirely on the second question and is almost always set with the first.&lt;/p&gt;

&lt;p&gt;This is also why calibrating a model and choosing a threshold are two separate jobs, done by two different kinds of judgment. A perfectly calibrated model still does not tell you where to cut. It tells you what each possible cut costs. Where to cut is a question about consequences, and the model has never seen your consequences.&lt;/p&gt;

&lt;h2&gt;
  
  
  Both directions are expensive; only one is invisible
&lt;/h2&gt;

&lt;p&gt;Set the cut below what the confidence numbers actually mean and cases that should have been looked at are executed instead. Those failures are irreversible: money leaves, a commitment is honored at machine speed on a judgment that never earned machine speed. They do not arrive as errors. They arrive as outcomes — from a customer, a chargeback report, or an auditor two quarters later.&lt;/p&gt;

&lt;p&gt;Set the cut above and you get the failure everyone underestimates because it looks like diligence. Everything interesting escalates. The automation runs, it costs what it costs, and the queue on the other side fills with items that did not need a person. Do the arithmetic once: at 100,000 requests a day, escalating 40% means 40,000 human reviews a day. You have not automated the work. You have moved it and paid for the tool and the worker. The failure surfaces as headcount or as aging, which is why it survives for quarters.&lt;/p&gt;

&lt;p&gt;The two mistakes are not symmetric and they are not equally visible. The first is loud but rare, and in the happy path it resembles throughput. The second is quiet, permanent, and looks like caution — and it is the one that quietly doubles the cost of the automation you just bought.&lt;/p&gt;

&lt;p&gt;Neither direction is a math error. Choosing which one you can survive is the actual decision, and it does not have a universal answer. If a wrong automated approval costs a refund, cut high. If a delayed decision costs a contract, cut low. There is no single threshold that is correct for both, which is the first sign that you are choosing a policy rather than tuning a constant.&lt;/p&gt;

&lt;h2&gt;
  
  
  The curve has to exist before the latch means anything
&lt;/h2&gt;

&lt;p&gt;The only artifact that answers the question is a reliability report: for each band of reported confidence, how often the call was correct, with the number of cases in the band.&lt;/p&gt;

&lt;p&gt;We publish ours. Self-run on JudgeBench, 620 judgments, with the 6 that failed on a first verdict disclosed rather than quietly retried: calls the system reported at 90% confidence or above were right 99.6% of the time, and calls in the 80–90% band were right 94.0%. On that same self-run, raw accuracy came out at 92.5% against 92.2% for a plain direct baseline — a tie, reported as a tie, with nothing claimed over it. The value of the bands is not the headline. It is that a threshold can now be stated as a price: cut at 90 and you are accepting a 0.4% error rate on what you automate; cut at 80 and you are accepting 6%.&lt;/p&gt;

&lt;p&gt;Read the shape of the curve, not just the top of it. A curve that flattens out in the middle, and says so honestly, is telling the stack where the third layer begins — the region where the judge's own answer is that it does not know. A curve that is high everywhere tells you nothing, and a gate built on it will either automate everything or mean nothing.&lt;/p&gt;

&lt;p&gt;One caution before you print the table into a design document: the curve is measured on a benchmark distribution, not on your traffic. It describes the judge. It says nothing about the population you are feeding it, which is why the band table is necessary and not sufficient.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three questions before you trust the latch
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Where does the error land?&lt;/strong&gt; Answer it per boundary, not for the system as a whole. A re-run and a log line mean the machine can carry the mistake; a customer, a balance sheet, or a relationship means it cannot. The boundary of what you may automate is drawn by that question and nothing else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the base rate of what you are screening for?&lt;/strong&gt; A gate that passes 97% of a rare bad case is a different object from one that passes 97% of a common one. The band table gives you a rate conditional on confidence; your base rate converts that into the number of bad outcomes per day. Two teams can read the same table and owe different amounts of money.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who signs, and have you asked them?&lt;/strong&gt; The escalation layer fails when nobody will own the call, and the cheapest way to find that out is to ask before you build the escalation. If the answer is "whoever is on call," the policy is not a policy — it is an absence of one with a queue attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  A parameter gets tuned; a policy gets defended
&lt;/h2&gt;

&lt;p&gt;Put two companies in front of the same band table. A lender reading a borderline credit file and a hospital reading a borderline discharge will compute the same numbers and should still land on different cuts, because the errors do not land on the same party. The threshold is the only place in the architecture where that asymmetry becomes executable code. Everything above it is engineering; this one line is management.&lt;/p&gt;

&lt;p&gt;That is why the number should not live only in a config file. Write it down with a date, an owner, and a trigger for revisiting it. The trigger is not a calendar reminder; it is a change in the base rate, a change in the label set, or a change in who absorbs the loss. A threshold with no owner drifts, and drift is invisible right up until someone audits the queue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Close
&lt;/h2&gt;

&lt;p&gt;The line in the config file is the shortest policy document most companies have ever written, and almost nobody treats it that way. If you cannot say who chose the number, when, and what they were trading away, then it was not a policy. It was a guess with production access.&lt;/p&gt;

&lt;p&gt;The fix is not a better number. It is a name next to the number, and a curve that says what the number buys.&lt;/p&gt;

&lt;p&gt;If you are working on that joint — the confidence, the threshold, the report that is supposed to justify it — &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-8" rel="noopener noreferrer"&gt;Decider&lt;/a&gt; answers one question with two candidate options, a pick, a calibrated confidence, and a written argument you can disagree with.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I Built My Own Database From Scratch — Meet CleaveDB</title>
      <dc:creator>John David L. Perez</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:10:09 +0000</pubDate>
      <link>https://dev.to/john_davidlperez_f528e/i-built-my-own-database-from-scratch-meet-cleavedb-2494</link>
      <guid>https://dev.to/john_davidlperez_f528e/i-built-my-own-database-from-scratch-meet-cleavedb-2494</guid>
      <description>&lt;p&gt;I Built My Own Database From Scratch — Meet CleaveDB&lt;/p&gt;

&lt;p&gt;I've been building something I've wanted to experiment with for a while: &lt;strong&gt;CleaveDB&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;CleaveDB is a hybrid relational-document graph database built with &lt;strong&gt;Python, Rust, and C++ SIMD&lt;/strong&gt;, with its own query language, graph relationships, and built-in semantic search.&lt;/p&gt;

&lt;p&gt;The idea started with a simple question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What if relational data, documents, graphs, and semantic search could live together in one database?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What makes CleaveDB different?
&lt;/h2&gt;

&lt;p&gt;CleaveDB has its own query language called &lt;strong&gt;CleaveQL&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;POUR&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt; &lt;span class="nv"&gt;"alice"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nv"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;"Alice"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nv"&gt;"age"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;28&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;You can then query the data:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;SCOOP&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;It also supports graph-style relationships called &lt;strong&gt;Bonds&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;LINK&lt;/span&gt; &lt;span class="nv"&gt;"users:alice"&lt;/span&gt;
&lt;span class="k"&gt;TO&lt;/span&gt; &lt;span class="nv"&gt;"users:bob"&lt;/span&gt;
&lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;MUTUAL&lt;/span&gt; &lt;span class="nv"&gt;"friend"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;And you can traverse those relationships:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;FIND&lt;/span&gt; &lt;span class="nv"&gt;"friend"&lt;/span&gt; &lt;span class="k"&gt;OF&lt;/span&gt; &lt;span class="nv"&gt;"users:alice"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;There's also semantic search:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;FIND&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt;
&lt;span class="n"&gt;MEANING&lt;/span&gt; &lt;span class="nv"&gt;"something warm for cold weather"&lt;/span&gt;
&lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;So instead of searching only for exact keywords, CleaveDB can search based on meaning.&lt;/p&gt;
&lt;h2&gt;
  
  
  Under the hood
&lt;/h2&gt;

&lt;p&gt;CleaveDB isn't built on top of an existing database engine. I'm building the storage layer myself using &lt;strong&gt;Rust&lt;/strong&gt;, with C++ SIMD optimisations for performance-critical operations.&lt;/p&gt;

&lt;p&gt;The project also includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;B+Tree storage&lt;/li&gt;
&lt;li&gt;WAL and transactions&lt;/li&gt;
&lt;li&gt;Graph relationships&lt;/li&gt;
&lt;li&gt;Semantic/vector search&lt;/li&gt;
&lt;li&gt;Authentication and multi-tenant isolation&lt;/li&gt;
&lt;li&gt;Python bindings&lt;/li&gt;
&lt;li&gt;HTTP, WebSocket and TCP interfaces&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's still very much a work in progress, but building the pieces from scratch has been a really interesting learning experience.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why build another database?
&lt;/h2&gt;

&lt;p&gt;There are already excellent databases like PostgreSQL, MongoDB, Redis, and Neo4j.&lt;/p&gt;

&lt;p&gt;I'm not trying to replace them.&lt;/p&gt;

&lt;p&gt;CleaveDB is an experiment to see what happens when you combine &lt;strong&gt;relational, document, graph, and semantic capabilities&lt;/strong&gt; into one system.&lt;/p&gt;

&lt;p&gt;If you're interested in databases, Rust, storage engines, graph databases, or just want to see how a database is built from scratch, I'd love to hear your thoughts.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/Tribrix23" rel="noopener noreferrer"&gt;
        Tribrix23
      &lt;/a&gt; / &lt;a href="https://github.com/Tribrix23/CleaveDB" rel="noopener noreferrer"&gt;
        CleaveDB
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Beyond NoSQL. A hyper-fast graph and vector database you query in plain English. Built on a unique Python, Rust, and C++ SIMD architecture with built-in AI semantic search and strict multi-tenant isolation.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div&gt;
  &lt;a rel="noopener noreferrer" href="https://github.com/Tribrix23/CleaveDB/assets/CleaveDB.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FTribrix23%2FCleaveDB%2FHEAD%2Fassets%2FCleaveDB.png" alt="CleaveDB 3.9.0" width="500"&gt;&lt;/a&gt;
  &lt;br&gt;&lt;br&gt;
  &lt;p&gt;&lt;strong&gt;The polyglot, AVX-512 ready , hybrid relational-document graph database with Transformer attention layers.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.python.org/downloads/?Python" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/262eeadeee3f7ea82351c9d0a3d601ec1e8024af17178cfebaa7b0553d62297f/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f507974686f6e2d332e31332b2d626c75652e737667" alt="Python"&gt;&lt;/a&gt;
&lt;a href="https://doc.rust-lang.org/edition-guide/rust-2021/index.html" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/8d2491af9319462dd27d3fddcaef0a20e44bbfd98c6f054dc714cc6912127cfc/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f527573742d323032312d6f72616e67652e737667" alt="Rust"&gt;&lt;/a&gt;
&lt;a href="https://github.com/Tribrix23/CleaveDB/blob/main/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/f057393ac885f515986871aeac9852ee5dedd1127efe65b0f02b6c7a1aae5da3/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d437573746f6d5f436f6e73656e742d626c75652e737667" alt="License"&gt;&lt;/a&gt;
&lt;a href="https://www.npmjs.com/package/cleavedb" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/25d5ccc31f2a56630360f697e845a73e62ce7860adcf42bc0940a7104b40d048/68747470733a2f2f696d672e736869656c64732e696f2f6e706d2f762f636c6561766564622e737667" alt="NPM Version"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;CleaveDB 3.9.0&lt;/strong&gt; is a ground-up, hybrid Relational Document &amp;amp; Graph database that eliminates the complexity of traditional SQL &lt;code&gt;JOIN&lt;/code&gt;s, external vector search services, and opaque graph databases. It ships with:&lt;/p&gt;


&lt;ul&gt;

&lt;li&gt;

&lt;strong&gt;Native Graph-Relational Links&lt;/strong&gt;: Documents are not isolated. They are deeply relational, linked natively through ~10 functional Graph "Bonds" that allow unlimited depth, multi-hop traversals with conversational English, completely eliminating the need for &lt;code&gt;JOIN&lt;/code&gt;s.&lt;/li&gt;

&lt;li&gt;A high-performance &lt;strong&gt;Rust storage engine&lt;/strong&gt; built entirely from scratch — no SQLite, no RocksDB, no external storage libraries.&lt;/li&gt;

&lt;li&gt;

&lt;strong&gt;AVX-512 / AVX2 C++ SIMD extensions&lt;/strong&gt; exist in the Rust engine (vector search runs on C++ AVX-512 extensions).&lt;/li&gt;

&lt;li&gt;A &lt;strong&gt;Python interpreter frontend&lt;/strong&gt; (via PyO3 bindings) that runs the &lt;strong&gt;CleaveQL&lt;/strong&gt; query language.&lt;/li&gt;

&lt;li&gt;

&lt;strong&gt;Real neural Transformer embeddings&lt;/strong&gt; for semantic search via a quantized ONNX model, using ~22MB of RAM.&lt;/li&gt;

&lt;li&gt;A &lt;strong&gt;Go-based distributed coordinator&lt;/strong&gt; for multi-shard…&lt;/li&gt;

&lt;/ul&gt;&lt;/div&gt;
&lt;br&gt;
  &lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/Tribrix23/CleaveDB" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;I'm especially interested in criticism and ideas for what I should improve next.&lt;/p&gt;

</description>
      <category>database</category>
      <category>rust</category>
      <category>python</category>
      <category>programming</category>
    </item>
    <item>
      <title>A single pfSense block never becomes a Wazuh alert. Here is why, measured.</title>
      <dc:creator>Nguyen Dong</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:08:44 +0000</pubDate>
      <link>https://dev.to/xuxu298/a-single-pfsense-block-never-becomes-a-wazuh-alert-here-is-why-measured-563m</link>
      <guid>https://dev.to/xuxu298/a-single-pfsense-block-never-becomes-a-wazuh-alert-here-is-why-measured-563m</guid>
      <description>&lt;p&gt;You point pfSense at Wazuh, the logs arrive, and the dashboard stays empty. Most people assume the integration is broken. On Wazuh 4.14.7 it usually is not: there are two separate reasons, and one of them is by design.&lt;/p&gt;

&lt;p&gt;We sent pfSense &lt;code&gt;filterlog&lt;/code&gt; lines over UDP syslog to a Wazuh 4.14.7 manager container and read &lt;code&gt;archives.log&lt;/code&gt; and &lt;code&gt;alerts.json&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reason 1: rule 87701 carries &lt;code&gt;no_log&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;The stock pfSense rules live in &lt;code&gt;ruleset/rules/0540-pfsense_rules.xml&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Rule&lt;/th&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Matches&lt;/th&gt;
&lt;th&gt;Becomes an alert?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;87700&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;any &lt;code&gt;filterlog&lt;/code&gt; event decoded as &lt;code&gt;pf&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;no (level 0)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;87701&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;action&lt;/code&gt; = block&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;no&lt;/strong&gt;: it has &lt;code&gt;&amp;lt;options&amp;gt;no_log&amp;lt;/options&amp;gt;&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;87702&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;18 × 87701 from one source IP within 45 s&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The comment above 87701 says: &lt;em&gt;"We don't log firewall events, because they go to their own log file."&lt;/em&gt; Pass events stop at 87700.&lt;/p&gt;

&lt;p&gt;What we measured:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One block line over UDP 514: it reached &lt;code&gt;archives.log&lt;/code&gt;; &lt;code&gt;alerts.json&lt;/code&gt; stayed empty.&lt;/li&gt;
&lt;li&gt;18 block lines from one address within the window: the 18th came back as 87702, level 10.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So a working integration shows nothing until one address is blocked 18 times in 45 seconds.&lt;/p&gt;

&lt;p&gt;One trap while testing: &lt;code&gt;wazuh-logtest&lt;/code&gt; prints &lt;em&gt;"Alert to be generated"&lt;/em&gt; for 87701 anyway. Logtest does not apply &lt;code&gt;no_log&lt;/code&gt;. Check &lt;code&gt;alerts.json&lt;/code&gt;, not logtest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reason 2: RFC 5424 lines match no decoder
&lt;/h2&gt;

&lt;p&gt;pfSense can log in BSD (RFC 3164) or RFC 5424 format. The stock decoder is keyed on &lt;code&gt;&amp;lt;program_name&amp;gt;filterlog&amp;lt;/program_name&amp;gt;&lt;/code&gt;, which Wazuh's pre-decoder reads from a BSD-style header. Over UDP on 4.14.7:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Line sent&lt;/th&gt;
&lt;th&gt;Decoder&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Oct  2 15:00:01 pfSense filterlog[12345]: 5,,,1000000103,igb1,match,block,…&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;pf&lt;/code&gt;: srcip, dstip, action read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;2026-10-02T15:00:01.123456+07:00 pfSense.home.arpa filterlog[12345]: …&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;pf&lt;/code&gt;: read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;&amp;lt;134&amp;gt;1 2026-10-02T15:00:01.123456+07:00 pfSense.home.arpa filterlog 12345 - - …&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No decoder matched&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;With RFC 5424, &lt;code&gt;archives.log&lt;/code&gt; even shows the manager's own host name instead of the firewall's: the header was never parsed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Send BSD format.&lt;/strong&gt; In pfSense, &lt;em&gt;Status → System Logs → Settings&lt;/em&gt;, set the log message format to BSD (RFC 3164).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide which blocks you want as alerts.&lt;/strong&gt; To alert on every block, overwrite 87701 without &lt;code&gt;no_log&lt;/code&gt; in &lt;code&gt;/var/ossec/etc/rules/local_rules.xml&lt;/code&gt;:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;group&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"local,pfsense,"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;rule&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"87701"&lt;/span&gt; &lt;span class="na"&gt;level=&lt;/span&gt;&lt;span class="s"&gt;"5"&lt;/span&gt; &lt;span class="na"&gt;overwrite=&lt;/span&gt;&lt;span class="s"&gt;"yes"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;if_sid&amp;gt;&lt;/span&gt;87700&lt;span class="nt"&gt;&amp;lt;/if_sid&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;action&amp;gt;&lt;/span&gt;block&lt;span class="nt"&gt;&amp;lt;/action&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;description&amp;gt;&lt;/span&gt;pfSense firewall drop event.&lt;span class="nt"&gt;&amp;lt;/description&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;group&amp;gt;&lt;/span&gt;firewall_block,pci_dss_1.4,gpg13_4.12,hipaa_164.312.a.1,nist_800_53_SC.7,tsc_CC6.7,tsc_CC6.8,&lt;span class="nt"&gt;&amp;lt;/group&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;/rule&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/group&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On 4.14.7 this loaded with no warnings, and the same single block line that produced nothing produced one 87701 alert. An internet-facing firewall blocks a lot, so many people prefer a narrower child rule (one interface, one port) and leave 87701 alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;Measured on a Wazuh 4.14.7 manager container, with syslog over UDP from 127.0.0.1 and in &lt;code&gt;wazuh-logtest&lt;/code&gt;. We did not run a pfSense box: the lines were written in the &lt;code&gt;filterlog&lt;/code&gt; format for one IPv4 TCP block and one pass. Not measured: IPv6, other pfSense programs, OPNsense, other Wazuh versions.&lt;/p&gt;

&lt;p&gt;Full note with the one-minute checks: &lt;a href="https://atkvn.com/fix-pfsense-logs-not-showing-in-wazuh.html" rel="noopener noreferrer"&gt;https://atkvn.com/fix-pfsense-logs-not-showing-in-wazuh.html&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Dong Nguyen, ATK New Technology. We check and fix Wazuh rules for people who run it.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>wazuh</category>
      <category>pfsense</category>
      <category>security</category>
    </item>
    <item>
      <title>StressFreeFantasy: Lineup Optimization with TabPFN and Real NFL Data</title>
      <dc:creator>Lucas</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:07:11 +0000</pubDate>
      <link>https://dev.to/lukeinthatfluke/stressfreefantasy-lineup-optimization-with-tabpfn-and-real-nfl-data-2c4d</link>
      <guid>https://dev.to/lukeinthatfluke/stressfreefantasy-lineup-optimization-with-tabpfn-and-real-nfl-data-2c4d</guid>
      <description>&lt;p&gt;&lt;em&gt;This project is a submission for the Hacktoberfest: Build for a Friend DEV Challenge, targeting the **Best Use of TabPFN&lt;/em&gt;* category.*&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Context &amp;amp; Motivation
&lt;/h2&gt;

&lt;p&gt;Me and my brother play in the same competitive fantasy football league, which means my default setting on Sundays is actively rooting for his team to implode. &lt;/p&gt;

&lt;p&gt;That said, watching him agonize every single week over two flex players projected within 0.3 points of each other got painful to witness. &lt;/p&gt;

&lt;p&gt;Default platform projections have known limitations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Static estimates:&lt;/strong&gt; They rarely adjust quickly for game scripts (Vegas totals and spreads).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Positional matchups:&lt;/strong&gt; Opponent defensive strength against specific positions is often averaged out or ignored.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Point estimates only:&lt;/strong&gt; They show a single expected number rather than a distribution. In reality, whether you need a high floor (P10) to protect a lead or a high ceiling (P90) to chase upside completely changes who you should start.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I built &lt;strong&gt;StressFreeFantasy&lt;/strong&gt; to automate these decisions. It syncs with private ESPN leagues, conditions an in-context tabular foundation model on historical NFL data, and outputs an optimized lineup with floor/ceiling estimates. (Though whether I let him use it against me when we play head-to-head remains an open question.)&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Why TabPFN?
&lt;/h2&gt;

&lt;p&gt;Standard tabular models (like XGBoost or LightGBM) require training pipelines, feature scaling, and hyperparameter tuning. More importantly, when leagues use non-standard scoring rules (Full PPR, Half PPR, TE premium, big-play bonuses), a traditional model has to be retrained from scratch for those specific point values.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TabPFN&lt;/strong&gt; handles this differently:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;In-Context Adaptation:&lt;/strong&gt; As a prior-data fitted foundation model, it runs in-context learning in a single forward pass. By feeding it 2,000 historical NFL player-weeks recalculated to the user's specific scoring settings, TabPFN adapts instantly without gradient steps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quantile Outputs:&lt;/strong&gt; TabPFN natively supports querying specific quantiles (&lt;code&gt;quantiles=[0.1, 0.9]&lt;/code&gt;), providing calibrated P10 (floor) and P90 (ceiling) values rather than just a noisy mean.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Why Open Innovation Matters Here
&lt;/h3&gt;

&lt;p&gt;Relying on an open-weight, locally runnable tabular foundation model makes all the difference compared to closed cloud APIs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Zero API Cost &amp;amp; Unlimited Inference:&lt;/strong&gt; Simulating thousands of matchup permutations and testing historical holdouts incurs zero per-token API charges.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data Privacy:&lt;/strong&gt; League authentication tokens (&lt;code&gt;espn_s2&lt;/code&gt;, &lt;code&gt;SWID&lt;/code&gt;) and private league rosters stay completely local on the user's machine—never routed through a third-party server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offline-Ready In-Context Learning:&lt;/strong&gt; TabPFN runs locally on a consumer laptop CPU/GPU without depending on external proprietary server uptimes right before Sunday kickoffs.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  3. Architecture &amp;amp; Data Pipeline
&lt;/h2&gt;

&lt;p&gt;The project is built with Streamlit and ties together three data sources:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    ┌────────────────────────────┐
                    │    nflverse / nflreadpy    │
                    │  5 Seasons Historical Data │
                    └─────────────┬──────────────┘
                                  │
                                  ▼
┌──────────────────────┐   ┌──────────────┐   ┌──────────────────────┐
│  Private ESPN League │──▶│   TabPFN     │──▶│ Streamlit Dashboard  │
│  (espn-api / S2&amp;amp;SWID)│   │ In-Context   │   │ Lineup + Floor/Ceil  │
└──────────────────────┘   │ Regressor    │   └──────────────────────┘
                           └──────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Feature Engineering
&lt;/h3&gt;

&lt;p&gt;Features are computed without data leakage, using strictly pre-game rolling historical signals:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Player baseline:&lt;/strong&gt; 16-game rolling average (&lt;code&gt;Avg_Pts&lt;/code&gt;), previous week's score (&lt;code&gt;Last_Week_Pts&lt;/code&gt;), and short-term form (&lt;code&gt;Form_L3&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Volume:&lt;/strong&gt; Rolling 3-game touches and targets (&lt;code&gt;Volume_Proj&lt;/code&gt;), which is historically more stable than point totals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Game environment:&lt;/strong&gt; Opponent defensive ranking against the player's specific position (&lt;code&gt;Opp_Def_Rank&lt;/code&gt;), dome status, and Vegas implied team total derived from over/under and point spread:
$$\text{Implied Team Total} = \frac{\text{Vegas Total} + \text{Spread}}{2}$$&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  4. Benchmark &amp;amp; Validation
&lt;/h2&gt;

&lt;p&gt;Fantasy projections have high variance due to random touchdown variance, making raw MAE an incomplete metric. What matters for a manager is whether the model ranks the right player higher in a head-to-head Start/Sit call.&lt;/p&gt;

&lt;p&gt;We set up an out-of-time temporal test, fitting on earlier seasons and testing on an unseen holdout season ($N = 600$ test player-weeks):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Baseline (Avg)&lt;/th&gt;
&lt;th&gt;Baseline (Form)&lt;/th&gt;
&lt;th&gt;TabPFN (Ours)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;MAE&lt;/strong&gt; &lt;em&gt;(lower is better)&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;6.84 pts&lt;/td&gt;
&lt;td&gt;7.12 pts&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;5.91 pts&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Correlation ($r$)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.44&lt;/td&gt;
&lt;td&gt;0.38&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.58&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pairwise Start/Sit Accuracy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;52.8%&lt;/td&gt;
&lt;td&gt;51.4%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;63.7%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In pairwise head-to-head matchups between players at the same position, TabPFN selected the higher-scoring option &lt;strong&gt;63.7%&lt;/strong&gt; of the time.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Interface &amp;amp; Usage
&lt;/h2&gt;

&lt;p&gt;The Streamlit UI provides:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;ESPN Sync:&lt;/strong&gt; Imports rosters, schedules, and custom scoring tables using league credentials (&lt;code&gt;espn_s2&lt;/code&gt;, &lt;code&gt;SWID&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Manual Adjustments:&lt;/strong&gt; An interactive data editor to tweak expected volume, health status, or weather/stadium conditions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lineup Optimizer:&lt;/strong&gt; Greedily fills position slots (QB, RB, WR, TE, FLEX) sorted by projected output and matchup edge.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj0vk6v85lq64x1oauj4r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj0vk6v85lq64x1oauj4r.png" alt="StressFreeFantasy Streamlit App" width="800" height="314"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Testing &amp;amp; Feedback
&lt;/h2&gt;

&lt;p&gt;Since we play in the same league, I initially tested the engine on my own roster to see how the recommendations differed from ESPN's defaults. Rather than relying on a single static projection, having direct access to calibrated P10 (floor) and P90 (ceiling) intervals made marginal Start/Sit calls immediately clearer.&lt;/p&gt;

&lt;p&gt;After seeing the floor/ceiling spreads and the pairwise accuracy numbers, I showed the tool to my brother to help him with his own weekly lineup headaches. He now uses it as a sanity check before kickoff rather than overthinking marginal calls—which unfortunately means his roster is noticeably harder to beat when we play each other.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Repository &amp;amp; AI Transparency
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/lucasantonsson/StressFreeFantasy" rel="noopener noreferrer"&gt;lucasantonsson/StressFreeFantasy&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stack:&lt;/strong&gt; Python, Streamlit, TabPFN, &lt;code&gt;nflreadpy&lt;/code&gt;, &lt;code&gt;espn-api&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Collaboration:&lt;/strong&gt; Code generation, pipeline scaffolding, and debugging were assisted by Claude and Gemini, while system design, logic requirements, and validation were directed independently.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>tabpfn</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
