<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: soy</title>
    <description>The latest articles on DEV Community by soy (@soytuber).</description>
    <link>https://dev.to/soytuber</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3812665%2F761376f9-10b8-4c2c-b6cb-af00f9fa48ab.jpeg</url>
      <title>DEV Community: soy</title>
      <link>https://dev.to/soytuber</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/soytuber"/>
    <language>en</language>
    <item>
      <title>Claude Code's Background-Agent Wait Ceiling: A $4.19 Lesson in Headless Design</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Mon, 28 Sep 2026 09:16:52 +0000</pubDate>
      <link>https://dev.to/soytuber/claude-codes-background-agent-wait-ceiling-a-419-lesson-in-headless-design-5d52</link>
      <guid>https://dev.to/soytuber/claude-codes-background-agent-wait-ceiling-a-419-lesson-in-headless-design-5d52</guid>
      <description>&lt;p&gt;A headless Claude Code pipeline that fans work out to background subagents can appear to succeed, exit code 0, a clean JSON result, while producing zero output and burning real API cost. This repository's own cron-driven news pipeline hit that failure mode once, and the fix was a single environment variable most headless setups never touch.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;scripts/agent_news.sh drives the daily news pipeline as a single headless call: &lt;code&gt;claude -p "$PROMPT" --model claude-sonnet-5 --allowedTools Read Write Edit Bash WebFetch WebSearch Glob Grep --permission-mode acceptEdits --output-format json&lt;/code&gt;. The prompt tells the agent to follow scripts/agent_news.md, which instructs it to fetch roughly 51 RSS/Atom/changelog sources in parallel, since fetching them one at a time takes too long. In practice, that instruction gets carried out by spawning multiple background Task subagents, one per group of sources, rather than issuing dozens of foreground tool calls inside a single turn.&lt;/p&gt;

&lt;p&gt;On 2026-09-15 18:00 JST, that run spawned 7 background subagents. The harness has a default wait ceiling on how long a &lt;code&gt;claude -p&lt;/code&gt; invocation will block on outstanding background work; six of the seven subagents finished, the seventh was still running when the ceiling was reached, and the harness terminated it (subagent_stats.killed.system: 1). The parent process then returned normally, exit_code=0, a well-formed JSON result, with the model's own final message reading only '最後のグループ(hardware B + Google系)の完了を待っています' ('still waiting on the last group'). Zero articles were saved. The run had already spent $4.19 (haiku-4-5 subagent calls: $1.67; sonnet-5 orchestration: $2.52) by the time it gave up.&lt;/p&gt;

&lt;p&gt;The fix, applied the same day, was to stop relying on the default wait ceiling and pin it explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 0 にすると無制限に待つので使わない。上限そのものが停止の歯止めになる。&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;1800000&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;30 minutes (1,800,000 ms) instead of the harness's implicit 600,000 ms default. A manual re-run at 18:22:45 that day spawned 5 background subagents, all of which completed inside the new window, and produced all 5 category articles plus the Deep Dive, at a cost of $5.76 for the successful run, on top of the $4.19 already spent on the failed one.&lt;/p&gt;

&lt;h2&gt;
  
  
  When this helps
&lt;/h2&gt;

&lt;p&gt;Anyone running claude -p headless with tool access to spawn background Task subagents, a common pattern for fanning out independent fetches, searches, or file scans, should pin CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS explicitly rather than trusting the harness default. The risk is specific to background subagents: a single foreground turn that calls many tools in parallel (as this same task did today, fetching over 50 sources directly inside one turn instead of spawning subagents) does not hit this ceiling at all, because there is no separate background wait to time out. The ceiling only matters once background Task calls enter the picture, nightly cron jobs, long-running research fan-outs, or any headless pipeline where a human isn't watching the terminal to notice a stall. Interactive sessions are lower risk: a stuck background agent is visible in the UI, and a human can wait, kill it, or retry immediately, instead of an exit=0 silently reaching a notification script that only checks the log for known failure markers, not whether the ceiling was hit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measured here
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;2026-09-15 18:00 JST(タイムアウト失敗): background subagent 7起動・6完了・1件がシステムによりkill、exit=0、記事0件保存、コスト$4.19(claude-haiku-4-5: $1.67 / claude-sonnet-5: $2.52)&lt;/li&gt;
&lt;li&gt;同日18:22:45(CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=1800000設定後の再実行): background subagent 5起動・5完了、記事5カテゴリ+Deep Dive保存、コスト$5.76(claude-haiku-4-5: $1.83 / claude-sonnet-5: $3.93)&lt;/li&gt;
&lt;li&gt;ハーネスの既定の待機上限は600,000ms(10分)。ログの実測メッセージ「Background tasks still running after 600s; terminating.」より。&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Is it worth it
&lt;/h2&gt;

&lt;p&gt;Setting the wait ceiling costs nothing and pays for itself the first time a background fan-out runs long: raising it from the harness default to 30 minutes turned a guaranteed-timeout failure into a normal, billable success on this pipeline's very next attempt. The harder lesson is architectural, not the env var itself, deciding whether a multi-source fetch belongs in background subagents at all. Today's run of this same procedure used direct, foreground parallel tool calls instead of Task subagents for the source fetch, and never touched the ceiling. Background subagents earn their keep when the work needs an isolated context window or must keep running after the orchestrating turn ends; for a bounded, known list of sources fetched once, foreground parallel calls remove the timeout risk entirely rather than just widening the window it can fail in. Worth adopting either way: pin the ceiling if background subagents are in use, and ask first whether they need to be.&lt;/p&gt;

&lt;p&gt;Written while running Claude Code unattended in production every day. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>agents</category>
      <category>automation</category>
      <category>cli</category>
    </item>
    <item>
      <title>A Missing Same-Day Guard Lets Two claude -p Cron Runs Duplicate Articles</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Sun, 27 Sep 2026 09:34:57 +0000</pubDate>
      <link>https://dev.to/soytuber/a-missing-same-day-guard-lets-two-claude-p-cron-runs-duplicate-articles-1495</link>
      <guid>https://dev.to/soytuber/a-missing-same-day-guard-lets-two-claude-p-cron-runs-duplicate-articles-1495</guid>
      <description>&lt;p&gt;A daily pipeline that runs &lt;code&gt;claude -p&lt;/code&gt; headless via cron, with a separate manual invocation of the same task, exposed a same-day duplicate-write hazard live: &lt;code&gt;save-article&lt;/code&gt; has no per-day-per-category uniqueness check, so two unrelated runs can each successfully write a different article to the same category on the same day. The pipeline's own agent noticed this exact hazard once before and could not act on it, because a headless run has no one to answer its own confirmation prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;The pipeline in question (&lt;code&gt;scripts/agent_news.sh&lt;/code&gt;) runs &lt;code&gt;claude -p&lt;/code&gt; under cron at 18:00 JST daily, with a fixed tool allowlist (&lt;code&gt;Read Write Edit Bash WebFetch WebSearch Glob Grep&lt;/code&gt;), &lt;code&gt;--permission-mode acceptEdits&lt;/code&gt;, and JSON output logged to &lt;code&gt;logs/agent_news_YYYYMMDD.log&lt;/code&gt;. The task instructs the agent to fetch dozens of release feeds in parallel using background subagents, then save up to five category articles via &lt;code&gt;python scripts/news_db.py save-article&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A first failure mode surfaced on 2026-09-15: the agent spawned 7 background subagents, but the default background-wait ceiling (600s) expired before they finished. One subagent was killed by the system, the run exited 0 having spent $4.19 and saved zero articles — a silent, billed no-op. The fix was raising the ceiling via &lt;code&gt;CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=1800000&lt;/code&gt; (30 minutes), set as an exported env var in the script itself so the limit is a deliberate stop, not an accident.&lt;/p&gt;

&lt;p&gt;A second, different failure mode was caught live on 2026-09-27: a manual &lt;code&gt;claude&lt;/code&gt; session ran the same daily procedure at the same time cron's own &lt;code&gt;claude -p&lt;/code&gt; (PID 26854, started 18:00:01 JST) was independently doing the same thing. &lt;code&gt;ps aux&lt;/code&gt; confirmed both processes coexisting more than 30 minutes in. Checking &lt;code&gt;sqlite3 data/media.db ".schema articles"&lt;/code&gt; showed the only uniqueness constraint is &lt;code&gt;UNIQUE(slug, content_type)&lt;/code&gt; — slugs are derived from the generated title, so two different picks for the same category on the same day get two different slugs and both insert cleanly. The pipeline's own agent hit this exact scenario on 2026-09-26 (visible in &lt;code&gt;logs/agent_news_20260926.log&lt;/code&gt;), correctly identified the risk, proposed killing the competing process, and then had no mechanism to get an answer, since a headless &lt;code&gt;-p&lt;/code&gt; run cannot pause for confirmation.&lt;/p&gt;

&lt;h2&gt;
  
  
  When this helps
&lt;/h2&gt;

&lt;p&gt;This affects anyone running &lt;code&gt;claude -p&lt;/code&gt; on a fixed schedule (cron, systemd timer) for a task that produces persisted, idempotency-assumed output — daily digests, reports, or scraped-and-saved content — where the same task can also be triggered ad hoc or manually. If the underlying save step has no natural key beyond a generated slug or filename, two overlapping runs will both succeed and both write, silently doubling output for that period.&lt;/p&gt;

&lt;p&gt;It does not affect single-operator interactive sessions with no scheduled counterpart, or pipelines where the storage layer already enforces a &lt;code&gt;(category, date)&lt;/code&gt;-style uniqueness constraint or checks existing rows before writing. It also does not affect the separate background-subagent timeout issue from 2026-09-15 — that one is about background wait limits inside a single run, not about two runs colliding.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measured here
&lt;/h2&gt;

&lt;p&gt;2026-09-27 18:31 JST時点の実測: &lt;code&gt;ps aux&lt;/code&gt;でcron起動の&lt;code&gt;claude -p&lt;/code&gt;(PID 26854)が18:00:01起動から約31分経過してもまだ実行中であることを確認。&lt;code&gt;sqlite3 data/media.db "select id, category, created_at from articles where date(created_at)=date('now')"&lt;/code&gt;で、対話セッション側が保存した5件(id 1065–1069、いずれもcreated_at 2026-09-27 09:29台UTC)のみが存在することを確認。2026-09-15の障害ログ(&lt;code&gt;logs/agent_news_20260915.log.failed-1800&lt;/code&gt;)には、7個のサブエージェント起動・1個がシステムにkillされる・exit=0・total_cost_usd 4.192997900000001という実測値が記録されている。2026-09-26のログ(&lt;code&gt;logs/agent_news_20260926.log&lt;/code&gt;)には、エージェント自身が並行実行を検知した旨のresultテキストと、total_cost_usd 6.929127000000001、spawned 5・completed 5・failed 0という実測値が記録されている。&lt;/p&gt;

&lt;h2&gt;
  
  
  Is it worth it
&lt;/h2&gt;

&lt;p&gt;The 2026-09-15 fix (raising the background-wait ceiling) was the right response to that specific failure, but it does not address this second, separate hazard. The gap here is architectural, not a matter of remembering not to run the script manually near 18:00: &lt;code&gt;save-article&lt;/code&gt; and &lt;code&gt;save-featured&lt;/code&gt; should check for an existing row with the same &lt;code&gt;(category, date)&lt;/code&gt; before doing any research, or the &lt;code&gt;articles&lt;/code&gt; table should carry a partial unique index on &lt;code&gt;(category, date(created_at), content_type)&lt;/code&gt; where content_type is the news-digest type, so a second concurrent run fails fast at the write instead of silently succeeding.&lt;/p&gt;

&lt;p&gt;Until that guard exists, treat any manual run of this pipeline as unsafe within the cron's daily window — check &lt;code&gt;ps aux&lt;/code&gt; for a running &lt;code&gt;agent_news.sh&lt;/code&gt;/&lt;code&gt;claude -p&lt;/code&gt; process first, or hold the manual run until the scheduled one has logged its result. This is a five-minute check that fully avoids the failure mode observed here.&lt;/p&gt;

&lt;p&gt;Written while running Claude Code unattended in production every day. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>automation</category>
      <category>agents</category>
      <category>cli</category>
    </item>
    <item>
      <title>Ising-Model Optimization Cuts LLM Depth-Pruning Accuracy Loss on Llama-3.3-70B by 23 MMLU Points</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Fri, 25 Sep 2026 09:16:49 +0000</pubDate>
      <link>https://dev.to/soytuber/ising-model-optimization-cuts-llm-depth-pruning-accuracy-loss-on-llama-33-70b-by-23-mmlu-points-5e7j</link>
      <guid>https://dev.to/soytuber/ising-model-optimization-cuts-llm-depth-pruning-accuracy-loss-on-llama-33-70b-by-23-mmlu-points-5e7j</guid>
      <description>&lt;p&gt;A block-pruning method published today reframes which transformer layers to remove as a constrained binary optimization problem mapped onto an Ising spin glass, rather than scoring each block independently. Tested on Llama-3.3-70B-Instruct at 50% depth compression, it holds MMLU at 76.9 without any retraining, against 54.0 for a standard block-importance baseline. The code is open source, which means teams already doing depth pruning for inference cost savings can test the claim directly rather than take it on faith.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;Researchers at Multiverse Computing released a new approach to transformer block pruning that treats the choice of which layers to keep as a constrained binary optimization (CBO) problem rather than a per-block scoring exercise. The method performs a second-order Taylor expansion of the model's loss around its current weights, producing a Hessian matrix over the set of blocks. Diagonal entries of that Hessian capture how much removing a single block alone would hurt performance -- the same information most existing pruning methods rely on -- but the off-diagonal entries capture how pairs of blocks interact when removed together, which single-block scoring throws away entirely.&lt;/p&gt;

&lt;p&gt;That Hessian maps directly onto the energy function of an Ising spin glass, where each block is a spin (kept or removed) and the off-diagonal terms become spin-spin couplings. Finding a low-energy spin configuration then becomes a proxy for finding a pruned model that will score well, and the team solves this with standard combinatorial optimization techniques rather than any physics-specific hardware.&lt;/p&gt;

&lt;h3&gt;
  
  
  Results
&lt;/h3&gt;

&lt;p&gt;The team tested four models -- Llama-3.1-8B-Instruct, Qwen3-14B, Llama-3.3-70B-Instruct, and the hybrid Mamba2/attention/MoE NVIDIA-Nemotron-3-Nano-30B -- across MMLU, AIME25, and GPQA. The headline result removes 40 of Llama-3.3-70B-Instruct's 80 blocks (50% depth compression) and reaches 76.9 on MMLU with zero retraining, against 54.0 for a standard block-influence baseline -- a 23-point gap at the same compression ratio.&lt;/p&gt;

&lt;p&gt;One counterintuitive finding: after retraining, some higher-energy ('excited state') spin configurations outperformed the lowest-energy solution the optimizer found, suggesting the energy landscape correlates with post-pruning quality but doesn't fully determine it. The implementation is open-sourced at &lt;code&gt;CompactifAI/Block_removal_through_constrained_binary_optimization&lt;/code&gt; on GitHub.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this affects
&lt;/h2&gt;

&lt;p&gt;Teams running large open-weight models (70B and up) who need to cut inference cost through depth pruning are the direct audience -- this method's main claim is retaining far more accuracy at the same compression ratio than the block-importance scoring most existing pruning tools use, with no retraining required to see the initial gain. It matters most at aggressive compression ratios (50% and beyond), where single-block scoring methods degrade sharply because they ignore how removed blocks interact with each other.&lt;/p&gt;

&lt;p&gt;It's less relevant for teams doing light pruning (10-20%) where baseline methods already perform adequately, or for teams working with models under roughly 8B parameters where the four tested model sizes don't provide direct evidence. Teams using architectures outside the tested set -- pure Mamba or pure MoE without the hybrid pattern tested here -- should treat the results as suggestive rather than confirmed for their specific architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;Adopt this for evaluation, not production, today. The reported numbers come from the authors' own benchmarks without independent replication, and the method's core claim -- that pairwise block interactions matter more than single-block importance -- is exactly the kind of result that benefits from a second team reproducing it on a different model family before teams bet inference infrastructure on it.&lt;/p&gt;

&lt;p&gt;That said, the bar to test it yourself is low: the code is open source, the method requires no retraining to get an initial pruned model, and it slots in as a drop-in alternative to whatever block-importance scoring a team's existing pruning pipeline already uses. Teams currently accepting the accuracy hit from standard block-importance pruning at 50%+ compression have little to lose by running this against their own model and benchmark suite before deciding whether to switch.&lt;/p&gt;

&lt;p&gt;Tracked daily from official release feeds and vendor changelogs. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Vulnerability Research Finds 76 Flaws in Managed PostgreSQL Hardening Extensions</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Thu, 24 Sep 2026 09:11:21 +0000</pubDate>
      <link>https://dev.to/soytuber/vulnerability-research-finds-76-flaws-in-managed-postgresql-hardening-extensions-34a2</link>
      <guid>https://dev.to/soytuber/vulnerability-research-finds-76-flaws-in-managed-postgresql-hardening-extensions-34a2</guid>
      <description>&lt;p&gt;Vulnerability researcher Mehmet Ince published Part 2 of a six-part series documenting 76 flaws across the security-hardening extensions that Azure, AWS Aurora, Aiven, Supabase, PlanetScale, Xata, Google AlloyDB and NeonDB use to fence in the superuser-like role they hand customers on managed PostgreSQL. The bugs share a small set of root causes, and most vendors have not fixed them yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;Ince's report groups the 76 findings into four recurring bug classes. The first is foreign data wrapper (FDW) validator abuse: vendor code typically switches &lt;code&gt;current_user&lt;/code&gt; to superuser, delegates to PostgreSQL core, then restores the original role, but never checks that the validator function's OID is actually safe — a malicious validator function can call &lt;code&gt;CREATE ROLE ... SUPERUSER&lt;/code&gt; while running in that elevated window. The second is a search_path OID-switching attack, where an attacker plants a same-named function (for example &lt;code&gt;postgres_fdw_validator&lt;/code&gt;) earlier in the search path, so the elevated context resolves a different OID than the one that was checked. The third is an lo_export() bypass: instead of the blocked function itself, the attacker updates &lt;code&gt;pg_catalog.pg_proc.prosrc&lt;/code&gt; and creates a new &lt;code&gt;LANGUAGE internal&lt;/code&gt; function that redirects to the real &lt;code&gt;lo_export&lt;/code&gt; implementation, writes an arbitrary file to disk, loads it as a PostgreSQL extension &lt;code&gt;.so&lt;/code&gt;, and reaches &lt;code&gt;LANGUAGE C&lt;/code&gt; code execution. The fourth class is broader: &lt;code&gt;LANGUAGE internal&lt;/code&gt; type-confusion bugs that give raw read/write access to PostgreSQL's process memory, which Ince argues makes the hardening extensions "effectively useless" as a security boundary. Patch status is uneven — Supabase fixed four of the FDW-related bugs, only two vendors properly protect &lt;code&gt;pg_proc&lt;/code&gt; against the lo_export bypass, and the &lt;code&gt;LANGUAGE internal&lt;/code&gt; memory bugs remain open almost everywhere. The PostgreSQL core team's position, per the report, is that when a provider does not intend to grant full superuser access, the resulting security gap belongs to the provider's service, not to PostgreSQL itself. Part 1 of the series covered a PostGIS memory-corruption bug; Part 3 onward is reported to cover more than 16 PostgreSQL core vulnerabilities that remain under embargo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this affects
&lt;/h2&gt;

&lt;p&gt;Anyone running production workloads on a managed PostgreSQL service that grants a restricted "superuser-like" role instead of true superuser — Azure Database for PostgreSQL, AWS Aurora PostgreSQL, Aiven, Supabase, PlanetScale, Xata, Google AlloyDB or NeonDB — should read this. That role's isolation guarantees rest entirely on the vendor's hardening extension, and this research shows those guarantees have concrete holes today, not hypothetical ones.&lt;/p&gt;

&lt;p&gt;Teams that manage their own PostgreSQL instances with real superuser access are not affected by these specific bugs, since there is no restricted-role boundary to break. Extension authors building similar "safe superuser" abstractions on &lt;code&gt;LANGUAGE internal&lt;/code&gt; or FDW validators should treat this as a design-pattern warning, independent of any single vendor's patch status.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;This is not an upgrade-or-wait decision about a piece of software — it is a trust decision about a security boundary. Given that only a fraction of the 76 findings have confirmed fixes (four FDW bugs at Supabase, &lt;code&gt;pg_proc&lt;/code&gt; protection at two vendors) and the &lt;code&gt;LANGUAGE internal&lt;/code&gt; memory-corruption class remains open almost everywhere, the restricted "superuser" role on most affected platforms should not currently be treated as an isolation guarantee against a determined attacker with access to that role.&lt;/p&gt;

&lt;p&gt;Teams on the named platforms should check with their vendor for a patch timeline before assuming the hardening extension protects them, and should avoid granting the restricted role to less-trusted users or automated agents in the meantime. Because Part 3 of the series is reported to disclose more than 16 additional PostgreSQL core vulnerabilities still under embargo, this is better read as the second chapter of an ongoing disclosure than a closed incident.&lt;/p&gt;

&lt;p&gt;Tracked daily from official release feeds and vendor changelogs. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>postgres</category>
      <category>database</category>
      <category>devops</category>
      <category>sql</category>
    </item>
    <item>
      <title>Claude Opus 5.5 Cuts API Pricing 40% While Beating Opus 5 on Every Benchmark</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Wed, 23 Sep 2026 09:09:38 +0000</pubDate>
      <link>https://dev.to/soytuber/claude-opus-55-cuts-api-pricing-40-while-beating-opus-5-on-every-benchmark-4f9c</link>
      <guid>https://dev.to/soytuber/claude-opus-55-cuts-api-pricing-40-while-beating-opus-5-on-every-benchmark-4f9c</guid>
      <description>&lt;p&gt;Anthropic released Claude Opus 5.5 on September 22, 2026, cutting Claude Platform API pricing from $5/$25 to $4/$20 per million input/output tokens while posting double-digit gains across agentic-coding and knowledge-work benchmarks. Claude Code v2.1.280 and the Anthropic Python SDK v1.8.0 shipped the same day, making Opus 5.5 the new default model and adding MCP tool-list pinning for production agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;Claude Opus 5.5 is the first release in Anthropic's 5.5 family, available immediately as &lt;code&gt;claude-opus-5-5&lt;/code&gt; on the Claude Platform API, Claude.ai, Claude Code, and through AWS, Google Cloud, and Microsoft Azure. Standard pricing drops from Opus 5's $5/$25 per million input/output tokens to $4/$20 — a 40% reduction for typical workloads — and cached-read pricing falls 60% to $0.20 per million tokens. A Fast mode, running roughly 2.5x faster, is priced at $8/$40 per million tokens, and standard output generation is itself more than 30% faster than Opus 5.&lt;/p&gt;

&lt;p&gt;Benchmark gains follow the same pattern across agentic and knowledge-work tasks. On Terminal-Bench 4.0, Opus 5.5 scores 66.4% against Opus 5's 52.3% and Claude Fable 5.1's 55.8%. It reaches 54.4% on FrontierCode v1.1 and 57.8% on CursorBench 4.0. On the knowledge-work benchmark GDPval-AA v2.1 it posts an Elo of 1846, up from Opus 5's 1708, and the AutomationBench pass rate rises to 40.0% from 26.9%. Additional results include 81.8% partial completion on the computer-use benchmark OSWorld 2.0, 67.7% on Humanity's Last Exam with tool use, and 89.0% on the chart-reading benchmark Chartography with tool use.&lt;/p&gt;

&lt;p&gt;Safety work shipped alongside the model: expanded third-party alignment testing by Frontier Design and METR, routing of cybersecurity-related tasks to Opus 4.8, a "preserved thinking" mechanism aimed at resisting distillation, and EU AI Act-compliant watermarking. Anthropic says Claude Sonnet 5.5 and Claude Haiku 5.5 will follow within weeks.&lt;/p&gt;

&lt;p&gt;Claude Code v2.1.280 makes Opus 5.5 the default model the same day, switching Pro and Team Standard plans from Sonnet to Opus by default, and fixes several agentic-loop bugs — including auto mode retrying safety-blocked actions indefinitely and a &lt;code&gt;role 'system' must precede an 'assistant' message&lt;/code&gt; API error that broke conversations every turn. Anthropic Python SDK v1.8.0 adds beta MCP tool-list pinning, letting applications freeze the tool list an MCP server exposes for the duration of a session.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this affects
&lt;/h2&gt;

&lt;p&gt;Teams already budgeting for Opus-class models on agentic coding or long-running tasks get a direct win: same tier, lower price, faster output, and measurably better scores on coding and computer-use benchmarks. Anyone running Claude Code on Pro or Team Standard will notice the default model change immediately, since it now defaults to Opus rather than Sonnet — a cost and latency trade-off worth checking against usage patterns. Developers operating MCP servers in production should look at the new tool-list pinning in the Python SDK, since it closes a class of bugs where a server silently changing its exposed tools could break an in-flight session. Anyone still routing cybersecurity-adjacent tasks through Opus should note Anthropic now routes those specifically to Opus 4.8 rather than 5.5. Teams on Sonnet or Haiku have nothing to act on yet; those tiers arrive in the coming weeks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;Upgrade now for agentic-coding and long-running-task workloads: Opus 5.5 costs 40% less than Opus 5 on standard tokens, generates output more than 30% faster, and outscores Opus 5 on every benchmark Anthropic published, including a jump from 52.3% to 66.4% on Terminal-Bench 4.0. There is no scenario where staying on Opus 5 is cheaper or faster, so the switch is close to a pure win for API users already on the Opus tier. Claude Code users on Pro or Team Standard should confirm the automatic default-model switch to Opus fits their budget, since it changes per-session cost even without any action taken. Teams waiting on Sonnet or Haiku pricing parity should hold off migrating other workloads until the 5.5 versions of those tiers ship, which Anthropic says will happen within weeks.&lt;/p&gt;

&lt;p&gt;Tracked daily from official release feeds and vendor changelogs. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>programming</category>
    </item>
    <item>
      <title>GitHub Actions Cache Can Leak Secrets Through Miri's target/ Directory</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Tue, 22 Sep 2026 09:20:07 +0000</pubDate>
      <link>https://dev.to/soytuber/github-actions-cache-can-leak-secrets-through-miris-target-directory-5gpc</link>
      <guid>https://dev.to/soytuber/github-actions-cache-can-leak-secrets-through-miris-target-directory-5gpc</guid>
      <description>&lt;p&gt;The Rust Security Response Team disclosed on September 21 that Miri writes every environment variable into its target/ build directory, and that caching target/ in GitHub Actions — a default pattern via actions/cache — can expose those variables to workflows running against pull requests. Any repository running cargo miri with a cached target/ directory and secrets in the same job is affected right now, not hypothetically.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;Miri, the interpreter the Rust project uses to catch undefined behavior that normal compilation misses, stores all environment variables into its &lt;code&gt;target/&lt;/code&gt; output directory as part of how it operates internally. That behavior has existed for a while, but it turned into a disclosed vulnerability once a common CI pattern collided with it: caching &lt;code&gt;target/&lt;/code&gt; between CI runs via &lt;code&gt;actions/cache&lt;/code&gt; to speed up builds.&lt;/p&gt;

&lt;p&gt;GitHub Actions caches created from a repository's default branch are readable by workflows triggered against pull requests from that same repository — and, depending on workflow configuration, from forks. That means a cache written by a trusted branch build, with secrets sitting inside &lt;code&gt;target/&lt;/code&gt; because Miri put them there, becomes readable to anyone who can open a pull request. The exploit chain needs three things to line up in the same job: &lt;code&gt;cargo miri&lt;/code&gt; running with secret environment variables present, the &lt;code&gt;target/&lt;/code&gt; directory from that run getting cached, and a PR-triggered workflow with read access to that cache. Where all three hold, an attacker with pull-request access can pull secrets straight out of the cache — and the disclosure notes they could plausibly cover their tracks afterward by overwriting the offending commits.&lt;/p&gt;

&lt;p&gt;The Rust team's own fix narrows what Miri writes into &lt;code&gt;target/&lt;/code&gt; to just &lt;code&gt;CARGO_*&lt;/code&gt; variables (with tokens excluded) plus &lt;code&gt;OUT_DIR&lt;/code&gt;, cutting off the specific leak path at its source. But the disclosure is explicit that this is a partial fix, not a complete one: it addresses what Miri itself writes, not what any other build step might already be caching, so the advisory pairs the code fix with three CI-configuration changes every affected repository owner needs to apply independently — none of which the Miri patch itself can substitute for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this affects
&lt;/h2&gt;

&lt;p&gt;This affects any repository that runs &lt;code&gt;cargo miri&lt;/code&gt; inside GitHub Actions and caches the &lt;code&gt;target/&lt;/code&gt; directory in that same job — a combination common enough that the Rust team treated it as worth a same-day public disclosure rather than a silent patch. It matters most to projects using Miri for CI-gated undefined-behavior checks on crates that also handle deployment secrets, package-registry tokens, or signing credentials somewhere in the same workflow file, since those are exactly the values that end up cached alongside Miri's output.&lt;/p&gt;

&lt;p&gt;It does not affect projects that don't run Miri in CI at all, or that already isolate secret-bearing steps into a separate job with no shared cache from the Miri job. Maintainers of forks that submit pull requests to affected upstream repositories aren't at direct risk themselves, but should be aware the read access their PR workflow gets is exactly the access path the advisory describes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;This is not a wait-and-see advisory. Any repository matching the affected pattern — &lt;code&gt;cargo miri&lt;/code&gt; plus a cached &lt;code&gt;target/&lt;/code&gt; plus secrets in the same job — should treat it as an incident to close today, not a changelog entry to read later. The concrete steps, in order: disable &lt;code&gt;actions/cache&lt;/code&gt; for jobs that invoke Miri, move any secret environment variables so they're scoped only to steps that never touch Miri, clear existing caches on the affected repository, and rotate every secret that could plausibly have been cached before this disclosure. Rust's own upstream fix — restricting what Miri writes to &lt;code&gt;target/&lt;/code&gt; — helps going forward but does not retroactively protect a cache that was already poisoned before the patch landed, so clearing and rotating is not optional even after updating.&lt;/p&gt;

&lt;p&gt;For repositories that don't run Miri in CI at all, there is nothing to do here beyond confirming that's actually true.&lt;/p&gt;

&lt;p&gt;Tracked daily from official release feeds and vendor changelogs. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>rust</category>
      <category>githubactions</category>
      <category>devops</category>
      <category>programming</category>
    </item>
    <item>
      <title>Inside AX: Google's Runtime for Running AI Agents at Kubernetes Scale</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Mon, 21 Sep 2026 09:21:16 +0000</pubDate>
      <link>https://dev.to/soytuber/inside-ax-googles-runtime-for-running-ai-agents-at-kubernetes-scale-4675</link>
      <guid>https://dev.to/soytuber/inside-ax-googles-runtime-for-running-ai-agents-at-kubernetes-scale-4675</guid>
      <description>&lt;p&gt;Google has open-sourced AX, a runtime purpose-built to run agentic tasks — not chat turns, but long-lived, stateful, potentially expensive-if-unwatched processes — at a scale existing infrastructure was never designed for. The release matters because most teams running agents in production today are stitching that infrastructure together themselves, out of cron jobs, containers, and ad hoc budget guards.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;AX treats an agentic task as a first-class unit of infrastructure, not an API call. It's built on Agent Substrate, an internal Google system for agentic runtimes, and organizes work around four primitives: &lt;strong&gt;Task&lt;/strong&gt;, a sandboxed unit of execution with explicit CPU and memory limits; &lt;strong&gt;Workspace&lt;/strong&gt;, declarative setup of a Git repo, an MCP server, and the skills a task needs — Google says AX can prepare this automatically from a plain-English description as well as a manifest; &lt;strong&gt;Gateway&lt;/strong&gt;, network policy for everything a task calls out to, handling allowlisting and credential injection so an agent doesn't hold raw secrets directly; and &lt;strong&gt;Model&lt;/strong&gt;, centralized configuration for which models a task can call, with what parameters, under which secrets.&lt;/p&gt;

&lt;p&gt;Developers declare a task in a YAML manifest and deploy it with &lt;code&gt;ax apply&lt;/code&gt;. Running tasks are then managed with &lt;code&gt;ax watch&lt;/code&gt;, &lt;code&gt;ax ssh&lt;/code&gt;, &lt;code&gt;ax suspend&lt;/code&gt;, and &lt;code&gt;ax delete&lt;/code&gt; — the same verbs an operator would expect from a container orchestrator, applied to a unit of work that can run for hours or days instead of milliseconds.&lt;/p&gt;

&lt;p&gt;The performance claim is the part worth scrutinizing: Google says a single cluster can host billions of concurrent tasks, with sub-second resumption and no cold-start delay, by treating an idle agent's reserved resources as reclaimable capacity rather than dead weight. That's the same idea that made container orchestration viable at scale — overcommit resources, multiplex idle time — applied to a workload that is long-lived, stateful, and bursty in a way the request/response assumptions behind most orchestration tooling were never built for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this affects
&lt;/h2&gt;

&lt;p&gt;Teams running more than a handful of agents in production — anything past a single script triggered by a cron job — are the direct audience: AX targets exactly the isolation, budget-control, and resumption problems that infrastructure teams currently solve with a mix of Kubernetes, custom sandboxing, and manual spend monitoring.&lt;/p&gt;

&lt;p&gt;It matters less for teams still prototyping a single agent workflow, where the overhead of adopting a new runtime primitive outweighs the isolation guarantees. It also isn't a replacement for Google's Agent Development Kit (ADK), which addresses agent logic and orchestration rather than the execution substrate underneath it — the two are complementary rather than competing.&lt;/p&gt;

&lt;p&gt;Platform and infrastructure engineers evaluating how to scale internal agent tooling beyond a proof of concept are the group with the clearest reason to read the source and try it now, rather than wait for a more mature ecosystem to form around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;AX is early: it's a fresh open-source release without a long production track record outside Google, and teams should expect rough edges in the tooling around &lt;code&gt;ax apply&lt;/code&gt;/&lt;code&gt;ax watch&lt;/code&gt; as the ecosystem catches up to the primitives. That argues for evaluating it now in a non-critical internal tool rather than betting a customer-facing agent pipeline on it immediately.&lt;/p&gt;

&lt;p&gt;The underlying idea — a standard runtime that isolates, meters, and resumes agent tasks the way Kubernetes does for services — is worth adopting in principle even for teams that don't pick AX specifically, because the alternative most teams are running today (cron jobs plus containers plus a spreadsheet tracking API spend) doesn't scale past a handful of agents. Teams already deep in a Kubernetes-based custom setup should pilot AX on a new project rather than migrate an existing one; teams starting fresh have less reason not to start here.&lt;/p&gt;

&lt;p&gt;Tracked daily from official release feeds and vendor changelogs. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>devops</category>
      <category>programming</category>
    </item>
    <item>
      <title>PostgreSQL 19 Delayed: Why 53 Features Got Reverted Before Beta 4</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Sun, 20 Sep 2026 09:12:52 +0000</pubDate>
      <link>https://dev.to/soytuber/postgresql-19-delayed-why-53-features-got-reverted-before-beta-4-1b9c</link>
      <guid>https://dev.to/soytuber/postgresql-19-delayed-why-53-features-got-reverted-before-beta-4-1b9c</guid>
      <description>&lt;p&gt;PostgreSQL 19 was set for its usual fall release, but the project has pushed the date back by weeks to months after committers pulled 53 features since last June's beta — including SQL/PGQ property graph queries and long-requested partition merge/split support. The pattern behind the reversions, more than the delay itself, is the part worth reading closely.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;Beta 4, originally scheduled for September 24, 2026, has slipped along with the rest of the fall timeline; the project now expects the final release "weeks and maybe even months" later than planned.&lt;/p&gt;

&lt;p&gt;The reverted feature list runs to 53 items pulled since the June 2025 beta, but two categories stand out. The first is functionality cut for design reasons: SQL/PGQ property graph query support went back over broader design and readiness concerns, and ALTER TABLE MERGE/SPLIT PARTITION(S) — a partition-management feature requested for years — was pulled after committers judged its design problems too late to fix within this cycle. UPDATE/DELETE FOR PORTION OF, aimed at temporal range columns, was reverted for similar reasons.&lt;/p&gt;

&lt;p&gt;The second category is more concerning for anyone tracking PostgreSQL's release discipline: correctness bugs caught after the code had already been committed. GROUP BY ALL is the clearest example — post-commit review found it "missed special handling of ORDER BY entries with nondefault equality semantics, producing wrong results," meaning queries using the feature could silently return incorrect data rather than merely fail to compile. A default of lz4 for TOAST compression was also reverted, in this case for missing buildfarm test coverage that would need to be rebuilt from scratch.&lt;/p&gt;

&lt;p&gt;The rest of the list is long: non-text output formats for pg_dumpall, cascading ON ERROR handling in JSON_TABLE, fast domain defaults, logical replication snapshots, nested query tracking, and online data checksums all got pulled as well. Coverage of the delay notes that AI-assisted bug-finding tools played a role in surfacing some of these problems — after the code was already merged, not before. PostgreSQL's release process runs through roughly 30 committers working the open commitfest queue and the pgsql-hackers mailing list; no individual committer is named for any specific reversion.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this affects
&lt;/h2&gt;

&lt;p&gt;Teams that had scoped a project around SQL/PGQ's graph queries or the new partition MERGE/SPLIT syntax should assume neither lands in PostgreSQL 19 and plan around PostgreSQL 20 instead, since both were cut for unresolved design issues rather than a fixable bug. Anyone testing GROUP BY ALL against pre-release builds should discard those results; the construct produced wrong answers under certain ORDER BY conditions and was pulled entirely, so no version of it ships in 19.&lt;/p&gt;

&lt;p&gt;Operators who don't touch any of the specific reverted features see little direct impact — the delay affects the release date, not the stability of features that did survive to the current beta. The likelier practical effect is indirect: a release cycle that let 53 changes reach beta before reversion, several for correctness rather than design reasons, is a data point worth factoring into how early an organization tests against PostgreSQL pre-releases.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;There's no PostgreSQL 19 to install yet, so the immediate decision isn't upgrade-or-wait — it's what to plan for. Anything built on SQL/PGQ or the new partition merge/split syntax should be re-scoped now, on the assumption that PostgreSQL 20 is the earliest realistic target for either. For everything else, the safer posture is to keep tracking beta builds for correctness issues rather than treating a beta tag as a stability signal, since this cycle demonstrated that features can reach beta with bugs serious enough to produce wrong query results.&lt;/p&gt;

&lt;p&gt;Once PostgreSQL 19 actually ships — now expected some weeks to months past its original fall date — the delay itself won't be the story worth remembering. The 53 reversions, and how many of them were correctness bugs rather than missed deadlines, are the more useful signal for judging how carefully to test the next PostgreSQL pre-release before relying on it.&lt;/p&gt;

&lt;p&gt;Tracked daily from official release feeds and vendor changelogs. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>postgres</category>
      <category>database</category>
      <category>sql</category>
      <category>opensource</category>
    </item>
    <item>
      <title>PgBouncer Transaction Mode Leaks Session State Between Tenants</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Sat, 19 Sep 2026 09:23:50 +0000</pubDate>
      <link>https://dev.to/soytuber/pgbouncer-transaction-mode-leaks-session-state-between-tenants-5hb7</link>
      <guid>https://dev.to/soytuber/pgbouncer-transaction-mode-leaks-session-state-between-tenants-5hb7</guid>
      <description>&lt;p&gt;A blog post from Seedfast, tested against PostgreSQL 18.6 and PgBouncer 1.25.2, shows that PgBouncer's transaction pooling mode hands a backend connection to a new client without clearing custom session parameters such as app.tenant, letting one client inherit another's row-level-security context. The failure is acute for AI agent workloads: since the MCP specification dropped session state from the protocol on 28 July 2026, each tool call now opens with a bare SET app.tenant outside any transaction, exactly the pattern the pooler mishandles. It is worth reading because the failure mode is silent — RLS policies keep evaluating, just against the wrong tenant.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;PgBouncer's transaction pooling mode returns a backend connection to the pool as soon as a transaction (or, for autocommit statements, a single statement) finishes, and reuses that same backend for the next client that needs one. At handover, PgBouncer only resets a fixed set of tracked parameters — client_encoding, datestyle, timezone, standard_conforming_strings, application_name — plus whatever the configured server_reset_query covers. Anything set with a plain SET on a custom GUC, such as the common app.tenant variable read by a row-level-security policy via current_setting('app.tenant', true), is left in place on the backend and is inherited verbatim by the next client PgBouncer assigns to it.&lt;/p&gt;

&lt;p&gt;The article demonstrates this directly: two client connections opened one after another through PgBouncer in transaction mode, each querying which rows they were allowed to see, and the second connection got the first connection's answer, because the backend still carried the first tenant's app.tenant setting.&lt;/p&gt;

&lt;p&gt;The same handover carries other state readers may not expect: SET ROLE (and therefore current_user), search_path, statement_timeout, temp tables created without ON COMMIT DROP, PREPAREd statements, DECLARE ... CURSOR WITH HOLD, LISTEN subscriptions, and session-level advisory locks taken with pg_advisory_lock(). Only transaction-scoped variants — SET LOCAL, set_config(..., is_local=true), SET LOCAL ROLE, pg_advisory_xact_lock() — are discarded at COMMIT and therefore safe under transaction pooling.&lt;/p&gt;

&lt;p&gt;A second, subtler effect: once a custom parameter has been touched at all on a backend, current_setting(name, true) returns an empty string instead of NULL for the remaining life of that backend, silently breaking IS NULL checks and COALESCE fallbacks that assume an unset variable.&lt;/p&gt;

&lt;p&gt;The root cause behind the current spike in exposure is architectural: the MCP specification removed session state from the protocol on 28 July 2026, so each agent tool call now arrives as an independent request that opens with its own SET app.tenant outside any transaction — precisely the shape PgBouncer's transaction mode cannot safely handle.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this affects
&lt;/h2&gt;

&lt;p&gt;Anyone running PostgreSQL behind PgBouncer (or a similarly behaved pooler) in transaction pooling mode, with multi-tenant row-level security keyed on session-level current_setting() values, is exposed. The article's proof of concept used PostgreSQL 18.6 and PgBouncer 1.25.2, but the mechanism is a documented property of transaction pooling itself, not a version-specific bug. It is especially relevant to teams exposing Postgres to AI agents through MCP tool calls, since the 28 July 2026 protocol change that removed session state encourages exactly the bare SET-then-query pattern that leaks. Teams using session pooling mode or dedicated per-tenant connections are not affected, since PgBouncer only skips the reset behavior in transaction mode. Setups that already wrap every request in an explicit transaction and use SET LOCAL for tenant context are also safe, as are those relying solely on application-level tenant filtering rather than session-variable-driven RLS.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;This is not a wait-and-see item: any multi-tenant deployment using PgBouncer transaction mode with session-variable-based RLS should audit its connection code today, specifically searching for bare SET app.tenant (or equivalent) calls that are not wrapped in BEGIN ... SET LOCAL ... COMMIT. The fix requires no PgBouncer upgrade or config flag — it is an application-side discipline change: issue an explicit transaction per request, set tenant context with SET LOCAL, and let COMMIT clear it, plus swapping session-level advisory locks and LISTEN usage for transaction-safe equivalents where pooled. Teams building or operating MCP-based agent tool servers against pooled Postgres should treat this as urgent, since the protocol's removal of session state pushes exactly the vulnerable pattern. Do not reach for server_reset_query_always as a blanket fix — it forces a DISCARD ALL after every transaction, breaking legitimate multi-statement workflows.&lt;/p&gt;

&lt;p&gt;Tracked daily from official release feeds and vendor changelogs. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>postgres</category>
      <category>security</category>
      <category>connectionpooling</category>
      <category>backend</category>
    </item>
    <item>
      <title>Bonsai 2 27B Squeezes Qwen3.8 27B Into 5.9GB, but Stock llama.cpp Can't Run It</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Fri, 18 Sep 2026 09:15:41 +0000</pubDate>
      <link>https://dev.to/soytuber/bonsai-2-27b-squeezes-qwen38-27b-into-59gb-but-stock-llamacpp-cant-run-it-4685</link>
      <guid>https://dev.to/soytuber/bonsai-2-27b-squeezes-qwen38-27b-into-59gb-but-stock-llamacpp-cant-run-it-4685</guid>
      <description>&lt;p&gt;PrismML has released Bonsai 2 27B, a ternary-quantized version of Qwen3.8 27B that compresses a 53.8GB FP16 model down to 5.9GB while retaining 98.2% of the original's benchmark score. The catch: the release only works with a patched llama.cpp fork carrying custom ternary hybrid-attention kernels — stock llama.cpp either refuses to load the files or, for one of the two quant formats, loads them silently and produces garbage output.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;PrismML has shipped Bonsai 2 27B, a ternary-quantized derivative of Qwen3.8 27B that stores weights as {-1, 0, +1} trits with FP16 group-wise scaling instead of standard integer quantization. The company describes the format as averaging 1.76 effective bits per weight across the two GGUF packings it publishes: PTQ1_0, a dense-trit packing at 5.95GB (1.75 bits/weight), and PQ2_0, a 2-bit-slot packing at 7.21GB (2.13 bits/weight). Both derive from the same underlying "ternary g128" format; a reference FP16 checkpoint (53.8GB) is also published for comparison. That puts the compressed model at roughly a ninth of the original's footprint.&lt;/p&gt;

&lt;p&gt;The model keeps Qwen3.8 27B's 262K-token context window and multimodal input (text plus image), and supports coding, reasoning, vision, and agentic tasks; the vision path adds a ~0.63GB Q8_0 component loaded only when an image is present. Across benchmarks spanning MMLU-Redux, HumanEval+, CharXiv, and others, PrismML reports an overall score of 83.9 against the FP16 baseline — 98.2% retention — with coding at 99.3% retention, math at 99.5%, and vision the weakest category at 96.3%. On inference speed, the company cites up to 143 tokens/second on an RTX 5090 and 46.8 tokens/second on an Apple M5 Max, with RTX 4090 energy use of 0.714 mWh/token, a figure PrismML says is 40% more efficient per token than a full-precision 8B model.&lt;/p&gt;

&lt;p&gt;None of this runs on stock llama.cpp. The ternary hybrid-attention kernels for CUDA and Metal live only in a PrismML-maintained llama.cpp fork; the vendor's own documentation warns that unmodified llama.cpp either refuses the custom quantization types outright or, for PQ2_0 specifically, loads the file without warning and produces garbage output. Ollama, LM Studio, vLLM, and Jan integrations are documented, but all route through that same fork rather than upstream code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this affects
&lt;/h2&gt;

&lt;p&gt;Bonsai 2 27B matters most to people running large open-weight models on a single consumer GPU or Apple silicon with limited memory, since 5.9GB puts a 27B-class model's quality within reach of hardware that could never hold the 53.8GB FP16 checkpoint. Developers already comfortable building llama.cpp from source, or willing to pull PrismML's prebuilt binaries or Docker image, get the compression without much extra friction.&lt;/p&gt;

&lt;p&gt;It matters less — or is actively risky — for anyone relying on a stock, unmodified llama.cpp install or a downstream tool that assumes standard GGUF quantization behavior. The PQ2_0 packing's silent failure mode on unpatched llama.cpp (loading without error and returning garbage) is the kind of bug that would be hard to notice in an automated pipeline. Teams that need predictable behavior across a heterogeneous set of inference backends should treat this as a single-fork dependency, not a drop-in GGUF file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;Worth trying, not worth adopting blindly. The compression numbers are real and unusually well-documented for a quantization release: 98.2% overall retention at roughly a ninth of the file size is a genuinely strong trade, and the RTX 5090 and M5 Max throughput figures make single-GPU or Mac-only deployment of a 27B model plausible for the first time at this quality bar.&lt;/p&gt;

&lt;p&gt;But the dependency on a PrismML-maintained llama.cpp fork — with a documented silent-corruption failure mode on stock builds — is the kind of detail that belongs in a pilot, not a production rollout. Readers who already run custom llama.cpp builds or who are comfortable pinning to a vendor fork should pull it down and benchmark against their current setup. Readers who need mainline llama.cpp, Ollama, or LM Studio compatibility without extra build steps should wait for either upstream support or a second quantization vendor to validate the approach.&lt;/p&gt;

&lt;p&gt;Tracked daily from official release feeds and vendor changelogs. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>Why PostgreSQL 19 Missed Its Fall Release Window</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Thu, 17 Sep 2026 09:20:25 +0000</pubDate>
      <link>https://dev.to/soytuber/why-postgresql-19-missed-its-fall-release-window-2jh</link>
      <guid>https://dev.to/soytuber/why-postgresql-19-missed-its-fall-release-window-2jh</guid>
      <description>&lt;p&gt;PostgreSQL 19 will not ship this fall as planned: 53 features were reverted during an unusually long review cycle that started with last year's first beta, including SQL/PGQ property graphs and the long-awaited ALTER TABLE ... MERGE/SPLIT PARTITION(S) commands. The reversions were driven partly by AI-assisted tooling that found defects, including a correctness bug in GROUP BY ALL, that earlier review cycles likely would have missed until after release. For anyone planning a PostgreSQL upgrade this year, the question is no longer whether to move to 19, but whether to wait for it at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;PostgreSQL 19 entered beta in June 2025 carrying an ambitious feature list, and the project spent the subsequent review cycle finding that several of the largest additions were not ready for a stable release. The 53 reversions announced ahead of Beta 4 (scheduled September 24, 2026) span everything from marquee features to small correctness fixes.&lt;/p&gt;

&lt;p&gt;The highest-profile casualty is SQL/PGQ property graph query support, a SQL standard feature that would have let PostgreSQL run graph queries natively, pulled over concerns that its design and implementation weren't ready as a unit, not over any single isolated bug. ALTER TABLE ... MERGE/SPLIT PARTITION(S), a partition-management capability administrators have asked for across multiple release cycles, was abandoned after design issues in how it interacted with existing partition strategies surfaced too late to fix safely before release. Temporal range support via UPDATE/DELETE ... FOR PORTION OF was reverted for similar readiness reasons.&lt;/p&gt;

&lt;p&gt;Some reversions are pure correctness fixes rather than scope cuts: GROUP BY ALL was found during testing to produce incorrect results in combination with certain ORDER BY clauses, the kind of bug that silently corrupts query results rather than crashing, exactly the category a database project cannot ship with. A planned default switch of TOAST compression to lz4 was pulled for unrelated build-infrastructure reasons rather than a defect in the feature itself.&lt;/p&gt;

&lt;p&gt;What makes this cycle unusual is the tooling behind the discoveries. The post crediting AI-assisted analysis with surfacing several of these defects and generating reproducible test cases points to PostgreSQL 18's August 2026 patch release, which fixed 28 CVEs against a historical average of a couple per release, as the clearest evidence that this isn't a one-off: the project's review process is catching more before release, which means more work before ship, not fewer total defects over the software's life.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this affects
&lt;/h2&gt;

&lt;p&gt;Anyone with a PostgreSQL 19 migration scheduled this year, for the property graph support, the new partition-management commands, or the temporal range features specifically, needs to replan around PostgreSQL 18, since none of those capabilities will ship until a still-undetermined future release. Teams that were simply planning to take the next major version for its usual mix of performance and replication improvements are less affected, since Beta 4 continues regardless and most of PostgreSQL 19's non-reverted work is intact.&lt;/p&gt;

&lt;p&gt;Anyone maintaining a PostgreSQL fork, extension, or tooling that depended on the reverted features' in-progress APIs, including early adopters who built against beta behavior for GROUP BY ALL or the partition commands, should expect that code to need rework regardless of when 19 eventually ships, since the reverted implementations are not simply delayed but pulled pending redesign.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;Wait. PostgreSQL 18, patched as recently as August 2026, remains the version to deploy for anything going to production this year, nothing about the PostgreSQL 19 delay makes 18 less stable, and the review process that caused the delay is, if anything, a point in the project's favor rather than against it.&lt;/p&gt;

&lt;p&gt;For teams specifically waiting on SQL/PGQ property graphs, the new partition-merge commands, or temporal range support, there is no substitute for those features on 18, and no committed date for when 19 will actually ship them, budget for months, not weeks, and don't schedule dependent work against a hard date.&lt;/p&gt;

&lt;p&gt;Everyone else evaluating a PostgreSQL major-version upgrade should treat this delay as routine project hygiene rather than a red flag: a database project finding and fixing correctness bugs before a stable release, rather than after, is the outcome to want.&lt;/p&gt;

&lt;p&gt;Tracked daily from official release feeds and vendor changelogs. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>postgres</category>
      <category>database</category>
      <category>sql</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Qwen3.8-27B Packs Native Vision-Language Understanding Into a Dense 27B Model</title>
      <dc:creator>soy</dc:creator>
      <pubDate>Wed, 16 Sep 2026 09:15:07 +0000</pubDate>
      <link>https://dev.to/soytuber/qwen38-27b-packs-native-vision-language-understanding-into-a-dense-27b-model-3a98</link>
      <guid>https://dev.to/soytuber/qwen38-27b-packs-native-vision-language-understanding-into-a-dense-27b-model-3a98</guid>
      <description>&lt;p&gt;Qwen released Qwen3.8-27B, a dense 27-billion-parameter vision-language model with native image and video understanding and a context window extending to 1,000,000 tokens. It beats its own predecessor by wide margins on agentic coding benchmarks and even edges out the larger Qwen3.7-Plus on several of them, arriving with same-day GGUF quantizations from Unsloth and ISTA-DASLab already sitting among Hugging Face's most-downloaded repositories.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;Qwen3.8-27B is built on the architectural foundation of Qwen3.5 but stays dense: every one of its 27B parameters activates on every token, unlike the family's larger mixture-of-experts variants. The model stacks 16 blocks, each containing three Gated DeltaNet layers followed by one Gated Attention layer, with every layer carrying its own feed-forward network sized at an intermediate dimension of 17,408. Gated DeltaNet uses 48 attention heads for V and 16 for QK at a head dimension of 128; Gated Attention uses 24 query heads and 4 key/value heads at a head dimension of 256, with rotary position embeddings at dimension 64. The model trains with multi-token prediction (MTP) across multiple steps, and context extends from a native 262,144 tokens up to 1,000,000.&lt;/p&gt;

&lt;p&gt;On Qwen's own benchmark suite, the gains over Qwen3.6-27B are consistent across agentic tasks: 61.7 vs. 53.5 on SWE-bench Pro, 79.0 vs. 49.3 on QwenSWEBench, 42.2 vs. 13.3 on DeepSWE 1.1, and 70.7 vs. 61.0 on CoWorkBench, a long-horizon office-work benchmark. On several of these same benchmarks, Qwen3.8-27B also beats the larger Qwen3.7-Plus — 61.7 vs. 57.6 on SWE-bench Pro, 79.0 vs. 59.2 on QwenSWEBench — though it trails Anthropic's Opus 4.6 Max on raw terminal-coding performance, 73.0 vs. 78.2 on Terminal Bench 2.1.&lt;/p&gt;

&lt;p&gt;The model ships under Apache 2.0 and works out of the box with Transformers, vLLM, SGLang, and TokenSpeed. Same-day releases from Unsloth (standard GGUF) and ISTA-DASLab (GSQ-RCO mixed-precision GGUF) put quantized versions on Hugging Face's trending list within hours, alongside a range of community fine-tunes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this affects
&lt;/h2&gt;

&lt;p&gt;Teams currently running Qwen3.6-27B, or a comparably sized open-weight vision-language model, get a straightforward upgrade path: same parameter class, better agentic-coding and long-horizon task scores, and a longer context window. Anyone building agents that need to reason over screenshots, diagrams, or video alongside text — rather than bolting a separate vision model onto a text-only LLM — gains a single dense model that handles both natively.&lt;/p&gt;

&lt;p&gt;It matters less for teams already committed to a mixture-of-experts deployment for throughput reasons, since Qwen3.8-27B's dense design trades that efficiency profile away, and less for teams whose workloads are purely text with no multimodal requirement, where a non-vision model in the same parameter range may run cheaper. Terminal-heavy coding workflows that need the absolute ceiling on agentic coding benchmarks should also weigh Opus 4.6 Max's lead on Terminal Bench 2.1 before switching.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;Upgrade if the current stack is Qwen3.6-27B or a similarly sized open-weight VLM: the benchmark gains are broad rather than cherry-picked to one category, the license and framework compatibility carry over cleanly, and quantized GGUF builds are already available for consumer-GPU deployment. Teams without an existing Qwen commitment should still evaluate it against whatever dense VLM they're running today, since beating a larger stablemate (Qwen3.7-Plus) on several benchmarks is a stronger signal than typical same-family point releases.&lt;/p&gt;

&lt;p&gt;Hold off if the workload is purely text, MoE-based deployment already meets throughput targets, or the ceiling on terminal-coding tasks specifically matters more than the rest of the benchmark spread — Opus 4.6 Max still leads there. There's no reason to wait for a more mature release: this is a full release with published weights, benchmarks, and day-one quantizations, not a preview.&lt;/p&gt;

&lt;p&gt;Tracked daily from official release feeds and vendor changelogs. Full archive: &lt;a href="https://media.patentllm.org" rel="noopener noreferrer"&gt;https://media.patentllm.org&lt;/a&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>selfhosted</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
