<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Qasim Parray</title>
    <description>The latest articles on DEV Community by Qasim Parray (@abyzgenic).</description>
    <link>https://dev.to/abyzgenic</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4120845%2Fdf746bdb-4b77-4db2-8b04-ed1d2da71842.png</url>
      <title>DEV Community: Qasim Parray</title>
      <link>https://dev.to/abyzgenic</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/abyzgenic"/>
    <language>en</language>
    <item>
      <title>Core Web Vitals: What Cloudflare's BEACON Data Says to Fix First</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Fri, 02 Oct 2026 16:38:14 +0000</pubDate>
      <link>https://dev.to/abyzgenic/core-web-vitals-what-cloudflares-beacon-data-says-to-fix-first-5gi0</link>
      <guid>https://dev.to/abyzgenic/core-web-vitals-what-cloudflares-beacon-data-says-to-fix-first-5gi0</guid>
      <description>&lt;p&gt;I have a bad habit with performance work. I open Lighthouse, see a red number, and start fixing whatever the report lists first. Usually that's an image that's 40KB too big. It feels productive, and it moves the score, and half the time it has nothing to do with what my client's users actually wait for.&lt;/p&gt;

&lt;p&gt;So when Cloudflare released a dataset of real-user measurements from about 10,000 of the biggest sites on its network, I read the &lt;a href="https://blog.cloudflare.com/how-fast-is-the-web/" rel="noopener noreferrer"&gt;BEACON announcement&lt;/a&gt; with a specific question: if I had one afternoon to improve Core Web Vitals on a normal client site, where should it go? Short version: fix the server response and the render delay before you touch a single image, and stop testing only on Chrome.&lt;/p&gt;

&lt;p&gt;I haven't run BEACON's BigQuery queries against my own client work yet, so treat what follows as a reading of Cloudflare's published numbers plus my opinions about them, not as a lab report.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the dataset actually is
&lt;/h2&gt;

&lt;p&gt;BEACON is anonymized real-user monitoring data, refreshed daily in Google BigQuery. Cloudflare says it covers roughly 10,000 large sites and every major browser engine. Domain names and URL paths are stripped, and any group with fewer than five data points is dropped. Records are published as histograms, not averages, so you can compute your own percentiles.&lt;/p&gt;

&lt;p&gt;That last detail matters more than it sounds. Most "average page speed" charts you see online are averages of averages. A histogram lets you ask the question Google itself asks: what happens at the 75th percentile? The &lt;a href="https://web.dev/articles/vitals" rel="noopener noreferrer"&gt;official thresholds on web.dev&lt;/a&gt; are LCP within 2.5 seconds, INP at 200 milliseconds or less, and CLS at 0.1 or less, all measured at the 75th percentile of page loads.&lt;/p&gt;

&lt;p&gt;One caveat before I lean on any of this. Ten thousand big sites on one CDN is not the web. A five-page brochure site on shared hosting has different problems from a site big enough to be in this sample. I'll come back to that at the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where LCP time actually goes
&lt;/h2&gt;

&lt;p&gt;Largest Contentful Paint gets split into four parts: time to first byte, load delay, load duration, and render delay. Cloudflare publishes the breakdown for pages that land in the good, needs-improvement and poor buckets, and the shape is what interested me.&lt;/p&gt;

&lt;p&gt;As I read the table, pages with good LCP have a document TTFB around 598ms, while poor ones sit near 1,891ms. Load delay goes from 76ms to about 1,485ms. Render delay goes from 157ms to about 2,002ms. Load duration, the time actually downloading the LCP resource, barely moves between buckets.&lt;/p&gt;

&lt;p&gt;Read that again, because it is the opposite of what most tutorials teach. The advice everywhere is "compress your hero image". But download time is the part that hardly differs between fast and slow pages. What differs is everything around the download: how late the server answers, how late the browser discovers the resource, and how long it sits idle before painting.&lt;/p&gt;

&lt;p&gt;So here's the order I'm going to work in from now on:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Check TTFB first. If the HTML takes more than a second to arrive, nothing else you do will get you under 2.5 seconds on a mid-range phone.&lt;/li&gt;
&lt;li&gt;Check load delay. This is the gap between the HTML arriving and the browser starting to fetch the LCP image or font.&lt;/li&gt;
&lt;li&gt;Check render delay. Usually a blocking script or a client-side framework that hasn't hydrated.&lt;/li&gt;
&lt;li&gt;Only then look at file size.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Measure the four parts on your own site
&lt;/h2&gt;

&lt;p&gt;You don't need BigQuery to find out which part hurts. The &lt;a href="https://github.com/GoogleChrome/web-vitals" rel="noopener noreferrer"&gt;web-vitals library&lt;/a&gt; has an attribution build that hands you the same four numbers per page view.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;onLCP&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;web-vitals/attribution&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nf"&gt;onLCP&lt;/span&gt;&lt;span class="p"&gt;(({&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;attribution&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nb"&gt;navigator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendBeacon&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/vitals&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;metric&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;LCP&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;total&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;ttfb&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;attribution&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;timeToFirstByte&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;loadDelay&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;attribution&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;resourceLoadDelay&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;loadDuration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;attribution&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;resourceLoadDuration&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;renderDelay&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;attribution&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;elementRenderDelay&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;element&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;attribution&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;}));&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Collect a week of that, sort by the 75th percentile of each field, and the biggest one is your afternoon. A very common cause of a fat load delay is a boring one: the hero image is set as a CSS background, so the browser can't see it until the stylesheet has been parsed. Moving it to a real image tag with a priority hint fixes most of it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="c"&gt;&amp;lt;!-- before: browser discovers this late, inside the CSS --&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;div&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"hero"&lt;/span&gt; &lt;span class="na"&gt;style=&lt;/span&gt;&lt;span class="s"&gt;"background-image:url(/hero.webp)"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&amp;lt;/div&amp;gt;&lt;/span&gt;

&lt;span class="c"&gt;&amp;lt;!-- after: discoverable in the HTML, fetched early --&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;img&lt;/span&gt; &lt;span class="na"&gt;src=&lt;/span&gt;&lt;span class="s"&gt;"/hero.webp"&lt;/span&gt; &lt;span class="na"&gt;width=&lt;/span&gt;&lt;span class="s"&gt;"1200"&lt;/span&gt; &lt;span class="na"&gt;height=&lt;/span&gt;&lt;span class="s"&gt;"600"&lt;/span&gt;
     &lt;span class="na"&gt;fetchpriority=&lt;/span&gt;&lt;span class="s"&gt;"high"&lt;/span&gt; &lt;span class="na"&gt;alt=&lt;/span&gt;&lt;span class="s"&gt;"Product dashboard on a laptop"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a two-line change and it attacks the load delay bucket, not the file size bucket. The width and height attributes also cover CLS, so you get a second win for free.&lt;/p&gt;

&lt;h2&gt;
  
  
  The browser engine gap nobody tests
&lt;/h2&gt;

&lt;p&gt;This was the finding that made me put my coffee down. Cloudflare reports that WebKit, the engine under every iPhone browser, is the best performer overall. Fine. But in 46 countries, together over 10% of traffic, WebKit's LCP or INP trails Blink-based browsers by at least 10%. Their example is Cambodia, where WebKit's LCP is 50% worse than Blink's while making up 17.5% of page views.&lt;/p&gt;

&lt;p&gt;It's a trap I fall into easily. Testing on desktop Chrome, maybe throttled, means looking at one engine, in one country, on good bandwidth. If your clients sell to people in Southeast Asia, Africa or South America, the browser you never open may be the one a fifth of the customers use.&lt;/p&gt;

&lt;p&gt;The practical fix is dull. Put the field data you collect (like the beacon above) behind a breakdown by browser engine and country, and look at the tail instead of the median. If you use a hosted RUM tool, check whether it lets you segment that way. If it doesn't, that's a real reason to switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Soft navigations change what "fast" means for SPAs
&lt;/h2&gt;

&lt;p&gt;The other table in the post compares hard navigations (a full page load) with soft navigations (client-side route changes in a single-page app). At the 75th percentile, hard navigations come in at 1,421ms and soft ones at 582ms. At the median it's 791ms against 274ms. Soft navigations render two to three times faster at every percentile Cloudflare lists.&lt;/p&gt;

&lt;p&gt;Before anyone tweets "SPAs are faster", look at the second half of that finding: the landing page in those apps is still heavy, with a median around 1,370ms. You pay the cost once, at the door, and then every click is cheap.&lt;/p&gt;

&lt;p&gt;That reframes a decision I get asked about a lot. If your users typically land on one page and leave, a heavy client-side framework buys you nothing and costs you the entry. If they land and then click through ten screens, the trade flips. I wrote about a related piece of this in my &lt;a href="https://abrarqasim.com/blog/react-19-3-view-transitions-stable-the-timeout-hack-i-deleted" rel="noopener noreferrer"&gt;post on React 19.3 view transitions&lt;/a&gt;, where the transition itself was the thing I'd been faking with timeouts. And the entry-cost side of the story showed up in my notes on &lt;a href="https://abrarqasim.com/blog/turbopack-chunking-nextjs-16-3-the-two-thirds-guess-i-never-questioned" rel="noopener noreferrer"&gt;Turbopack chunking in Next.js 16.3&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Chrome measures soft navigations experimentally at the moment, so I'd treat these numbers as a preview of where measurement is going, not a metric you can put in a client report today.&lt;/p&gt;

&lt;h2&gt;
  
  
  INP: where the 200ms goes
&lt;/h2&gt;

&lt;p&gt;For INP, Cloudflare splits the interaction into input delay, processing time and presentation delay. The good-bucket values are about 18ms, 55ms and 56ms. The poor bucket has 84ms of input delay, 284ms of processing and 217ms of presentation delay.&lt;/p&gt;

&lt;p&gt;Processing time is where your own code lives, and it's the bucket that grows fastest. The usual cause is one click handler doing too much before the browser gets to paint. The smallest fix I know is to yield to the main thread after the visible work is done.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;button&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;addEventListener&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;click&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;showSpinner&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;                       &lt;span class="c1"&gt;// the thing the user must see now&lt;/span&gt;

  &lt;span class="c1"&gt;// let the browser paint before the heavy part runs&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;scheduler&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="nb"&gt;window&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;yield&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;scheduler&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;scheduler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;yield&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nf"&gt;runExpensiveFilter&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;                &lt;span class="c1"&gt;// the slow part&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;scheduler.yield()&lt;/code&gt; is not in every browser yet, hence the fallback. It won't make the expensive function faster. It makes the interaction feel faster, which is what INP is measuring in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do this week
&lt;/h2&gt;

&lt;p&gt;Ten thousand large sites is the right sample to learn from and the wrong one to copy blindly. Use it to decide what to measure, then measure your own.&lt;/p&gt;

&lt;p&gt;Here's the afternoon plan. Add the attribution snippet above to one production site and let it run for a few days. Look at the 75th percentile of TTFB, load delay and render delay, and fix the biggest one. Then split the same data by browser engine and check whether Safari on iPhone is worse than Chrome for your audience. If you'd like to see the sort of client work where I do this kind of audit, my &lt;a href="https://abrarqasim.com" rel="noopener noreferrer"&gt;portfolio&lt;/a&gt; has examples. And if you're curious about the raw data, open BEACON in BigQuery and run one query against the LCP histograms for a browser and country you care about. It takes about ten minutes and it will change what you look at first.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/how-to-improve-core-web-vitals-cloudflare-beacon-what-to-fix-first/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>corewebvitals</category>
      <category>webperf</category>
      <category>cloudflare</category>
      <category>lcp</category>
    </item>
    <item>
      <title>Vite+ 1.0 and Oxlint: The Speedup Claims I'll Test Before Switching</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Thu, 01 Oct 2026 16:38:10 +0000</pubDate>
      <link>https://dev.to/abyzgenic/vite-10-and-oxlint-the-speedup-claims-ill-test-before-switching-3jj9</link>
      <guid>https://dev.to/abyzgenic/vite-10-and-oxlint-the-speedup-claims-ill-test-before-switching-3jj9</guid>
      <description>&lt;p&gt;Confession: I have been putting off a lint cleanup for four months. Not because it's hard, but because every time I run the type-aware ESLint pass on a mid-sized TypeScript project, I go make coffee, come back, and it's still going. So when Evan You published a &lt;a href="https://blog.cloudflare.com/voidzero-update/" rel="noopener noreferrer"&gt;four-month progress report on VoidZero at Cloudflare&lt;/a&gt; with the line "up to 18x faster than ESLint in large codebases", my first reaction was suspicion, and my second was to open a terminal.&lt;/p&gt;

&lt;p&gt;Short version for the impatient: the numbers are vendor numbers, the direction is believable, and Vite+ 1.0 is the first time the whole Rust-based JavaScript toolchain comes in one box. I haven't measured any of it on my own repos yet, so this post is a reading of the announcement plus the exact checks I'm going to run before I let any of it near a client project. If you want the checks without the commentary, skip to the last section.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually shipped in four months
&lt;/h2&gt;

&lt;p&gt;The post lists five things, and they are worth separating because they are different kinds of claims. The Oxc React Compiler compiles React apps 10x faster than the Babel version. Vitest 5 is up to 50% faster than Vitest 4. tsgolint, the type-aware engine behind Oxlint, is stable and runs 12 to 18 times faster than ESLint with typescript-eslint. Oxfmt is 7x faster than Prettier now that its JSON, CSS, SCSS, Less, GraphQL and YAML formatters are written in Rust. And Vite+ hit 1.0.&lt;/p&gt;

&lt;p&gt;Notice the qualifiers. "Up to 50%" is a best case. "Up to 18x" is a best case on large codebases. "10x" for the React Compiler is a compile-step number, and compile time is rarely the slowest part of your build. Every one of these is a real speedup on some project, and none of them tells me what happens on mine. That's not a knock on VoidZero, every vendor writes this way. It's a reminder that a benchmark from the people who wrote the tool answers "can this be fast?" and not "will this be fast for me?".&lt;/p&gt;

&lt;p&gt;The other detail I liked is the stack argument. Oxc is the compiler, Rolldown is the bundler, Vite sits on top, Oxlint and Vitest use the same parts. When one layer gets faster, the layers above get faster for free. The logic holds up on paper, since a shared parser and shared AST mean one fix pays off in several tools, so I'm inclined to believe it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I care about: type-aware linting
&lt;/h2&gt;

&lt;p&gt;Linting with type information is the slow one, and it's the one that catches real bugs. Floating promises, misused async callbacks, unsafe &lt;code&gt;any&lt;/code&gt; flowing into a function. These need the TypeScript program in memory, and that's where ESLint plus typescript-eslint spends its time.&lt;/p&gt;

&lt;p&gt;According to the announcement, tsgolint supports 59 of the 61 typescript-eslint type-aware rules, and Oxlint can share a single TypeScript program between linting and type-checking instead of building it twice. That second part is the one I'd look at first, because most CI pipelines I've seen run &lt;code&gt;tsc --noEmit&lt;/code&gt; and ESLint as two separate steps, and both of them pay to parse the whole project.&lt;/p&gt;

&lt;p&gt;The config change they show is small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;defineConfig&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;oxlint&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nf"&gt;defineConfig&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;options&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;typeAware&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;typeCheck&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compare that with what I have today, which is an &lt;code&gt;eslint.config.js&lt;/code&gt; with a parser option pointing at a tsconfig, a project service setting I had to look up twice, and a comment explaining why the slow rules are slow. Fewer moving parts is a win even if the speedup turns out to be 6x instead of 18x.&lt;/p&gt;

&lt;p&gt;The two missing rules matter, though. If your team leans on one of the two typescript-eslint rules tsgolint doesn't cover, you'll be running both linters for a while. I don't know which two they are, and the announcement doesn't say, so that's the first thing I'd check in the docs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Oxfmt and the Prettier question
&lt;/h2&gt;

&lt;p&gt;Formatters are a religious topic, so let me keep this practical. Oxfmt claims Prettier-compatible output, and the migration is two commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pnpm add &lt;span class="nt"&gt;-D&lt;/span&gt; oxfmt
pnpm oxfmt &lt;span class="nt"&gt;--migrate&lt;/span&gt; prettier
pnpm oxfmt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The risk with any Prettier replacement is the diff. If the output differs on even 0.5% of lines, you get a giant formatting commit that wrecks &lt;code&gt;git blame&lt;/code&gt;, and I've already lived through that once with a different tool (I wrote about the &lt;a href="https://abrarqasim.com/blog/laravel-pint-the-one-commit-that-broke-my-git-blame/" rel="noopener noreferrer"&gt;one commit that broke my git blame&lt;/a&gt; with Laravel Pint). So the plan is boring: run Oxfmt on a branch, count changed lines, and add the formatting commit to &lt;code&gt;.git-blame-ignore-revs&lt;/code&gt; if I keep it.&lt;/p&gt;

&lt;p&gt;Speed on a formatter is also the least interesting win for me. On a small project, formatting time is already a few seconds. I'd switch for the single toolchain, not for the seconds saved.&lt;/p&gt;

&lt;h2&gt;
  
  
  Vite+ 1.0 and the decision fatigue argument
&lt;/h2&gt;

&lt;p&gt;Vite+ bundles Vite 8, Vitest 5, Rolldown, Oxlint, Oxfmt and task caching with defaults chosen for you. Evan You's pitch is that agents and humans both waste time on "which linter shall I use?" and that fewer choices means faster shipping. I half agree.&lt;/p&gt;

&lt;p&gt;The half I agree with: a new project shouldn't need six config files before the first commit. The half I don't: defaults are great until the day you need to deviate, and then you're reading a unified tool's docs instead of five well-known ones. I have client projects pinned to specific ESLint plugins for framework reasons, and a "great defaults" toolchain doesn't help those. I'd use Vite+ for new projects and leave the old ones alone until there's a reason.&lt;/p&gt;

&lt;p&gt;There's also a framing I'm skeptical of. The post says that in 2026 the mission is making developers' agents faster, because once inference is quick, a slow linter or type-checker becomes the bottleneck. That's true as far as it goes. But if your agent loop spends 40 seconds waiting on a build, the fix is often to run less, not to run the same thing faster. Scoping a lint run to the changed files beats a 10x faster full lint on every iteration. Faster tools and smarter invocations stack, so do both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bundled dev and the React Compiler
&lt;/h2&gt;

&lt;p&gt;Two more items deserve a mention. The first is Bundled Dev, formerly Full Bundle Mode, which uses Vite's production bundler in development. It's behind an experimental flag today, and the post says Cloudflare's own dashboard already uses it for all internal developers. That's a useful data point, since a big dashboard app is the kind of project where unbundled dev servers struggle with thousands of module requests.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;defineConfig&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;vite&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nf"&gt;defineConfig&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;experimental&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;bundledDev&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second is the Oxc React Compiler, enabled through &lt;code&gt;@vitejs/plugin-react&lt;/code&gt; with &lt;code&gt;compiler: true&lt;/code&gt; plus the &lt;code&gt;oxc-transform-react&lt;/code&gt; package. I've written about how &lt;a href="https://abrarqasim.com/blog/react-19-3-view-transitions-stable-the-timeout-hack-i-deleted/" rel="noopener noreferrer"&gt;React 19.3 view transitions&lt;/a&gt; let me delete some hand-rolled timing code, and the compiler is the same category of change: less manual work in components. My hesitation is that a compiler is a correctness risk in a way a linter isn't. A slower linter wastes time. A wrong compile output ships a bug. I'd want a full test run and a visual check before trusting it on anything with money attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'll run before switching
&lt;/h2&gt;

&lt;p&gt;Here is the checklist, in the order I'd do it on a real project. It takes about an hour, and it's the honest way to test a vendor claim.&lt;/p&gt;

&lt;p&gt;First, time your current baseline. Run your existing lint, type-check and test commands three times each and write down the median. Without a baseline, "18x faster" is a number you can't verify.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;i &lt;span class="k"&gt;in &lt;/span&gt;1 2 3&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt; /usr/bin/time &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s2"&gt;"%e s"&lt;/span&gt; pnpm eslint &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done
for &lt;/span&gt;i &lt;span class="k"&gt;in &lt;/span&gt;1 2 3&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt; /usr/bin/time &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s2"&gt;"%e s"&lt;/span&gt; pnpm tsc &lt;span class="nt"&gt;--noEmit&lt;/span&gt; &lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done
for &lt;/span&gt;i &lt;span class="k"&gt;in &lt;/span&gt;1 2 3&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt; /usr/bin/time &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s2"&gt;"%e s"&lt;/span&gt; pnpm vitest run &lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Second, do the same with the new tools on a branch, with warm and cold caches. Cold-cache numbers are what CI sees, and warm-cache numbers are what you see locally. Third, diff the outputs: lint findings from both linters on the same commit, and formatter output on the same files. A faster tool that misses a class of bug is not a win. Fourth, check the two type-aware rules that tsgolint doesn't cover against your config.&lt;/p&gt;

&lt;p&gt;If you want to try it this week, pick your slowest repo, run the baseline loop above today, and post the three medians somewhere you'll find them again. Then install Oxlint on a branch and compare. I'll be doing the same on my own projects, and I'd rather you tell me the announcement's numbers didn't hold up for you than take my word or theirs. You can see the kind of projects I'd test this on in my &lt;a href="https://abrarqasim.com" rel="noopener noreferrer"&gt;portfolio&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Sources for the numbers above are in the &lt;a href="https://blog.cloudflare.com/voidzero-update/" rel="noopener noreferrer"&gt;VoidZero announcement on the Cloudflare blog&lt;/a&gt;, and the tools themselves live at &lt;a href="https://vite.dev/" rel="noopener noreferrer"&gt;Vite&lt;/a&gt; and &lt;a href="https://oxc.rs/" rel="noopener noreferrer"&gt;Oxc&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/vite-plus-1-0-oxlint-tsgolint-the-benchmarks-i-will-run-before-switching/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>vite</category>
      <category>oxlint</category>
      <category>typescript</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Laravel Wayfinder: Typed Routes and the Deploy Bug I Hit</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Thu, 01 Oct 2026 08:38:10 +0000</pubDate>
      <link>https://dev.to/abyzgenic/laravel-wayfinder-typed-routes-and-the-deploy-bug-i-hit-3agl</link>
      <guid>https://dev.to/abyzgenic/laravel-wayfinder-typed-routes-and-the-deploy-bug-i-hit-3agl</guid>
      <description>&lt;p&gt;Three weeks ago a teammate renamed a route in a Laravel app we were shipping together, &lt;code&gt;posts.show&lt;/code&gt; became &lt;code&gt;posts.view&lt;/code&gt; because he hated the old name, and nothing in PHP complained. The break showed up on the frontend instead, quietly, because somebody (me, a few months earlier) had hardcoded &lt;code&gt;/posts/${id}&lt;/code&gt; directly into a fetch call two files away from any route definition. Nobody noticed until a client clicked a broken link in production. That is the exact category of bug Laravel Wayfinder exists to make impossible, and after running it on a real Inertia project for a few weeks, it mostly delivers. Mostly, because it also handed me a new way to break a deploy that I had never hit before.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Wayfinder actually generates
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/laravel/wayfinder" rel="noopener noreferrer"&gt;Wayfinder&lt;/a&gt; is a Laravel package that reads your registered routes and controllers, then generates fully typed TypeScript functions for each one. Install it with Composer and pair it with the &lt;a href="https://github.com/laravel/vite-plugin-wayfinder" rel="noopener noreferrer"&gt;official Vite plugin&lt;/a&gt; so it regenerates automatically as you work:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;composer require laravel/wayfinder
npm i &lt;span class="nt"&gt;-D&lt;/span&gt; @laravel/vite-plugin-wayfinder
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add the plugin to &lt;code&gt;vite.config.js&lt;/code&gt;, and every time you touch a controller or a route file, a fresh set of TypeScript definitions lands in &lt;code&gt;resources/js/actions&lt;/code&gt; and &lt;code&gt;resources/js/routes&lt;/code&gt;. You can also run it by hand:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;php artisan wayfinder:generate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That command is doing something conceptually simple but tedious to write by hand: it walks the router, matches parameter bindings, resolves which controller method actually owns each route, and produces a function that knows its own URL and the shape of the arguments it expects. The HTTP method comes along for free too, which matters more than it sounds like once you're calling &lt;code&gt;destroy()&lt;/code&gt; and &lt;code&gt;show()&lt;/code&gt; from the same file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Swapping hardcoded URLs for typed functions
&lt;/h2&gt;

&lt;p&gt;Here's the pattern I was using before, which is probably familiar if you've built an Inertia or React frontend against a Laravel API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// before: a string someone typed by hand, nowhere near the route definition&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;loadPost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`/posts/&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing here checks that &lt;code&gt;/posts/${id}&lt;/code&gt; still matches what's in &lt;code&gt;routes/web.php&lt;/code&gt;. Rename the route, or swap the parameter binding from an id to a slug, and this line keeps compiling right up until it 404s in front of a user. With Wayfinder generating functions from the same PostController the route actually points to, the call becomes this instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;show&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@/actions/App/Http/Controllers/PostController&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;loadPost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;show&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the route changes, &lt;code&gt;wayfinder:generate&lt;/code&gt; regenerates &lt;code&gt;show&lt;/code&gt;, and my code either keeps compiling because nothing meaningful changed, or TypeScript flags the call site because the parameters shifted. That second case is the entire point: a rename that used to fail silently in the browser now fails loudly at build time, on my machine, before anyone else sees it.&lt;/p&gt;

&lt;p&gt;The part I didn't expect to like as much is how it handles routes that share a controller method. If two named routes point at the same action, Wayfinder can't guess which URL you mean, so it hands back a dictionary keyed by the route instead of a plain function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;index&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@/actions/App/Http/Controllers/ClientPaymentsController&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// clients.payments.index and clients.payments.archive both hit ClientPaymentsController@index&lt;/span&gt;
&lt;span class="nx"&gt;index&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/clients/{client}/payments&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]({&lt;/span&gt; &lt;span class="na"&gt;client&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first time I saw that shape in an editor autocomplete, I assumed it was a bug in the generator. It isn't. It's Wayfinder refusing to silently guess which of two ambiguous routes you meant, which is a more honest failure mode than most route helpers I've used.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this compares to what I was doing before
&lt;/h2&gt;

&lt;p&gt;Before Wayfinder, my go-to for this problem was Ziggy, which exposes a &lt;code&gt;route()&lt;/code&gt; helper on the frontend that mirrors Laravel's own named-route syntax. Ziggy solves a chunk of the same problem: you stop hardcoding URLs, and a route rename gets picked up automatically. What it doesn't give you is typed arguments. Calling &lt;code&gt;route('posts.show', { id: 'not-a-number' })&lt;/code&gt; compiles fine and fails at runtime, because Ziggy's helper takes a loosely typed object and doesn't know or care what the controller expects. Wayfinder's generated functions come from the same source of truth, the router, but because they're generated per-controller-method rather than interpreted at runtime, TypeScript actually understands the parameter shape. I didn't rip Ziggy out of the project; the two coexist fine, and there are still a few places where Ziggy's simpler &lt;code&gt;route()&lt;/code&gt; call is genuinely less code for a one-off link. But every new fetch call I write goes through Wayfinder now, because the difference between "this compiles" and "this actually matches the controller" is exactly the gap that caused the bug I opened this post with.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotcha nobody mentions until you hit it
&lt;/h2&gt;

&lt;p&gt;The bug I actually lost an afternoon to wasn't in the generated code at all. It was in the deploy pipeline. Wayfinder reads routes from the application's live router at generation time, and our deploy script ran &lt;code&gt;php artisan optimize&lt;/code&gt;, which caches the route table, before the frontend build step. On the release where we added a new controller action, the cached route table from the previous release was still in memory when &lt;code&gt;wayfinder:generate&lt;/code&gt; ran during &lt;code&gt;npm run build&lt;/code&gt;. The new action was invisible to it. Vite's build failed with an error about a module that "didn't exist," pointing at a &lt;code&gt;resources/js/actions&lt;/code&gt; file that had simply never been generated, because the router it asked didn't know the route existed yet.&lt;/p&gt;

&lt;p&gt;The fix, once I found where to look, is one line in the deploy script, run before the frontend build and after the code is in place:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;php artisan route:clear
npm run build
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Clearing the route cache before regenerating TypeScript definitions and only re-caching it afterward closed the gap. I want to be annoyed that this isn't the default behavior, but I also understand why it isn't: route caching and TypeScript generation are two features written by different people for different reasons, and the order they run in is a deploy-script decision, not a package decision. It just means you now own one more ordering constraint in a script that was already fragile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it still falls short
&lt;/h2&gt;

&lt;p&gt;Wayfinder only knows about routes that exist when you run the generator, so anything built dynamically at runtime, a route registered from a database-driven plugin system, for instance, won't show up no matter how carefully you order your deploy steps. I also had to opt in explicitly to generate the &lt;code&gt;.form&lt;/code&gt; variants for plain HTML form submissions, which cost me twenty minutes of confusion before I found the flag:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;php artisan wayfinder:generate &lt;span class="nt"&gt;--with-form&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the package itself is still labeled beta, with the maintainers warning the API can change before a 1.0 release. That's a reasonable thing to be cautious about if you're gluing generated code into a large codebase you don't want to touch twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do this week
&lt;/h2&gt;

&lt;p&gt;If you're already running Laravel with an Inertia or React frontend and you've ever hardcoded a URL string that later broke silently, install Wayfinder on a single low-risk feature branch. Wire in the Vite plugin, then replace the fetch calls on just one page and see how the generated types feel before you touch anything else. Then check your own deploy script for the same &lt;code&gt;optimize&lt;/code&gt; before &lt;code&gt;build&lt;/code&gt; ordering that bit me. It takes about ten minutes to confirm, and it's a lot cheaper to fix before a release than after one. I write up more of this kind of Laravel tooling friction, alongside the &lt;a href="https://abrarqasim.com/blog/laravel-pint-the-one-commit-that-broke-my-git-blame" rel="noopener noreferrer"&gt;Pint formatting change that broke a client's git blame&lt;/a&gt;, on this blog, and I take on exactly this flavor of frontend-backend integration work through &lt;a href="https://abrarqasim.com/about" rel="noopener noreferrer"&gt;my consulting practice&lt;/a&gt; when teams want a second set of eyes before they ship it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/laravel-wayfinder-typed-routes-the-deploy-bug-i-hit-on-day-one/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>laravel</category>
      <category>typescript</category>
      <category>inertia</category>
      <category>wayfinder</category>
    </item>
    <item>
      <title>Testing the 9x Smaller Local LLM Claim on a GPU-less VPS</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Wed, 30 Sep 2026 16:38:09 +0000</pubDate>
      <link>https://dev.to/abyzgenic/testing-the-9x-smaller-local-llm-claim-on-a-gpu-less-vps-4026</link>
      <guid>https://dev.to/abyzgenic/testing-the-9x-smaller-local-llm-claim-on-a-gpu-less-vps-4026</guid>
      <description>&lt;p&gt;Last month I did something no sane person should try on a Tuesday night: I downloaded a custom fork of llama.cpp from a company I'd never heard of, just to run a 27 billion parameter model that supposedly fits in under 6 gigabytes. I'd read about it in &lt;a href="https://simonwillison.net/2026/Sep/17/hn-49747390/" rel="noopener noreferrer"&gt;Simon Willison's comment on the release&lt;/a&gt;, a ternary-quantized model called Bonsai 2 27B, pitched as "near-lossless compression in a 9x smaller footprint." Near-lossless is a big claim for something throwing away most of the bits in every weight. I wanted to know whether that phrase meant what I hoped, or whether it was the kind of number that only survives contact with a benchmark table. So instead of running it on a beefy Mac like Willison did, I pointed it at the cheapest box I own: the small Hetzner VPS that runs this very blog's publishing pipeline. Here's what actually happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "ternary" buys you and what it costs
&lt;/h2&gt;

&lt;p&gt;Most quantized models you download are 4-bit or 8-bit: each weight still gets a handful of possible values. Ternary quantization is far more aggressive. Every weight collapses to one of three values, roughly minus one, zero, or plus one, and the model leans on scaling factors to recover something close to the original behavior. Prism's Ternary-Bonsai-2-27B-gguf takes a 27 billion parameter model and compresses it down to a single GGUF file just under 6 gigabytes, using a PTQ1_0 format that most mainstream llama.cpp builds can't even read yet. That's why you need &lt;a href="https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15" rel="noopener noreferrer"&gt;Prism's llama.cpp fork&lt;/a&gt; instead of the stock binary: the 1-bit quantization support isn't merged upstream.&lt;/p&gt;

&lt;p&gt;The tradeoff is obvious once you say it out loud. You're asking a model to represent everything it knows using something closer to a light switch than a dial. Whether that "near-lossless" label holds up depends entirely on what you're asking it to do, and I was not about to take a marketing line at face value without running it myself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting it running on a box with no GPU
&lt;/h2&gt;

&lt;p&gt;Willison's setup used the macOS Apple Silicon build and Metal acceleration, and reported 20 to 44 tokens per second depending on the run. I don't have an M-series Mac sitting around for this kind of experiment, and I was more curious about the worst case anyway: what happens on a machine with no GPU at all. Prism ships CPU-only Ubuntu binaries alongside the CUDA and Metal ones, so I grabbed that instead.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fL&lt;/span&gt; https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b10685-7dffb15/llama-prism-b10685-7dffb15-bin-ubuntu-x64.tar.gz &lt;span class="nt"&gt;-o&lt;/span&gt; bonsai-runtime.tar.gz
&lt;span class="nb"&gt;tar&lt;/span&gt; &lt;span class="nt"&gt;-xzf&lt;/span&gt; bonsai-runtime.tar.gz

curl &lt;span class="nt"&gt;-fL&lt;/span&gt; https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/main/Ternary-Bonsai-2-27B-PTQ1_0.gguf &lt;span class="nt"&gt;-o&lt;/span&gt; bonsai.gguf

./llama-prism-b10685-7dffb15/llama-server &lt;span class="nt"&gt;-m&lt;/span&gt; bonsai.gguf &lt;span class="nt"&gt;--port&lt;/span&gt; 8331 &lt;span class="nt"&gt;-c&lt;/span&gt; 8192
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The download itself was the fast part, five and a half minutes on the VPS's connection for a model that would normally arrive as forty-plus gigabytes uncompressed. The server started fine and the built-in web UI came up at localhost, same as Willison described. Where it stopped being fun was throughput: without a GPU to lean on, I was getting something in the neighborhood of two to three tokens a second on simple prompts. That's the kind of speed where you start a request, go refill your coffee, and come back to find it's maybe a third done. Usable for a background batch job. Not usable for anything where a person is sitting there waiting.&lt;/p&gt;

&lt;p&gt;I also hit a smaller version of the same oddity Willison mentioned: restarting the server sometimes changed the throughput noticeably, once by close to 40 percent, with no configuration change on my end. I don't have a clean explanation for that, and neither did the release notes.&lt;/p&gt;

&lt;p&gt;Memory turned out to be the easier problem to reason about. The GGUF itself sits under 6 gigabytes, but once you add the context window and the server's own overhead, plan on closer to 8 or 9 gigabytes of RAM before you touch a single prompt. My VPS has 16, so there was headroom, but I've seen enough OOM-killed processes on smaller boxes to say: check &lt;code&gt;free -h&lt;/code&gt; before you start the server, not after it silently dies mid-download.&lt;/p&gt;

&lt;h2&gt;
  
  
  So is it actually near lossless
&lt;/h2&gt;

&lt;p&gt;Here's where I have to be honest about the limits of a one-evening test. I did not run a proper benchmark suite against the full-precision version of this model. What I did was throw a handful of the kinds of questions I actually ask models day to day: summarizing a paragraph of documentation, writing a short regex, explaining a stack trace, renaming a batch of variables to something less embarrassing. On those, the quantized model held up better than I expected for something described in bits rather than bytes. It didn't fall apart or produce garbage. But "didn't fall apart on five casual prompts" and "near-lossless" are different claims, and I'm not going to pretend I verified the second one just because the first one felt true.&lt;/p&gt;

&lt;p&gt;What convinces me the underlying technique deserves attention, rather than dismissal, is the pace of the surrounding field this year. Willison's own year-in-review post on 2026's LLM developments makes the case that most of the meaningful progress hasn't been raw capability, it's been models crossing a threshold from "often makes mistakes" to "reliable enough to use daily" at a given size and cost. A 9x compression ratio that gets even 80 percent of the way to lossless changes which threshold a given piece of hardware can clear. That's a more interesting question than whether any single benchmark number is exactly right.&lt;/p&gt;

&lt;h2&gt;
  
  
  When self-hosting something like this actually makes sense
&lt;/h2&gt;

&lt;p&gt;I already wrote about &lt;a href="https://abrarqasim.com/blog/hetzner-vs-digitalocean-what-this-blog-actually-runs-on" rel="noopener noreferrer"&gt;what this blog's own infrastructure costs to run&lt;/a&gt;, and the honest answer is that a hosted API call is almost always cheaper than the electricity and hardware amortization of running your own inference, unless you're doing enough volume to change that math or you have a hard requirement that the data never leaves your machine. Ternary quantization doesn't flip that equation for most people. What it does is lower the floor: a model that used to need a $2,000 GPU to run at a usable speed might now limp along, slowly, on a $5 VPS, or run at a genuinely usable speed on a laptop that couldn't fit the full-precision version in memory at all.&lt;/p&gt;

&lt;p&gt;That's a narrow use case. If you're prototyping something that has to work entirely offline, or you're testing whether a smaller footprint model is "good enough" before committing to a paid deployment, this is worth twenty minutes of your time. If you just want a model that answers questions quickly, pay for the API and skip the custom fork.&lt;/p&gt;

&lt;p&gt;There's also a version of this that has nothing to do with speed. If you're building something where the data genuinely can't leave a customer's premises, a compression scheme that turns a 27 billion parameter model into a file smaller than a typical Blu-ray disc is a real answer to "can we even fit this on the hardware they gave us," even before you ask how fast it runs. I've had exactly one client conversation this year where that constraint was non-negotiable, and it wasn't fast, but it worked, which mattered more than tokens per second in that specific case. Most of the client infrastructure work I take on through &lt;a href="https://abrarqasim.com/work" rel="noopener noreferrer"&gt;my consulting practice&lt;/a&gt; still ends up on a hosted API for exactly this reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd actually try this week
&lt;/h2&gt;

&lt;p&gt;If you want to reproduce this without the CPU-only pain, grab whichever prebuilt binary matches your own hardware from Prism's release page, not necessarily the Ubuntu CPU one I used, and time a handful of real prompts against whatever you're currently paying for. Write down the tokens per second and the answer quality side by side. That fifteen-minute comparison will tell you more about whether "near-lossless" and "9x smaller" apply to your workload than any number in a release announcement, mine included.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/local-llm-models-bonsai-2-27b-the-9x-claim-i-actually-tested/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>localllm</category>
      <category>quantization</category>
      <category>selfhosting</category>
      <category>gguf</category>
    </item>
    <item>
      <title>How I Stopped Opening Ports With Cloudflare Tunnel</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Wed, 30 Sep 2026 08:38:08 +0000</pubDate>
      <link>https://dev.to/abyzgenic/how-i-stopped-opening-ports-with-cloudflare-tunnel-47ok</link>
      <guid>https://dev.to/abyzgenic/how-i-stopped-opening-ports-with-cloudflare-tunnel-47ok</guid>
      <description>&lt;p&gt;I almost opened port 3000 on this very server to show a client her staging dashboard. Old habit: spin up the app, open the port, send a link, forget about it until a scanner finds it. I caught myself, and instead spent twenty minutes wiring up a Cloudflare Tunnel, which is now how every internal tool on this box reaches the outside world. Zero open inbound ports on the firewall except SSH.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a tunnel actually replaces
&lt;/h2&gt;

&lt;p&gt;Before tunnels, exposing something self-hosted meant one of three options: open a port and hope your reverse proxy config is airtight, run a VPN and make every viewer install a client, or pay for a static IP and deal with port forwarding on top of it. The &lt;a href="https://github.com/cloudflare/cloudflared" rel="noopener noreferrer"&gt;cloudflared daemon&lt;/a&gt; runs as an outbound-only agent on your server, holds a persistent connection to Cloudflare's edge, and routes traffic to your app through that connection instead of through anything listening on a public port. Nothing needs to accept inbound connections except cloudflared itself, and it initiates the connection, so there's nothing for a port scanner to find.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting it up for a real service
&lt;/h2&gt;

&lt;p&gt;Here's the config I actually run for the staging dashboard, a Grafana instance living on the same Docker network:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# config.yml&lt;/span&gt;
&lt;span class="na"&gt;tunnel&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;8f2a1c9e-staging-dashboard&lt;/span&gt;
&lt;span class="na"&gt;credentials-file&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/etc/cloudflared/8f2a1c9e.json&lt;/span&gt;

&lt;span class="na"&gt;ingress&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;hostname&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;staging.clientdomain.com&lt;/span&gt;
    &lt;span class="na"&gt;service&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://grafana:3000&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;hostname&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pipeline.abrarqasim.com&lt;/span&gt;
    &lt;span class="na"&gt;service&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://localhost:8000&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;service&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http_status:404&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last catch-all rule matters more than it looks. Without it, any hostname that hits the tunnel but isn't explicitly listed falls through to cloudflared's default behavior, which in older versions meant an unhelpful blank response. An explicit 404 at the bottom means anything you haven't wired up fails obviously instead of quietly.&lt;/p&gt;

&lt;p&gt;Running it as a Docker service instead of a bare systemd unit keeps it on the same network as the containers it's routing to, which avoids a whole class of "why can't cloudflared reach my app" debugging:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;cloudflared&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cloudflare/cloudflared:latest&lt;/span&gt;
    &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;tunnel --config /etc/cloudflared/config.yml run&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;./cloudflared:/etc/cloudflared&lt;/span&gt;
    &lt;span class="na"&gt;networks&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;app_network&lt;/span&gt;
    &lt;span class="na"&gt;restart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;unless-stopped&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The DNS side is one command once the tunnel exists: &lt;code&gt;cloudflared tunnel route dns 8f2a1c9e-staging-dashboard staging.clientdomain.com&lt;/code&gt; creates a CNAME pointing at your tunnel's unique subdomain on Cloudflare's edge, and it propagates almost immediately since you're not waiting on your own registrar. I've set up a dozen or so of these now and the DNS step has never once been the part that went wrong; it's always the ingress config or a container sitting on the wrong Docker network.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gating it behind a login instead of a shared link
&lt;/h2&gt;

&lt;p&gt;A public hostname is still a public hostname, so for anything more sensitive than a demo dashboard I put &lt;a href="https://developers.cloudflare.com/cloudflare-one/policies/access/" rel="noopener noreferrer"&gt;Cloudflare Access&lt;/a&gt; in front of the same tunnel. It's a policy layer, not a separate piece of infrastructure: you tell it which email addresses or domains are allowed, and a visitor has to click through a one-time code sent to their inbox before the ingress rule above ever forwards the request. For a client dashboard, I scope the policy to her exact email address and mine, which means the tunnel hostname can leak into a Slack message or a bookmark without becoming a real exposure. It took me longer to write this paragraph than it did to set the policy up in the dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cloudflare Tunnel vs Tailscale, for the specific case of "show a client something"
&lt;/h2&gt;

&lt;p&gt;I use both, for different jobs. Tailscale is the right tool when the audience is me and maybe one other engineer, because it puts everyone on a private mesh network and nothing is reachable from the open internet at all, full stop. That's the wrong shape for a client demo: she doesn't want to install a VPN client to look at her own dashboard. A Cloudflare Tunnel terminates at a normal HTTPS hostname anyone can open in a browser, with Access rules in front of it when I want that email gate. For anything meant for a non-technical person to click a link and see, tunnel wins. For anything meant only for me and my own machines, Tailscale wins, and I'd rather run both than force one tool to do the other's job badly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure mode worth planning for
&lt;/h2&gt;

&lt;p&gt;Cloudflared is a single outbound process, and if it crashes or the box reboots without it restarting, everything behind that tunnel silently goes dark from the outside, with no local symptom at all, because the app itself is still running fine on its port. I lost about forty minutes the first time this happened to me, checking Grafana's own logs and the app container's health before it occurred to me to check whether the tunnel process was even still alive. &lt;code&gt;restart: unless-stopped&lt;/code&gt; in the compose file above is not optional, and I added a simple curl-based healthcheck that hits each tunneled hostname from an external monitor every five minutes specifically because a dead tunnel produces no error anywhere on the server itself, only silence on the client's end.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to set up this week
&lt;/h2&gt;

&lt;p&gt;If you're still opening ports for anything you self-host, the current cloudflared release installs in about ten minutes and the config above is close to a complete starting point for a single service. Start with whatever internal tool you're most tempted to expose "just for a minute," wire the ingress rule, add an Access policy if more than one person needs to see it, and check your firewall rules afterward to confirm you can actually close that port instead of just adding an unused alternative next to it. I run this exact setup on the box behind &lt;a href="https://abrarqasim.com/blog/hetzner-vs-digitalocean-what-this-blog-actually-runs-on" rel="noopener noreferrer"&gt;what this blog actually runs on&lt;/a&gt;, and the rest of the self-hosting and client infrastructure work I do is at &lt;a href="https://abrarqasim.com/work" rel="noopener noreferrer"&gt;abrarqasim.com/work&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/cloudflare-tunnel-how-i-stopped-opening-ports-for-clients/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cloudflaretunnel</category>
      <category>selfhosting</category>
      <category>devops</category>
      <category>networking</category>
    </item>
    <item>
      <title>Coolify vs Dokploy: What Actually Differs After a Real Migration</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Tue, 29 Sep 2026 16:38:05 +0000</pubDate>
      <link>https://dev.to/abyzgenic/coolify-vs-dokploy-what-actually-differs-after-a-real-migration-cm1</link>
      <guid>https://dev.to/abyzgenic/coolify-vs-dokploy-what-actually-differs-after-a-real-migration-cm1</guid>
      <description>&lt;p&gt;A client asked me to stop deploying her staging environment by SSHing in and running &lt;code&gt;git pull &amp;amp;&amp;amp; docker compose up -d --build&lt;/code&gt; like it's 2019. Fair complaint. I'd been meaning to try a self-hosted PaaS instead of hand-rolled Caddy and compose files for a while, so I spun up two identical Hetzner CPX21 boxes and installed &lt;a href="https://github.com/coollabsio/coolify" rel="noopener noreferrer"&gt;Coolify&lt;/a&gt; on one, &lt;a href="https://dokploy.com" rel="noopener noreferrer"&gt;Dokploy&lt;/a&gt; on the other, then deployed the same small Laravel app plus a Postgres database to both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting either one running takes about the same four minutes
&lt;/h2&gt;

&lt;p&gt;Both install with a single curl-to-bash command, which I'll admit made me nervous the first time and doesn't anymore, mostly because I've read both scripts now:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Coolify&lt;/span&gt;
curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://cdn.coollabs.io/coolify/install.sh | bash

&lt;span class="c"&gt;# Dokploy&lt;/span&gt;
curl &lt;span class="nt"&gt;-sSL&lt;/span&gt; https://dokploy.com/install.sh | sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both scripts do roughly the same thing: check Docker is installed, pull a handful of containers, wire up Traefik or their own proxy, and hand you a dashboard URL. Coolify's install finished in about three and a half minutes on my box. Dokploy's took closer to five, mostly because it also stands up Docker Swarm mode even for a single-node install, which Coolify only does once you add a second server.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the two actually diverge: build strategy
&lt;/h2&gt;

&lt;p&gt;This is the part that matters if you deploy anything that isn't a plain Dockerfile. Coolify defaults to &lt;a href="https://railpack.com" rel="noopener noreferrer"&gt;Railpack&lt;/a&gt;, its newer buildpack successor to Nixpacks, though Nixpacks is still there as a fallback. Dokploy sticks with Nixpacks and Heroku Buildpacks side by side, and lets you pick per application. For my Laravel app, both auto-detected PHP correctly and produced a working image without a Dockerfile, but Coolify's Railpack build was noticeably faster, about 40 seconds against Dokploy's Nixpacks build at just over a minute, largely because Railpack caches Composer layers more aggressively by default.&lt;/p&gt;

&lt;p&gt;Neither win matters if you already have a Dockerfile you trust, which is what I ended up using for the client's app anyway once I needed a custom PHP extension neither buildpack shipped.&lt;/p&gt;

&lt;h2&gt;
  
  
  Database management is where Coolify pulls ahead for me
&lt;/h2&gt;

&lt;p&gt;Coolify ships managed databases (Postgres, MySQL, MariaDB, Redis, and a handful of others) with one-click provisioning, automatic backups to S3-compatible storage, and a UI for restoring a specific backup without touching a shell. Dokploy supports the same list of engines, and its backup story is solid too, but the restore flow currently routes through the CLI rather than a full point-and-click restore in the dashboard. For a client who isn't going to SSH into anything, that's the difference between "she can restore last night's backup herself" and "she has to text me."&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# the compose service definition both platforms happily imported as-is&lt;/span&gt;
&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;.&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;APP_ENV&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
      &lt;span class="na"&gt;DB_HOST&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgres&lt;/span&gt;
    &lt;span class="na"&gt;depends_on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;postgres&lt;/span&gt;
  &lt;span class="na"&gt;postgres&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgres:17&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;pgdata:/var/lib/postgresql/data&lt;/span&gt;
&lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pgdata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Networking and multi-server access work differently too
&lt;/h2&gt;

&lt;p&gt;Custom domains and automatic HTTPS worked the same way on both within a couple of minutes of pointing DNS at the box, so that's a wash. The gap shows up once you add a second server. Coolify treats additional servers as first-class from the start, meaning any app can target any connected server from the same dashboard, and it layers team roles and permissions on top so I can give a client's other contractor deploy access without handing over root. Dokploy's multi-server story leans harder on Docker Swarm itself once you're past a single node, which is more powerful if you already think in Swarm terms and more friction if you don't. Coolify also ships an MCP server alongside its API and CLI, which is a strange thing to care about until you realize it means an AI coding agent can query deployment status or trigger a redeploy directly, something I've started leaning on for routine restarts instead of opening the dashboard at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Both projects ship faster than you can keep up with
&lt;/h2&gt;

&lt;p&gt;Coolify is at v4.3.23 as I write this, four patch releases in the last two weeks alone, and it carries 62,000 GitHub stars. Dokploy is close behind on velocity at v0.30.7 with a similar two-week cadence, and sits at about 37,500 stars, roughly a year and a half younger as a project. Neither release cadence is a red flag; if anything it means real bugs get fixed in days rather than sitting in a backlog. It does mean you should read the changelog before you click update on a client's production box, the same rule I'd apply to n8n or anything else that ships this often. I got burned once treating a Coolify minor version bump as routine and losing a custom Traefik label override that a new default had quietly replaced.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd actually recommend
&lt;/h2&gt;

&lt;p&gt;For a single Hetzner box running two or three client apps, either tool removes real pain over hand-rolled compose and Caddy files, and I wouldn't talk anyone out of either one. I moved this specific client to Coolify, mainly for the one-click backup restore and because Railpack's build speed adds up when I'm redeploying a dozen times a day during active work. If she ever needs true multi-node Docker Swarm scaling from day one rather than as an add-on, Dokploy's Swarm-first design would be the better starting point. I wrote about the box underneath both of these in &lt;a href="https://abrarqasim.com/blog/hetzner-vs-digitalocean-what-this-blog-actually-runs-on" rel="noopener noreferrer"&gt;what this blog actually runs on&lt;/a&gt;, and if you want help picking or migrating between any of this, that's the kind of work I take on, described at &lt;a href="https://abrarqasim.com/work" rel="noopener noreferrer"&gt;abrarqasim.com/work&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you're running either tool right now, the one thing worth doing this week regardless of which you picked: manually trigger a database restore into a throwaway app once, before you need it for real. I found Dokploy's CLI restore step only after I needed it under pressure, and that's a bad time to be reading docs for the first time.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/coolify-vs-dokploy-what-actually-differs-after-migrating-a-client/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>coolify</category>
      <category>dokploy</category>
      <category>selfhosting</category>
      <category>devops</category>
    </item>
    <item>
      <title>The n8n Workflow That Broke Silently for Six Weeks</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Tue, 29 Sep 2026 06:42:45 +0000</pubDate>
      <link>https://dev.to/abyzgenic/the-n8n-workflow-that-broke-silently-for-six-weeks-167h</link>
      <guid>https://dev.to/abyzgenic/the-n8n-workflow-that-broke-silently-for-six-weeks-167h</guid>
      <description>&lt;p&gt;A client asked me in July to stop forwarding her contact form leads to a spreadsheet and just handle them. Score the lead, draft a reply, ping her on Slack if it looked like a real budget, archive it if it looked like a recruiter spamming her about a job she didn't want. I built the whole thing in n8n in an afternoon. It ran clean for six weeks. Then it silently stopped tagging anything as high-priority, and nobody noticed until she asked why she hadn't heard from a promising lead in eleven days.&lt;/p&gt;

&lt;h2&gt;
  
  
  The workflow, roughly
&lt;/h2&gt;

&lt;p&gt;Here's the shape of it, trimmed to the part that matters. A webhook catches the form submission, a Function node scores it, an If node branches on that score:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"nodes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Webhook"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"n8n-nodes-base.webhook"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Score Lead"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"n8n-nodes-base.function"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"High Priority?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"n8n-nodes-base.if"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Slack Alert"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"n8n-nodes-base.slack"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Draft Reply"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"n8n-nodes-base.httpRequest"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the actual scoring logic, which is where the whole thing quietly broke:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;budget&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;$input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;budget&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;$input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;$&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;json&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;$input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;score&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;body.budget&lt;/code&gt; field existed because the form on her site sent it as a top-level key. In August she redesigned the form with a page builder plugin, and the plugin nested every field under a &lt;code&gt;formData&lt;/code&gt; object instead of putting it at the root. The webhook still fired. The payload still arrived. &lt;code&gt;$input.item.json.body.budget&lt;/code&gt; just quietly evaluated to &lt;code&gt;undefined&lt;/code&gt;, &lt;code&gt;"".includes("$")&lt;/code&gt; returned false every single time, and every lead scored zero regardless of budget. No error, no failed execution in the n8n logs, because nothing actually threw.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is worse than a normal outage
&lt;/h2&gt;

&lt;p&gt;A crashed workflow is loud. n8n's execution list turns red, and if you've wired up error workflows (you should, and I hadn't on this one) you get a notification within seconds. A workflow that runs to completion and produces a wrong answer is silent by design, because as far as n8n is concerned, nothing went wrong. The Function node executed, returned valid JSON, and passed it along. The bug lived entirely in an assumption about the shape of somebody else's data, which is exactly the kind of thing that survives every test you write yourself because you're the one who wrote the test data too.&lt;/p&gt;

&lt;p&gt;I now add a one-line sanity check at the top of every scoring function that touches an external payload:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;budget&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="k"&gt;in &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;$input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="p"&gt;{})))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;payload shape changed: no budget field&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cheap, ugly, and it turns a silent miscategorization into a loud failed execution that shows up in n8n's error workflow. I'd rather get paged for a broken assumption than find out from a client eleven days later.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I actually caught it
&lt;/h2&gt;

&lt;p&gt;She noticed the drought before I did, which tells you everything about how I was monitoring this thing at the time: not at all. I'd built the workflow, watched it work for a week, and moved on to the next client project, the classic freelancer mistake of treating a shipped automation like a shipped feature instead of a running service. Once she flagged it, finding the bug took ten minutes with n8n's execution history, since every run was sitting right there with its input and output JSON. Finding out it had been broken for six weeks took a client noticing a gap in her own pipeline, and that part I actually feel bad about.&lt;/p&gt;

&lt;p&gt;What I run now on every client workflow that touches money or leads: a second, tiny n8n workflow on a daily cron that queries the main workflow's execution count for the last 24 hours through n8n's own API, and pings me on Slack if that count drops to zero or if the average score across all of yesterday's leads is suspiciously flat. A flat score distribution was exactly the fingerprint of my bug: every lead scoring 0, day after day. A simple standard-deviation check on that score would have caught it on day one instead of day forty-two. It's maybe fifteen lines of Function-node code, and it's a lot cheaper than the one honest conversation I had to have about why a warm lead went cold for over a month.&lt;/p&gt;

&lt;h2&gt;
  
  
  n8n's pace is part of the story here
&lt;/h2&gt;

&lt;p&gt;n8n ships fast enough that "which version am I even running" is a real question worth asking before you debug anything else. Checking the &lt;a href="https://github.com/n8n-io/n8n/releases" rel="noopener noreferrer"&gt;n8n releases page&lt;/a&gt; while writing this, the project cut nine separate tagged releases across its stable and beta lines in the four days leading up to September 25 alone. If you're self-hosting via Docker and you pull &lt;code&gt;n8nio/n8n:latest&lt;/code&gt; on a cron job the way a lot of tutorials tell you to, you can wake up to genuinely different node behavior with zero warning. I pin a specific tag now and read the changelog before bumping it, the same discipline I'd apply to any other dependency, which n8n absolutely is even though the interface makes it feel like a website you're clicking around in rather than software you're running.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where n8n actually earns its keep over Zapier or Make
&lt;/h2&gt;

&lt;p&gt;For a client on a budget, n8n's self-hosted option removes the per-task pricing that makes Zapier expensive the moment a workflow runs more than a few hundred times a month. I've moved three client automations off Zapier this year purely on cost, not features. The tradeoff is that you now own the box it runs on, the Postgres database behind it, and the debugging when a node behaves differently than the docs suggest. That's a fair trade for a workflow that runs thousands of times a month; it's a bad trade for something a client runs twice a week, where Zapier's hosted reliability is worth the per-task fee and the lack of a server to patch.&lt;/p&gt;

&lt;p&gt;I wrote about the harder end of this same tradeoff, an automation that actively hurt a client relationship instead of just quietly under-delivering, in &lt;a href="https://abrarqasim.com/blog/when-not-to-use-ai-automation-the-refund-bot-that-cost-me-a-client" rel="noopener noreferrer"&gt;the refund bot that cost me a client&lt;/a&gt;. Silent miscategorization and an overconfident refund bot are the same root problem wearing different clothes: automation that never tells you when its assumptions stopped holding.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to check this week if you run anything on n8n
&lt;/h2&gt;

&lt;p&gt;Open your busiest workflow and look for any Function node that reaches into a nested field from a webhook, form, or third-party API payload. Add the one-line guard above to each one. Then go check whether you've actually wired an error workflow to catch failed executions, not just left the default "do nothing" behavior, because that's the notification that would have caught my lead-scoring bug in week one instead of week six. If you're doing freelance automation work for clients, this is also the kind of detail worth mentioning up front; I cover how I scope that conversation in &lt;a href="https://abrarqasim.com/blog/how-i-actually-price-freelance-web-development-work" rel="noopener noreferrer"&gt;my writeup on pricing this work honestly&lt;/a&gt;, and you can see the rest of what I build at &lt;a href="https://abrarqasim.com" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/n8n-workflow-automation-the-bug-that-cost-a-client-eleven-days/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>n8n</category>
      <category>workflowautomation</category>
      <category>automation</category>
      <category>freelance</category>
    </item>
    <item>
      <title>How I Actually Price Freelance Web Development Work</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Mon, 28 Sep 2026 14:42:38 +0000</pubDate>
      <link>https://dev.to/abyzgenic/how-i-actually-price-freelance-web-development-work-o8d</link>
      <guid>https://dev.to/abyzgenic/how-i-actually-price-freelance-web-development-work-o8d</guid>
      <description>&lt;p&gt;A client asked me for a "quick rate" over email three years ago, and I typed a number I'd basically made up on the spot, based on nothing more than what felt reasonable for a Tuesday. I got the job. I also spent the next six weeks resenting every hour of it, because the number I'd picked assumed a simple integration and the actual scope turned out to involve migrating a decade of legacy data out of a system nobody had documented. I didn't renegotiate. I just worked longer hours and told myself that's what freelancing meant.&lt;/p&gt;

&lt;p&gt;It isn't. I've changed how I price work twice since then, and the current approach is the first one that's actually held up under a messy, real project instead of just looking clean on a spreadsheet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hourly rates aren't wrong, they're just answering the wrong question
&lt;/h2&gt;

&lt;p&gt;There's nothing broken about hourly billing as a concept. The problem is what it optimizes for. If I bill hourly, I get paid more the slower I work and less the faster I solve the client's problem, which is backwards from what the client actually wants and backwards from what I want too, since efficient work is the entire reason clients hire an experienced freelancer instead of training a junior in-house. &lt;a href="https://jonathanstark.com/hbin" rel="noopener noreferrer"&gt;Jonathan Stark has been making this argument for years&lt;/a&gt;, and the part that finally landed for me wasn't the theory, it was noticing my own behavior. I caught myself, more than once, taking a slightly longer path through a problem because I knew a faster fix would mean a smaller invoice. That's not a character flaw. It's what the incentive structure produces, in me and in basically anyone billing by the hour.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://survey.stackoverflow.co/2025/work" rel="noopener noreferrer"&gt;2025 Stack Overflow Developer Survey&lt;/a&gt; still shows a wide compensation spread among independent and freelance developers, wider than salaried roles show, and I don't think that's purely a skill gap. A meaningful chunk of it is pricing method. Two developers with comparable skill can land in very different places depending on whether one is trading hours for dollars and the other is pricing the outcome.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually quote now
&lt;/h2&gt;

&lt;p&gt;For anything with a defined scope, a new feature, a migration, a redesign, I quote a fixed project price, not a rate. I still track my hours privately so I can calibrate future quotes, but the client never sees an hourly number. The quote is built from three questions I ask myself before I write it: what does this project need to actually solve for the client, what's the real risk that scope grows once I'm inside the codebase, and what would I charge if I could only send one invoice for the whole thing.&lt;/p&gt;

&lt;p&gt;That third question changes the number more than the first two combined. When I imagine sending exactly one invoice with no do-overs, I stop lowballing the parts of the project I'm least certain about, the parts where "should take a day" has a real chance of taking four.&lt;/p&gt;

&lt;p&gt;For ongoing work, a retainer or an indefinite maintenance relationship, I do still use a day rate rather than a project price, because the scope genuinely isn't fixed and pretending otherwise would just mean padding a fixed number to cover uncertainty I can't quantify yet. Day rate, not hourly. Clients stop watching the clock and start describing what they need, which produces better conversations than a stopwatch relationship ever did.&lt;/p&gt;

&lt;h2&gt;
  
  
  The discovery call is where the pricing actually happens
&lt;/h2&gt;

&lt;p&gt;I used to treat the discovery call as a formality before sending a quote. It's the opposite. The call is where I find the information that makes the quote accurate instead of a guess. I ask what happens if this project is late, not because I want to threaten anyone with lateness, but because the answer tells me how much risk I'm actually taking on. "Nothing happens, we'll just launch next quarter instead" is a different project than "our biggest client renews in six weeks and this has to be live before then." Same scope, same code, genuinely different price, because the second one is buying certainty as much as it's buying software.&lt;/p&gt;

&lt;p&gt;I also ask directly who else has quoted this and what they said, not to undercut a competitor, but because a wildly different range from another freelancer usually means we're scoping different things, and finding that mismatch before I quote saves both of us from a bad first few weeks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scope creep is a pricing problem wearing a project management costume
&lt;/h2&gt;

&lt;p&gt;Most scope creep advice focuses on process: change orders, signed amendments, tracked hours against a budget. Useful, but it treats the symptom. The actual cause, in my experience, is that the original quote never accounted for the parts of the project that were genuinely unknown at quoting time. You can't change-order your way out of a quote that was wrong from the start; you can only patch it, and patched quotes are where resentment lives on both sides.&lt;/p&gt;

&lt;p&gt;Now I build a specific line into every proposal: a fixed number of hours reserved for "things we find once we're inside the code that weren't visible from outside it." I call it exactly that, in plain language, not a vague buffer percentage. Clients respond well to it because it's honest about the actual nature of the work instead of pretending I can see through a black box from a kickoff call.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I do when a client pushes back on the number
&lt;/h2&gt;

&lt;p&gt;I used to fold immediately, because I hate confrontation about money more than I hate most things, and my quotes at the time were often a first offer I expected to negotiate down from anyway. Two changes fixed this. First, I stopped quoting a padded first number, so there's genuinely less room to negotiate downward without cutting the project's scope along with the price. Second, I started answering pushback with a scope question instead of a price question: "what part of this would you want to cut to hit that number?" That reframes the conversation from "convince me I'm worth this" to "let's agree on what we're actually building," which is a conversation I'm far more comfortable having, and one that usually produces a smaller, cleaner project instead of an awkward discount on the original one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The math I actually run before I type a number
&lt;/h2&gt;

&lt;p&gt;When I'm building a fixed quote, I still start from a rough hours estimate, I just don't stop there. I take my target day rate, multiply by my honest estimate of days, then apply a multiplier based on how well I understand the codebase going in. A greenfield project where I'm writing every line gets close to a 1.1x multiplier, barely any padding, because the unknowns are mine to control. A project where I'm the fourth developer to touch an unfamiliar codebase gets 1.4x or higher, because the discovery call can't show me what a previous developer left half-finished three directories deep. That multiplier isn't padding in the dishonest sense. It's pricing the actual risk I'm taking on by agreeing to a fixed number against work I can't fully see yet, and naming it explicitly, even just to myself, keeps me from either underpricing real uncertainty or overcharging a client whose codebase turns out to be cleaner than I expected.&lt;/p&gt;

&lt;p&gt;I also stopped rounding quotes to numbers that feel comfortable, the way $5,000 feels safer to type than $5,400. A specific number reads as calculated rather than guessed, and in three years of sending both kinds, the specific ones get fewer "can we talk about the price" emails, not more.&lt;/p&gt;

&lt;h2&gt;
  
  
  One thing to try this week
&lt;/h2&gt;

&lt;p&gt;Take the last project you quoted hourly and rewrite the quote as a single fixed number, using the "one invoice, no do-overs" question above. Don't send it anywhere. Just see what the number does to how you'd approach the work, and notice which parts of the original scope you'd suddenly want nailed down before you started. That gap between the two numbers is usually where the real pricing problem was hiding.&lt;/p&gt;

&lt;p&gt;I wrote a longer breakdown of the actual proposal document I send after a discovery call in &lt;a href="https://abrarqasim.com/blog/web-development-proposal-template-the-section-that-wins-the-job" rel="noopener noreferrer"&gt;an earlier post on the section that wins the job&lt;/a&gt;, if you want the template alongside the pricing logic here. More about how I run this as a business is on &lt;a href="https://abrarqasim.com/about" rel="noopener noreferrer"&gt;my about page&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/how-i-actually-price-freelance-web-development-work/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>freelance</category>
      <category>pricing</category>
      <category>business</category>
      <category>consulting</category>
    </item>
    <item>
      <title>Postgres Performance Tuning Parameters I Actually Touch</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Mon, 28 Sep 2026 06:42:35 +0000</pubDate>
      <link>https://dev.to/abyzgenic/postgres-performance-tuning-parameters-i-actually-touch-4aak</link>
      <guid>https://dev.to/abyzgenic/postgres-performance-tuning-parameters-i-actually-touch-4aak</guid>
      <description>&lt;p&gt;My VPS ran an autovacuum job for six straight hours last winter on a table with maybe four million rows, which is not a lot of rows, and I spent that evening convinced I'd found some exotic Postgres bug. I hadn't. I'd just never tuned autovacuum past its defaults, because the defaults work fine right up until they don't, and "fine" had quietly become "grinding through a table lock during my only maintenance window."&lt;/p&gt;

&lt;p&gt;That's the thing about Postgres performance tuning parameters: almost nobody touches them until something hurts. I want to walk through the handful I actually adjust on every server I run, plus what caught my eye in &lt;a href="https://www.postgresql.org/about/news/postgresql-19-beta-4-released-3386/" rel="noopener noreferrer"&gt;PostgreSQL 19 Beta 4&lt;/a&gt;, released this week, because a couple of the changes land directly on the knobs I care about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Autovacuum is the one that actually bites people
&lt;/h2&gt;

&lt;p&gt;The default &lt;code&gt;autovacuum_vacuum_cost_delay&lt;/code&gt; throttles vacuum aggressively so it doesn't compete with your live traffic for I/O. That's a reasonable default for a shared host you don't control. It's a bad default for a dedicated VPS where you're the only tenant and dead tuples are piling up faster than the throttled vacuum can clear them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;SYSTEM&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;autovacuum_vacuum_cost_limit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;SYSTEM&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;autovacuum_vacuum_cost_delay&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'2ms'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;SYSTEM&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;autovacuum_naptime&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'15s'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;pg_reload_conf&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Raising the cost limit and dropping the delay lets vacuum work through more pages per cycle, which matters a lot once a table has any real write volume. &lt;code&gt;autovacuum_naptime&lt;/code&gt; controls how often the launcher checks whether a table needs vacuuming at all; the default 60 seconds is fine for most workloads, but a table that gets bursts of updates benefits from checking more often.&lt;/p&gt;

&lt;p&gt;PostgreSQL 19 Beta 4 actually touches this directly. The release notes mention "several fixes to the new autovacuum scoring system," which replaced the older threshold-based trigger with something that weighs tables by how urgently they need attention instead of a flat percentage. I haven't run it in production yet, betas don't belong there, but I've been testing it against a staging clone of the same table that gave me that six-hour vacuum, and the scoring system picked it up for vacuum noticeably earlier than the old threshold would have.&lt;/p&gt;

&lt;h2&gt;
  
  
  work_mem, and the query that taught me to stop guessing
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;work_mem&lt;/code&gt; sets how much memory a single sort or hash operation gets before it spills to disk. The default, 4MB, is conservative enough to be safe on a shared box and too small for almost any real analytical query.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;EXPLAIN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;ANALYZE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;BUFFERS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;4821&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If that plan shows &lt;code&gt;Sort Method: external merge Disk: 84000kB&lt;/code&gt;, your sort spilled to disk instead of staying in memory, and that's your signal to raise &lt;code&gt;work_mem&lt;/code&gt;, not to add an index blindly. I used to jump straight to indexing every slow query. Sometimes the query is fine and the sort just needed more headroom.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;work_mem&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'64MB'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;-- per session, test first&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;SYSTEM&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;work_mem&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'32MB'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;-- global, be careful&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The global setting multiplies per connection and per sort operation within a query, so a server running fifty connections each doing a two-way hash join can burn through memory fast if you set this too high. I test with a session-level &lt;code&gt;SET&lt;/code&gt; against a copy of production data before I touch the global value.&lt;/p&gt;

&lt;h2&gt;
  
  
  shared_buffers is not "give Postgres all your RAM"
&lt;/h2&gt;

&lt;p&gt;The advice to set &lt;code&gt;shared_buffers&lt;/code&gt; to 25% of system RAM gets repeated everywhere, and it's a reasonable starting point, but I've seen people push it to 60 or 70% assuming more is strictly better. It isn't. Postgres relies on the operating system's page cache as a second layer, and a &lt;code&gt;shared_buffers&lt;/code&gt; value that's too aggressive starves that cache, which can make things slower, not faster, especially on a box that's also running your application.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;SYSTEM&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;shared_buffers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'2GB'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;-- on an 8GB box&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;SYSTEM&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;effective_cache_size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'6GB'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;-- OS cache estimate&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;effective_cache_size&lt;/code&gt; doesn't allocate anything. It's a hint to the query planner about how much data is likely already cached somewhere, and it affects whether the planner favors an index scan over a sequential scan. I set it to roughly 75% of total RAM and haven't had a reason to touch it since.&lt;/p&gt;

&lt;h2&gt;
  
  
  REPACK, the new command that replaces a maintenance script I've run for years
&lt;/h2&gt;

&lt;p&gt;This is the change from PG19 that actually excited me. Reclaiming bloated table space has meant &lt;code&gt;VACUUM FULL&lt;/code&gt;, which takes an exclusive lock and blocks reads and writes for the duration, or &lt;code&gt;pg_repack&lt;/code&gt;, a well-maintained extension that does it online but still requires installing and trusting a third-party tool. PostgreSQL 19 adds &lt;code&gt;REPACK&lt;/code&gt; as a first-party command.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;REPACK&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Beta 4's changelog lists several fixes to REPACK specifically: crashes on invalid indexes, incorrect behavior with materialized views, and permission and error-reporting corrections. That tells me the feature is getting real testing pressure before GA, which is exactly what I want to see before I trust a command that rewrites a whole table. I'm not running it against anything that matters yet, but it's the first thing I'm testing once 19 goes stable, because it would let me delete a &lt;code&gt;pg_repack&lt;/code&gt; cron job I've maintained across three different servers.&lt;/p&gt;

&lt;h2&gt;
  
  
  checkpoint_completion_target, the setting nobody mentions until a write spike hits
&lt;/h2&gt;

&lt;p&gt;Postgres writes dirty pages to disk during a checkpoint, and by default it tries to finish that work quickly, which can cause a burst of I/O contention right when your application is also trying to write. &lt;code&gt;checkpoint_completion_target&lt;/code&gt; tells Postgres how much of the interval between checkpoints it should spread that writing across.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;SYSTEM&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;checkpoint_completion_target&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;9&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;SYSTEM&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;max_wal_size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'4GB'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Raising it toward 0.9 spreads the write load across nearly the whole checkpoint interval instead of compressing it into a spike near the end. I found this one the hard way, watching &lt;code&gt;iostat&lt;/code&gt; during a nightly batch import and seeing disk write throughput spike in a pattern that lined up exactly with &lt;code&gt;checkpoint_timeout&lt;/code&gt;, which defaults to five minutes. &lt;code&gt;max_wal_size&lt;/code&gt; works alongside it; a bigger WAL budget means fewer checkpoints overall, which is the other lever if checkpoints themselves, not their spread, are the problem.&lt;/p&gt;

&lt;p&gt;I test changes like this with &lt;code&gt;pg_stat_bgwriter&lt;/code&gt; before and after, specifically the ratio of &lt;code&gt;buffers_checkpoint&lt;/code&gt; to &lt;code&gt;buffers_clean&lt;/code&gt;. A checkpoint doing most of the writing, instead of the background writer handling it gradually, is the signal that this setting needs attention on a given server.&lt;/p&gt;

&lt;h2&gt;
  
  
  max_connections is usually the wrong lever
&lt;/h2&gt;

&lt;p&gt;When an app starts throwing "too many connections" errors, the instinct is to raise &lt;code&gt;max_connections&lt;/code&gt;. I did this on a client project two years ago, doubled it from 100 to 200, and the server got slower, not faster. Each connection carries its own memory overhead and its own backend process, and Postgres doesn't scale linearly past a few hundred active connections on modest hardware. The actual fix was a connection pooler.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="c"&gt;# pgbouncer.ini
&lt;/span&gt;&lt;span class="nn"&gt;[databases]&lt;/span&gt;
&lt;span class="py"&gt;app&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;host=127.0.0.1 port=5432 dbname=app&lt;/span&gt;

&lt;span class="nn"&gt;[pgbouncer]&lt;/span&gt;
&lt;span class="py"&gt;pool_mode&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;transaction&lt;/span&gt;
&lt;span class="py"&gt;max_client_conn&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;1000&lt;/span&gt;
&lt;span class="py"&gt;default_pool_size&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;20&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PgBouncer in transaction mode lets a thousand application-side connections share twenty real Postgres backends, handing each one back to the pool the moment its transaction commits. That's a very different problem from raising &lt;code&gt;max_connections&lt;/code&gt;, and it's the one that actually fixes the error message instead of just delaying it. I run PgBouncer on every Postgres box I manage now, even the small ones, because retrofitting it under load is a worse afternoon than setting it up in advance.&lt;/p&gt;

&lt;h2&gt;
  
  
  The features that got pulled, and why that's a good sign
&lt;/h2&gt;

&lt;p&gt;Beta 4 also reverted a handful of things that were planned for 19: SQL/PGQ property graph query support, online toggling of data checksums, and the &lt;code&gt;ALTER TABLE ... MERGE PARTITIONS&lt;/code&gt; and &lt;code&gt;SPLIT PARTITIONS&lt;/code&gt; commands. Watching a release cut scope this late used to worry me, it feels like something going wrong. Reading through the Postgres project's own reasoning changed my mind: they're explicit that reliability comes before the release calendar, and features that aren't ready get pushed to a later major version instead of shipping half-finished. That's a genuinely rare property in software release culture, and it's a big part of why I still run Postgres on a VPS instead of reaching for a managed service with a bigger feature list and less of that discipline behind it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually do this week
&lt;/h2&gt;

&lt;p&gt;Run &lt;code&gt;EXPLAIN (ANALYZE, BUFFERS)&lt;/code&gt; against your three slowest queries and check for &lt;code&gt;external merge Disk&lt;/code&gt; in the sort output before you touch &lt;code&gt;work_mem&lt;/code&gt;. Then check &lt;code&gt;pg_stat_user_tables&lt;/code&gt; for &lt;code&gt;n_dead_tup&lt;/code&gt; on your largest tables; if that number is climbing faster than autovacuum is clearing it, that's your &lt;code&gt;autovacuum_vacuum_cost_delay&lt;/code&gt; problem before it becomes a six-hour vacuum on a Tuesday night.&lt;/p&gt;

&lt;p&gt;I've written before about a &lt;a href="https://abrarqasim.com/blog/laravel-postgres-unindexed-foreign-keys-vacuum-lint-in-ci" rel="noopener noreferrer"&gt;Postgres foreign key mistake&lt;/a&gt; that made this exact vacuum problem worse than it needed to be, worth a read if you're touching autovacuum settings anyway. The full 19 release notes are at &lt;a href="https://www.postgresql.org/docs/19/release-19.html" rel="noopener noreferrer"&gt;postgresql.org/docs/19/release-19.html&lt;/a&gt; if you want the complete changelog, and more of the infrastructure work behind this blog is documented on &lt;a href="https://abrarqasim.com/work" rel="noopener noreferrer"&gt;my portfolio&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/postgres-performance-tuning-parameters-i-actually-touch/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>postgres</category>
      <category>devops</category>
      <category>performance</category>
    </item>
    <item>
      <title>LLM Pricing Comparison: The Cached-Token Math Behind the Price War</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Sat, 26 Sep 2026 06:42:18 +0000</pubDate>
      <link>https://dev.to/abyzgenic/llm-pricing-comparison-the-cached-token-math-behind-the-price-war-5abd</link>
      <guid>https://dev.to/abyzgenic/llm-pricing-comparison-the-cached-token-math-behind-the-price-war-5abd</guid>
      <description>&lt;p&gt;Short version for the impatient: after this week's releases, the cheapest model on the price sheet is not always the cheapest model on your invoice. Which one wins depends on how much of your input is cached, and most LLM pricing comparisons leave that column out of the conversation.&lt;/p&gt;

&lt;p&gt;Here's how I got here. On Monday Grok 4.7 came out. On Tuesday Anthropic shipped Claude Opus 5.5, and about an hour later OpenAI shipped GPT-6 Sol and GPT-6 Luna. Simon Willison put &lt;a href="https://simonwillison.net/2026/Sep/22/opus-and-sol-and-luna/" rel="noopener noreferrer"&gt;all the new prices in one table&lt;/a&gt;, which saved me an evening of tab-hopping. Then I did what I suspect half of you did. I sorted by output price and started mentally moving client projects around.&lt;/p&gt;

&lt;p&gt;I stopped myself, because I've been burned by that exact sort before. I wrote about the general problem in &lt;a href="https://abrarqasim.com/blog/llm-cost-optimization-after-the-free-lunch-ended/" rel="noopener noreferrer"&gt;LLM cost optimization after the free lunch ended&lt;/a&gt;. This time I wanted to do the boring arithmetic first. It changed the ranking, and not in the direction I expected.&lt;/p&gt;

&lt;h2&gt;
  
  
  The price sheet everyone is sharing
&lt;/h2&gt;

&lt;p&gt;Prices per million tokens, from Simon's table:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Cached input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Luna&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$0.01&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Sol&lt;/td&gt;
&lt;td&gt;$2&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4.7&lt;/td&gt;
&lt;td&gt;$2&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5.5&lt;/td&gt;
&lt;td&gt;$4&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5.1&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;td&gt;$1&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Look at only the first and last columns and two conclusions pop out. Grok 4.7 has by far the cheapest output of the mid-tier models. Fable 5.1 and GPT-6 Astra cost the same. For the work I actually bill for, both conclusions are wrong.&lt;/p&gt;

&lt;p&gt;The Opus change is also bigger than the headline. Anthropic's &lt;a href="https://www.anthropic.com/claude-opus-5-5" rel="noopener noreferrer"&gt;Opus 5.5 page&lt;/a&gt; lists input and output at $4 and $20, down 20% from Opus 5's $5 and $25. The line people skim past is cache reads: $0.20 per million, down from $0.50. That's a 60% cut on the column that dominates agent workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two workloads, two different winners
&lt;/h2&gt;

&lt;p&gt;I priced two request shapes, because they're the two I run most often.&lt;/p&gt;

&lt;p&gt;The first is a plain chat call. A 2,000-token prompt with nothing cached and a 500-token answer. Think of a support-reply draft, or a classification call with a longish instruction block.&lt;/p&gt;

&lt;p&gt;The second is one turn of an agent loop. 200,000 tokens of context, of which 180,000 are a cache hit from the previous turn, plus 8,000 tokens of output. Simon mentions that in longer agentic conversations 90% or more of input tokens go through at the cached price, so 90% felt like a fair middle guess.&lt;/p&gt;

&lt;p&gt;Here's the tiny function I used. It ignores cache writes (Anthropic charges extra to write the cache, $5 per million on Opus 5.5), which slightly flatters the Claude numbers on the first turn of a session and makes no real difference after that.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;PRICES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;  &lt;span class="c1"&gt;# per million tokens: input, cached input, output
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-6-luna&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;      &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.50&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-6-sol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;       &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;2.00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;10.00&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;grok-4.7&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;2.00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;6.00&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5.5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;4.00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;20.00&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-fable-5.1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:(&lt;/span&gt;&lt;span class="mf"&gt;10.00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.25&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;50.00&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-6-astra&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;     &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;10.00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;1.00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;50.00&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;input_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cached_ratio&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;output_tokens&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;p_in&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p_cached&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p_out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;PRICES&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;cached&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;input_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;cached_ratio&lt;/span&gt;
    &lt;span class="n"&gt;fresh&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;input_tokens&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;cached&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fresh&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;p_in&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;cached&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;p_cached&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;output_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;p_out&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1_000_000&lt;/span&gt;

&lt;span class="n"&gt;chat&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;call_cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;PRICES&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;call_cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;200_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;PRICES&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the results, per call:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Chat call&lt;/th&gt;
&lt;th&gt;Agent turn&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Luna&lt;/td&gt;
&lt;td&gt;$0.0004&lt;/td&gt;
&lt;td&gt;$0.0078&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4.7&lt;/td&gt;
&lt;td&gt;$0.0070&lt;/td&gt;
&lt;td&gt;$0.1780&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Sol&lt;/td&gt;
&lt;td&gt;$0.0090&lt;/td&gt;
&lt;td&gt;$0.1560&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5.5&lt;/td&gt;
&lt;td&gt;$0.0180&lt;/td&gt;
&lt;td&gt;$0.2760&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5.1&lt;/td&gt;
&lt;td&gt;$0.0450&lt;/td&gt;
&lt;td&gt;$0.6450&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;$0.0450&lt;/td&gt;
&lt;td&gt;$0.7800&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For the chat call, Grok 4.7 comes in about 22% under GPT-6 Sol. For the agent turn the order flips, and Sol is about 12% cheaper than Grok. The whole difference is the cached column. Grok charges $0.50 per million for cache hits and Sol charges $0.20, and when 180,000 tokens per turn are cache hits, that column is most of your input bill.&lt;/p&gt;

&lt;p&gt;Fable and Astra split the same way. Identical chat cost, but on the agent turn Fable is about 17% cheaper because its cache reads cost a quarter of Astra's.&lt;/p&gt;

&lt;p&gt;Scale it up and the gaps stop being rounding errors. At 1,000 agent turns a day for 30 days, Sol runs about $4,680 and Grok about $5,340. Opus 5.5 lands around $8,280 against roughly $11,700 for the same traffic on Opus 5 pricing. That's a 29% drop from price alone, which is more than the 20% headline. Anthropic claims 40% on typical workloads because Opus 5.5 also uses fewer tokens per task. I can't verify the token-efficiency half yet, so I'm only counting the part I can compute.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cost nobody puts in a table: runaway output
&lt;/h2&gt;

&lt;p&gt;Here's the number that worried me more than any row above. In Simon's usual pelican-on-a-bicycle test, Opus 5.5 at the "max" thinking level hit the 128,000-token output cap while still reasoning and returned nothing. He tried twice. Each failure cost $2.56 and took nearly 20 minutes.&lt;/p&gt;

&lt;p&gt;That $2.56 is just 128,000 times $20 per million. It's the worst-case price of a single call, and you pay it for an empty response. On my agent-turn numbers above, one of those failures costs as much as about nine normal Opus 5.5 turns. The same blowout on GPT-6 Luna would cost about six cents, which is its own argument for using cheap models on anything open-ended.&lt;/p&gt;

&lt;p&gt;I haven't run max effort on real client work, so I don't know how often this happens outside a deliberately silly prompt. One test is one test. But I don't need to know the frequency to put a ceiling on it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;max_call_cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_output_tokens&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p_out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;PRICES&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;max_output_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;p_out&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1_000_000&lt;/span&gt;

&lt;span class="c1"&gt;# pick the output cap from a budget, not the other way round
&lt;/span&gt;&lt;span class="n"&gt;BUDGET_PER_CALL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.40&lt;/span&gt;
&lt;span class="n"&gt;cap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BUDGET_PER_CALL&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;PRICES&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5.5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# 20,000 tokens
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If a task legitimately needs more than that, I'd rather find out from a truncated response I can inspect than from the invoice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm still unsure about
&lt;/h2&gt;

&lt;p&gt;Price per token says nothing about how many turns a model needs. If a cheap model takes three tries to do what a pricier one does in one, a 3x price gap disappears. Luna is about 20 times cheaper than Sol per agent turn in my numbers, and I have no idea yet what its retry rate looks like on my workloads. The one real data point I have is secondhand: Simon moved his Datasette Agent demo to GPT-6 Luna and found it fast and competent at SQL and at building HTML and JavaScript. That's encouraging. It isn't my workload.&lt;/p&gt;

&lt;p&gt;The other thing I'm unsure about is how long any of this holds. Simon notes GPT-5.6 has a 25% price increase scheduled for November, and Anthropic says Sonnet 5.5 and Haiku 5.5 are coming. Haiku 4.5 currently sits at $1 and $5, ten times Luna. If Haiku 5.5 lands anywhere near Luna, half of this table changes again. That's the practical argument for keeping prices in config rather than hardcoding a model name in six places.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I'm setting this up on client projects
&lt;/h2&gt;

&lt;p&gt;Every provider I use returns cached-token counts in the usage block of each response. I log those next to the model name and the feature that made the call. Then the pricing comparison stops being a guess, because the cached ratio comes from real traffic instead of a blog post (including this one).&lt;/p&gt;

&lt;p&gt;Routing then gets simple. Short, uncached, high-volume calls go to whichever model is cheapest on fresh input and output. Long agent sessions go to whichever model is cheapest on cached input at acceptable quality. Anything open-ended gets a hard output cap derived from a per-call budget. It's the same setup I build into the AI automation work on &lt;a href="https://abrarqasim.com" rel="noopener noreferrer"&gt;my portfolio&lt;/a&gt;, and the price table is the only part that changes when a vendor ships something new.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do this week
&lt;/h2&gt;

&lt;p&gt;Pull the last seven days of usage logs from your LLM provider. Divide cached input tokens by total input tokens, per feature, and write that ratio down. Plug your real averages into the &lt;code&gt;call_cost&lt;/code&gt; function above and re-rank the models for each feature. Then set a &lt;code&gt;max_tokens&lt;/code&gt; value on every call that doesn't have one, picked from a dollar budget. If the ranking doesn't change, you've lost twenty minutes. If it does, you've probably found money.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/llm-pricing-comparison-cached-token-math-opus-5-5-gpt-6-grok/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llmpricing</category>
      <category>llmapi</category>
      <category>promptcaching</category>
      <category>claudeopus55</category>
    </item>
    <item>
      <title>GitHub Actions Secrets Leaked Through a Cache: The Miri Lesson</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Sat, 26 Sep 2026 02:42:16 +0000</pubDate>
      <link>https://dev.to/abyzgenic/github-actions-secrets-leaked-through-a-cache-the-miri-lesson-4de5</link>
      <guid>https://dev.to/abyzgenic/github-actions-secrets-leaked-through-a-cache-the-miri-lesson-4de5</guid>
      <description>&lt;p&gt;Short version for the impatient: if a CI job caches &lt;code&gt;target/&lt;/code&gt; and that same job can see a secret as an environment variable, assume the secret can end up in the cache. The Rust Security Response Team just published a concrete case of this with Miri, and the fix they recommend is mostly about how you write your workflow file.&lt;/p&gt;

&lt;p&gt;I'll admit my first reaction to the headline was a bit smug. I don't run Miri in CI on most client projects, so I figured this one wasn't mine. Then I read the threat model section of the advisory and the smugness wore off. The line that got me is the warning that many tools "assume the entire environment can be written to the filesystem." That isn't a Miri quirk. That's a fair description of half the build tooling I've ever touched.&lt;/p&gt;

&lt;p&gt;So this post is less about Miri and more about the habit that made the Miri issue exploitable: putting GitHub Actions secrets in a top-level &lt;code&gt;env:&lt;/code&gt; block because it's convenient, and then caching build output from the same job.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually leaked, in plain terms
&lt;/h2&gt;

&lt;p&gt;Here's the chain from the &lt;a href="https://blog.rust-lang.org/2026/09/21/github-actions-leaking-secrets-when-miri-output-is-cached/" rel="noopener noreferrer"&gt;Rust blog advisory&lt;/a&gt;, boiled down.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;cargo miri&lt;/code&gt; invokes Miri several times per build, and it needs to remember build-relevant environment variables between those runs. The code that did this stored every environment variable into &lt;code&gt;target/&lt;/code&gt;. Every one. If your job had &lt;code&gt;AWS_SECRET_ACCESS_KEY&lt;/code&gt; or a deploy token in its environment, it went into a file under &lt;code&gt;target/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;On its own, that's a file on a throwaway runner. The trouble is caching. Rust projects cache &lt;code&gt;target/&lt;/code&gt; heavily because cold builds are slow, usually with &lt;code&gt;actions/cache&lt;/code&gt; or &lt;a href="https://github.com/Swatinem/rust-cache" rel="noopener noreferrer"&gt;Swatinem/rust-cache&lt;/a&gt;. A typical setup lets runs on &lt;code&gt;main&lt;/code&gt; write the cache and lets pull requests only read it. That read-only rule stops cache poisoning. It does nothing to stop a PR from reading what &lt;code&gt;main&lt;/code&gt; wrote.&lt;/p&gt;

&lt;p&gt;And who can run a PR workflow? GitHub asks a maintainer to approve CI for a first-time contributor. After someone has landed one change, their later PRs run CI on every push. The advisory spells out the nasty part: a past contributor could push a commit that dumps the cached &lt;code&gt;target/&lt;/code&gt;, grab the secret from the job output, then push a second commit over the first. GitHub sometimes hides overwritten commits in the UI, and run logs get deleted after a few months.&lt;/p&gt;

&lt;p&gt;The team scanned public GitHub repositories and found one that was vulnerable, plus seven that didn't look vulnerable but were close enough to get a heads-up. Small numbers. I don't read that as "rare", though. The scan could only see public workflow files, and the post itself says the scan was probably imperfect.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short-term fix, and why I wouldn't lean on it
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://github.com/rust-lang/miri/pull/5337" rel="noopener noreferrer"&gt;Miri patch&lt;/a&gt; narrows what gets saved: only &lt;code&gt;CARGO_*&lt;/code&gt; variables, minus anything matching &lt;code&gt;CARGO_*_TOKEN&lt;/code&gt;, plus &lt;code&gt;OUT_DIR&lt;/code&gt;. The nightly dated 2026-09-22 is the first one that carries it. If you pin an older nightly in your toolchain file, you don't have the fix yet.&lt;/p&gt;

&lt;p&gt;It's a good patch. But the advisory is blunt that Cargo, Miri and Rust in general make no promise to keep environment variables out of &lt;code&gt;target/&lt;/code&gt;. Build scripts are arbitrary code. Any &lt;code&gt;build.rs&lt;/code&gt; in your dependency tree can read the environment and write whatever it likes into &lt;code&gt;OUT_DIR&lt;/code&gt;. Almost none do. You're still trusting every crate author never to do it, including the one who adds a debug dump in a patch release at 2am.&lt;/p&gt;

&lt;p&gt;So I treat the Miri change as closing one known hole. The workflow change is the real fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  The workflow pattern behind the leak
&lt;/h2&gt;

&lt;p&gt;This is roughly what I see in a lot of Rust repos, and in a couple of my own older ones too. Secrets live at the top because some step near the end needs them, and the cache step sits at the start of the same job.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# before: every step in the job can see both secrets&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ci&lt;/span&gt;
&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;push&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;pull_request&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

&lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;SENTRY_AUTH_TOKEN&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.SENTRY_AUTH_TOKEN }}&lt;/span&gt;
  &lt;span class="na"&gt;CRATES_IO_TOKEN&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.CRATES_IO_TOKEN }}&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;dtolnay/rust-toolchain@nightly&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;components&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;miri&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Swatinem/rust-cache@v2&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cargo test&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cargo miri test&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;./scripts/upload-sourcemaps.sh&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing in that file looks wrong at a glance, which is exactly why it spreads. The workflow-level &lt;code&gt;env:&lt;/code&gt; block means &lt;code&gt;cargo miri test&lt;/code&gt; inherits both tokens, and the rust-cache step saves &lt;code&gt;target/&lt;/code&gt; when the job finishes. GitHub's own page on &lt;a href="https://docs.github.com/en/actions/how-tos/write-workflows/choose-what-workflows-do/use-secrets" rel="noopener noreferrer"&gt;using secrets in a workflow&lt;/a&gt; passes secrets per step in its examples. I think a lot of us skim past that detail because the top-level block is less typing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd change: split jobs, scope secrets to steps
&lt;/h2&gt;

&lt;p&gt;Two rules. A job that writes a shared cache gets no secrets at all. A step that needs a secret gets it in its own &lt;code&gt;env:&lt;/code&gt;, and no other step sees it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# after: the cached job has no secrets&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ci&lt;/span&gt;
&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;push&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;pull_request&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;dtolnay/rust-toolchain@nightly&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;components&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;miri&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Swatinem/rust-cache@v2&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cargo test&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cargo miri test&lt;/span&gt;

  &lt;span class="na"&gt;release-artifacts&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;needs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;test&lt;/span&gt;
    &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;github.ref == 'refs/heads/main'&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;./scripts/upload-sourcemaps.sh&lt;/span&gt;
        &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;SENTRY_AUTH_TOKEN&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.SENTRY_AUTH_TOKEN }}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second job never touches the Rust cache, only runs on &lt;code&gt;main&lt;/code&gt;, and hands one secret to one step. It'll be slower if it has to compile anything. I'd take that trade every time.&lt;/p&gt;

&lt;p&gt;If you can't split the job today, the advisory lists quicker options: turn off caching for that job, move the secrets onto steps that don't call Miri, or switch Miri off for a while. Any of those is fine as a stopgap while you do the proper split.&lt;/p&gt;

&lt;p&gt;There's one sneaky path the advisory mentions almost in passing: a previous step can persist a secret into the environment. The usual way is &lt;code&gt;echo "TOKEN=..." &amp;gt;&amp;gt; "$GITHUB_ENV"&lt;/code&gt;. Once that runs, every later step in the job has the value, including the one that runs Miri. So I'd grep for &lt;code&gt;GITHUB_ENV&lt;/code&gt; as part of this audit.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# quick audit across a repo's workflows&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rnE&lt;/span&gt; &lt;span class="s1"&gt;'secrets\.|GITHUB_ENV|rust-cache|actions/cache'&lt;/span&gt; .github/workflows/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one command shows where secrets enter, where they get promoted to job-wide env, and which jobs cache. Cross-reference the three and you've found your risky jobs.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you think you were exposed
&lt;/h2&gt;

&lt;p&gt;Fixing the YAML doesn't remove what's already sitting in the cache. The Rust team's order is: fix the workflow, clear the cache, then consider rotating anything that might have leaked. GitHub's docs on &lt;a href="https://docs.github.com/en/actions/how-tos/manage-workflow-runs/manage-caches" rel="noopener noreferrer"&gt;managing caches&lt;/a&gt; cover deleting entries from the UI or with the &lt;code&gt;gh&lt;/code&gt; CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gh cache list &lt;span class="nt"&gt;--limit&lt;/span&gt; 100
gh cache delete &lt;span class="nt"&gt;--all&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On rotation, I'd drop the word "consider". If a token sat in a cache that PRs could read, and you can't prove nobody read it, rotate it. Rotating a Sentry token takes five minutes. Explaining to a client why their crates.io account published a strange version takes a lot longer.&lt;/p&gt;

&lt;p&gt;I'm less sure what to say about the logs. Old CI logs expire, so you may not be able to confirm or rule out access at all. I don't have a clever answer for that one. Rotate and move on.&lt;/p&gt;

&lt;h2&gt;
  
  
  This isn't really a Rust problem
&lt;/h2&gt;

&lt;p&gt;I keep coming back to this. Miri got caught because someone went looking (Predrag Gruevski reported it, and the team credits him in the post). The same pattern can show up anywhere a tool writes its environment into an output directory and that directory gets cached. I'd look hard at anything that snapshots env for reproducible builds, and at Docker layer caches built with secrets passed as build args. I haven't checked specific tools for this exact behavior, and I won't pretend I have. The point is that you shouldn't need to check, because the job writing the cache shouldn't hold anything worth stealing.&lt;/p&gt;

&lt;p&gt;I wrote about a related habit in &lt;a href="https://abrarqasim.com/blog/github-actions-reusable-workflows-the-bug-i-fixed-eleven-times/" rel="noopener noreferrer"&gt;the reusable workflows bug I fixed eleven times&lt;/a&gt;, where copy-pasted workflow blocks kept bringing back the same mistake. Secrets in a top-level &lt;code&gt;env:&lt;/code&gt; spread the same way. Someone adds one for a single step, the next person copies the file into a new repo, and a year later every job in the org can see every token.&lt;/p&gt;

&lt;p&gt;This is one of the first things I check when I take over a client's CI, and it's part of the DevOps cleanup work you'll find on &lt;a href="https://abrarqasim.com" rel="noopener noreferrer"&gt;my portfolio&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do this week
&lt;/h2&gt;

&lt;p&gt;Open &lt;code&gt;.github/workflows/&lt;/code&gt; in your busiest repo and run the grep above. For every job with a cache step, confirm it has no &lt;code&gt;secrets.&lt;/code&gt; reference and no &lt;code&gt;GITHUB_ENV&lt;/code&gt; write that carries a secret. When you find one, move the secret to the single step that needs it, or into a separate job that doesn't cache. Then run &lt;code&gt;gh cache delete --all&lt;/code&gt; and rotate the token. It's maybe half an hour of work, and I'd bet you find at least one job you'd forgotten existed.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/github-actions-secrets-miri-cache-leak-the-env-block-at-the-top/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>githubactions</category>
      <category>githubactionssecrets</category>
      <category>rust</category>
      <category>miri</category>
    </item>
    <item>
      <title>How to Export a Trello Board to Excel, CSV, PDF or Sheets</title>
      <dc:creator>Qasim Parray</dc:creator>
      <pubDate>Fri, 25 Sep 2026 22:42:14 +0000</pubDate>
      <link>https://dev.to/abyzgenic/how-to-export-a-trello-board-to-excel-csv-pdf-or-sheets-3n04</link>
      <guid>https://dev.to/abyzgenic/how-to-export-a-trello-board-to-excel-csv-pdf-or-sheets-3n04</guid>
      <description>&lt;p&gt;Someone asked me a few weeks ago whether they could get their Trello board into a spreadsheet with the card comments attached. I said yes, because I maintain a Trello export Power-Up and I assumed I had built that years ago. Then I went and checked. I hadn't. The Power-Up SDK that board add-ons run against does not hand you comments at all, so the answer was a week of work against Trello's REST API: a token flow, a rate limiter, twenty pages of card actions, and a stack of edge cases I did not want to meet.&lt;/p&gt;

&lt;p&gt;So here is the post I wanted when I started. What Trello exports by itself, what it quietly drops, and what your options are when the built-in download is not enough. The first two thirds apply no matter what tool you end up using. I talk about my own Power-Up at the end, because I built it to fill these exact gaps, and you can judge whether that is fair.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Trello exports on its own
&lt;/h2&gt;

&lt;p&gt;Open any board, hit the board menu, and pick "Print, Export, and Share". You get two formats.&lt;/p&gt;

&lt;p&gt;JSON is available to every board member on every plan, including free. It is the complete record of the board: cards, lists, labels, members, checklists, custom fields, and the board's most recent actions. Comments live inside those actions. &lt;a href="https://support.atlassian.com/trello/docs/exporting-data-from-trello/" rel="noopener noreferrer"&gt;Atlassian's export documentation&lt;/a&gt; puts the ceiling at the 1,000 most recent actions per board. On an active board you will hit that number, and whatever falls off the end is simply not in your file. The docs also say the JSON cannot be re-imported to rebuild a board, which surprises people who assume export means backup.&lt;/p&gt;

&lt;p&gt;CSV needs a Premium workspace. It is the format most people actually want, because it opens in Excel or Google Sheets with nobody writing code. Two limits come straight from Atlassian's own page: CSV export only includes cards that are not archived, and comments are excluded from spreadsheet exports.&lt;/p&gt;

&lt;p&gt;There is also a workspace-level export, admin only, Premium only, which can bundle raw attachment files into a ZIP instead of just linking to them. If you are leaving Trello or want a cold archive of everything, that is the one to reach for.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the built-in CSV leaves out
&lt;/h2&gt;

&lt;p&gt;I keep a list of the things people ask for that the native CSV does not carry. It has stayed stable for about two years.&lt;/p&gt;

&lt;p&gt;Archived cards. Teams archive instead of delete, so on a year-old board half the history is invisible to the export. Every "why does my report show 40 cards when we shipped 300" question traces back to this.&lt;/p&gt;

&lt;p&gt;Comments. The decision log of most boards lives in comments, not descriptions. They are in the JSON and not in the CSV, which is the single most annoying split in the whole product.&lt;/p&gt;

&lt;p&gt;Movement history. When did this card enter the list it is in now? How many days has it been sitting there? Who moved it to Done? None of that is a card field. It has to be derived from the board's action feed, which you reach through the &lt;a href="https://developer.atlassian.com/cloud/trello/rest/" rel="noopener noreferrer"&gt;REST API&lt;/a&gt;. No menu will give it to you.&lt;/p&gt;

&lt;p&gt;Checklist items as rows. A CSV gives you one row per card. If you run a board where each card holds a twelve-item checklist, and you want one row per item, you are writing a script.&lt;/p&gt;

&lt;p&gt;Formatting. Not a data problem, but a real one. A raw CSV has no column widths, no frozen header, no label colours, no filters, and no totals. Anyone building a weekly status report ends up doing the same twenty minutes of spreadsheet work every Friday.&lt;/p&gt;

&lt;h2&gt;
  
  
  Exporting a Trello board to Excel
&lt;/h2&gt;

&lt;p&gt;Four routes, roughly in order of effort.&lt;/p&gt;

&lt;p&gt;Copy and paste works for one list on one board on one afternoon. It stops working the moment somebody wants it again next week.&lt;/p&gt;

&lt;p&gt;The JSON route is the free one. Download the JSON, then parse it with Python, Power Query, or whatever you already know. This is genuinely fine for a technical user with a one-off question, and it is the only free path to comments. The cost is that you are now maintaining a script, and Trello's JSON shape is nested enough that flattening it well takes longer than you think.&lt;/p&gt;

&lt;p&gt;Browser extensions read the board through the page you are looking at. A few have been around for years and do a decent job. The tradeoff is that an extension sees whatever your browser session can see, updates on its own schedule, and cannot run when your laptop is shut.&lt;/p&gt;

&lt;p&gt;Power-Ups run inside Trello itself. You enable one from the board's Power-Ups menu, it appears as a button on the board, and it uses Trello's own APIs with the permissions you grant. This is the route I picked when I built mine, mostly because nothing has to be installed on the person's machine and the free tier of Trello supports Power-Ups.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where my Power-Up fits
&lt;/h2&gt;

&lt;p&gt;I run &lt;a href="https://exportly.abrarqasim.com" rel="noopener noreferrer"&gt;Exportly&lt;/a&gt;, a Trello Power-Up for exporting boards. It is the successor to an earlier export Power-Up of mine that has been sitting on a thousand-odd boards for a while, and the rewrite exists because of the gap list above. You can install it from &lt;a href="https://trello.com/power-ups/6aaef31b1b790e7087a6a484" rel="noopener noreferrer"&gt;Trello's Power-Up directory&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The short version of what it does. Excel workbooks that are already formatted, with a frozen header, coloured labels, dates that Excel treats as dates, and an optional Summary sheet with live COUNTIF formulas that keep working after you edit the file. CSV with a delimiter and date format that match your locale, which matters more than it sounds if you are in Europe and your spreadsheet keeps reading dates as text. PDF for the people who want to print a board or email a client something that is not editable. Google Sheets, pushed straight into your Drive.&lt;/p&gt;

&lt;p&gt;On the data side it covers what the native CSV does not: archived lists and cards, comment text, the date a card entered its current list, days in list, who moved it to Done, and a checklist-item row mode where every item becomes its own row. Filters for list, label, member, and due window run before the export, so a "what is overdue this sprint" report becomes a saved preset you click once. You can export up to ten boards at once into one workbook.&lt;/p&gt;

&lt;p&gt;The comment and history columns need you to connect Trello once, because they come from the REST API rather than the Power-Up SDK. That connection is the part that took the week I mentioned at the top, and it is also the part I would have skipped if I had known how fiddly Trello's authorisation window is inside an iframe.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scheduling the export so you stop doing it by hand
&lt;/h2&gt;

&lt;p&gt;This is the feature I use myself and the reason the product has a server at all.&lt;/p&gt;

&lt;p&gt;You pick a board, a format, a time, and a timezone, and the export arrives by email on a schedule. Monday 08:00 in Asia/Dubai, every weekday, month-end, whatever the cron expression allows. Files above eight megabytes come as a download link instead of an attachment, and links expire after seven days.&lt;/p&gt;

&lt;p&gt;The interesting part of building it was storage. A scheduled export needs a Trello token sitting on a server, which is a thing I did not want to be careless about, so tokens are sealed with AES-256-GCM and only ever decrypted inside the worker process that runs the exports. The API process cannot read them. It runs on a small Hetzner box, the same one I wrote about when I compared &lt;a href="https://abrarqasim.com/blog/hetzner-vs-digitalocean-what-this-blog-actually-runs-on" rel="noopener noreferrer"&gt;Hetzner and DigitalOcean for this blog&lt;/a&gt;, which has turned out to be more than enough machine for the job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Backing up a board with its attachments
&lt;/h2&gt;

&lt;p&gt;Trello's workspace export can include raw attachments, but it is Premium and admin only, and it gives you everything at once rather than the one board you care about.&lt;/p&gt;

&lt;p&gt;Exportly has a board backup that produces a ZIP: the export file itself, every uploaded attachment sorted into folders by list and card, and a manifest CSV that records which files were included, which were skipped, and why. Link-only attachments cannot be fetched, files over fifty megabytes are skipped by design, and the manifest says so, so nobody has to guess. I added the manifest after my own first test run quietly dropped four files and I spent an hour working out which ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this costs and what I would pick
&lt;/h2&gt;

&lt;p&gt;Free tier is three exports a month, which is enough to decide whether the thing is useful. Then it is a fourteen day trial, and after that 6.99 US dollars a month or 49 a year. Refunds are fourteen days after any charge, no argument.&lt;/p&gt;

&lt;p&gt;If you are on Trello Premium, you export once a quarter, and you do not care about comments or archived cards, the built-in CSV is fine. Use it and skip the rest of this.&lt;/p&gt;

&lt;p&gt;If you are on the free plan, or you need comments, archived cards, formatted Excel, PDF, or a report that shows up in your inbox every Monday without you doing anything, a Power-Up is the shortest path. You can &lt;a href="https://trello.com/power-ups/6aaef31b1b790e7087a6a484" rel="noopener noreferrer"&gt;add Exportly to a board&lt;/a&gt;, run a free export, and see whether the workbook is the one you wanted before you pay for anything.&lt;/p&gt;

&lt;p&gt;Here is the one thing to do this week, whichever way you go. Open your busiest board, run the native JSON export, and search the file for a comment you remember writing three months ago. If it is missing, you have just found the 1,000-action ceiling, and you now know why the spreadsheet you were about to build would have been wrong. I write about this kind of small, expensive detail fairly often over on &lt;a href="https://abrarqasim.com" rel="noopener noreferrer"&gt;my site&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abrarqasim.com/blog/how-to-export-a-trello-board-excel-csv-pdf-sheets/" rel="noopener noreferrer"&gt;abrarqasim.com&lt;/a&gt;. I write there about React, PHP, Rust, Go and the AI tooling around them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>trello</category>
      <category>export</category>
      <category>excel</category>
      <category>csv</category>
    </item>
  </channel>
</rss>
