<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Billie M</title>
    <description>The latest articles on DEV Community by Billie M (@billiem).</description>
    <link>https://dev.to/billiem</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4026133%2F5dd55f62-43ed-45e3-b3bc-98692f37a24d.gif</url>
      <title>DEV Community: Billie M</title>
      <link>https://dev.to/billiem</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/billiem"/>
    <language>en</language>
    <item>
      <title>What GPT-6.1 Sol chose to build</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Thu, 01 Oct 2026 12:38:32 +0000</pubDate>
      <link>https://dev.to/billiem/what-gpt-61-sol-chose-to-build-495a</link>
      <guid>https://dev.to/billiem/what-gpt-61-sol-chose-to-build-495a</guid>
      <description>&lt;p&gt;A cutting layout can look convincing while leaving its cut order ambiguous. A document comparison can work while putting its navigation offscreen. A traffic solver can return the right numbers without making them understandable.&lt;/p&gt;

&lt;p&gt;Those were concrete review problems in three autonomous GPT-6.1 Sol experiments. The subjects were freely chosen: sheet cutting, the editing of an official record, and congestion. What changed during review was how someone could inspect and use the result.&lt;/p&gt;

&lt;p&gt;I wanted to continue the &lt;a href="https://billiem.uk/posts/llm-choice-as-an-idea-engine/" rel="noopener noreferrer"&gt;Choice exercise&lt;/a&gt;: let a new model choose what to build, then examine its outputs and style. This batch ran at Extra High reasoning. The coordinating agent assigned compact, substantial and flagship ambition levels. One fresh builder chose all three topics and implemented the first two; a second fresh Sol builder took the flagship before public code existed.&lt;/p&gt;

&lt;h2&gt;
  
  
  A packing diagram needs a cutting sequence
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://billiem.uk/choice/experiments/2026-10-01-sol61-xhigh-substantial/outputs/offcut" rel="noopener noreferrer"&gt;Try Offcut&lt;/a&gt; with sheet dimensions, part quantities, grain constraints and blade width. It produces a layout and a sequence of full cuts through the remaining rectangles. Playback shows when a piece is actually released, rather than merely showing where it ends up.&lt;/p&gt;

&lt;p&gt;The supplied job places nine parts on a 1,220 × 610 mm sheet. The blade's lost material changes feasibility: two 50 × 100 mm parts fit a 100 × 100 mm sheet with zero blade width, but only one fits with a 3 mm blade.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgroqspk0z7xh1jarfcj6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgroqspk0z7xh1jarfcj6.png" alt="A 1,220 by 610 mm sheet contains three labelled shelves, two sides, smaller pieces and hatched remainder areas, with grain running horizontally." width="768" height="460"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Offcut's supplied nine-part layout. The diagram accounts for blade loss and grain; no physical cutting was tested.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The review problem was in the CSV cut order. Two physical copies of identically named stock were indistinguishable. Stable physical sheet IDs now distinguish them. Mobile feedback also changed because solving could leave the result offscreen.&lt;/p&gt;

&lt;p&gt;The outputs include a dimensioned SVG, CSV instructions, editable JSON and remaining-stock inventory. Their usefulness still depends on the model's bounds: axis-aligned rectangular parts and full guillotine cuts. Damaged edges, trimming allowance, clamping and tool clearance are excluded, and no physical cutting was tested. An unplaced part is not proof that every possible arrangement fails; this is a bounded search.&lt;/p&gt;

&lt;h2&gt;
  
  
  A text comparison needs a way back to its source
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://billiem.uk/choice/experiments/2026-10-01-sol61-xhigh-jewel/outputs/minutes" rel="noopener noreferrer"&gt;Explore The Minutes&lt;/a&gt; by selecting sentences across six authored versions of a fictional exhibition incident. It exposes earlier wording and lets you recover facts omitted from the current record.&lt;/p&gt;

&lt;p&gt;“The issue was managed promptly” traces back to a request to isolate the power, refused because the display lighting had to stay on. The final reassurance is not established by that first account. The sequence also contains a useful correction: an uncertain pump-stoppage time becomes a time supported by the controller log. Editing is not uniformly treated as concealment.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8i6wfqkti7xvtgjaq58z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8i6wfqkti7xvtgjaq58z.png" alt="The earlier isolation request and refusal are struck through above the introduced reassurance, ‘The issue was managed promptly.’" width="349" height="130"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The final record compared with the first account in The Minutes. This ancestry belongs to an explicitly authored fictional incident.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The initial navigation began below the visible page. Review moved it above the document; on mobile, selecting a sentence brings its provenance into view and supplies a return to the sentence. The explanation distinguishes removed facts from introduced assurances. The corpus and its ancestry are authored, so this is an inspectable argument, not a dishonesty detector for uploaded documents.&lt;/p&gt;

&lt;h2&gt;
  
  
  A solver needs more than its raw output
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://billiem.uk/choice/experiments/2026-10-01-sol61-xhigh-flagship/outputs/shortcut" rel="noopener noreferrer"&gt;Change the network in The Shortcut Tax&lt;/a&gt;. At the default west demand of six flow units, average travel is 80 minutes without the shortcut and 100 with it under individual route choice. No driver can save time by switching routes alone.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4my4f2krrtgbs689rpmp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4my4f2krrtgbs689rpmp.png" alt="The open shortcut carries four flow units, and average travel is 100 minutes beside an 80-minute outer-roads comparison. Individual route choice is selected." width="388" height="930"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The default six-unit demand with the shortcut open. These are results of the teaching network, not measured traffic times.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Coordinated routing restores the 80-minute average. Pricing each driver's imposed delay can reproduce it too, with tolls expressed in equivalent minutes. Changing demand can make the same shortcut helpful, harmful or unused.&lt;/p&gt;

&lt;p&gt;The prototype initially exposed calculations as raw JSON. Review asked for a legible network, road and route costs, a demand sweep and a toll ledger. A second origin makes the shared bottleneck's effects inspectable by neighbourhood. These are teaching results under fixed demand, instantaneous equilibrium and simple congestion costs, not measured traffic or a forecast with queues.&lt;/p&gt;

&lt;p&gt;The published surfaces still resemble &lt;a href="https://billiem.uk/posts/gpt-6-astra-choice/" rel="noopener noreferrer"&gt;September's Astra batch&lt;/a&gt; in their warm paper, serif headings and muted colours. The working conditions differ: Astra used three separately assigned builders at Max, with different subjects and a different brief. These portfolios cannot establish a model's taste or superiority.&lt;/p&gt;

&lt;p&gt;The concrete comparison is between the prototype and what a visitor can now do. Follow a sentence back, distinguish physical sheets in a kept cut order, or inspect why the network's travel time changes. Review altered those routes into the underlying work, not just its finish.&lt;/p&gt;




&lt;p&gt;Want to talk about something I’ve written or built? &lt;a href="https://billiem.uk/contact/" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/gpt-6-1-sol-choice/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Free compute in October 2026: new backends and temporary VMs</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Thu, 01 Oct 2026 11:42:01 +0000</pubDate>
      <link>https://dev.to/billiem/free-compute-in-october-2026-new-backends-and-temporary-vms-2mil</link>
      <guid>https://dev.to/billiem/free-compute-in-october-2026-new-backends-and-temporary-vms-2mil</guid>
      <description>&lt;p&gt;Two of October’s useful development routes let you postpone signup: a temporary Postgres project and a Linux machine reached over SSH. The database lasts three days; the machine gives you an hour to build. The claim links and account handoffs have their own limits.&lt;/p&gt;

&lt;p&gt;This is the &lt;strong&gt;1 October 2026&lt;/strong&gt; check, based on official pricing and documentation. No accounts or workloads were created for it. The &lt;a href="https://billiem.uk/free-compute/" rel="noopener noreferrer"&gt;maintained free-compute guide&lt;/a&gt; carries the broader shortlist, and the &lt;a href="https://billiem.uk/posts/free-compute-september-2026/" rel="noopener noreferrer"&gt;September edition&lt;/a&gt; is the previous snapshot. The focus here is getting a prototype started, then keeping its expiry and credit conditions attached to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A database before its owner signs up
&lt;/h2&gt;

&lt;p&gt;An agent can create a &lt;a href="https://neon.com/docs/reference/claimable-neon" rel="noopener noreferrer"&gt;Claimable Neon&lt;/a&gt; Postgres project before a human has an account. Neon &lt;a href="https://neon.com/docs/changelog#2026-09-11" rel="noopener noreferrer"&gt;announced that route on 11 September&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The unclaimed project gets &lt;strong&gt;100 MB storage and 1 GB transfer&lt;/strong&gt;. There are two deadlines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The project expires after &lt;strong&gt;72 hours&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Its claim code expires after &lt;strong&gt;15 minutes&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A database can still exist after its original claim code has expired. Treating those clocks as one deadline would miss the shorter one.&lt;/p&gt;

&lt;p&gt;The Data API and Auth can be requested at creation. Functions, Object Storage and AI Gateway require claiming first. That leaves a useful temporary route for generating and testing a database-backed prototype, with a clear boundary around which backend products are available before onboarding. It does not replace the permanent Free plan or a backup.&lt;/p&gt;

&lt;h2&gt;
  
  
  A shell for the development session
&lt;/h2&gt;

&lt;p&gt;Railway’s equivalent entry point is a command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh railway.new
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://railway.com/changelog/2026-09-25-free-vms-without-an-account" rel="noopener noreferrer"&gt;25 September announcement&lt;/a&gt; describes a Linux VM with &lt;strong&gt;2 vCPU and 2 GB RAM&lt;/strong&gt;, coding agents already installed, and a private preview URL. It needs neither an account nor a payment card.&lt;/p&gt;

&lt;p&gt;Here the active development window is &lt;strong&gt;60 minutes&lt;/strong&gt;, followed by &lt;strong&gt;24 hours to claim the VM&lt;/strong&gt;. Claiming transfers the machine, files and URL into an account and makes the preview shareable. Unclaimed machines and files are deleted. Save or claim work before the clock runs out.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://docs.railway.com/pricing/free-trial" rel="noopener noreferrer"&gt;trial documentation&lt;/a&gt; adds three anonymous boxes per IP address per day and regional capacity restrictions. Claiming also changes the billing context: the workload moves into the account’s ordinary credits and billing. The separate signup trial is $5 for 30 days, followed by $1 of recurring monthly Free credit. An anonymous hour and an account’s monthly allowance are different offers.&lt;/p&gt;

&lt;h2&gt;
  
  
  After claiming, read the backend meters separately
&lt;/h2&gt;

&lt;p&gt;Neon’s &lt;a href="https://neon.com/docs/changelog#2026-09-18" rel="noopener noreferrer"&gt;backend became generally available on 18 September&lt;/a&gt;. Functions and Object Storage are available in Ohio, Northern Virginia, Frankfurt and Singapore, replacing September’s Ohio-only beta description with published Free allowances.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://neon.com/pricing" rel="noopener noreferrer"&gt;Free pricing page&lt;/a&gt; lists up to 100 Postgres projects, each with &lt;strong&gt;100 CU-hours per month and 0.5 GB database storage&lt;/strong&gt;. Object Storage gets &lt;strong&gt;5 GB per project&lt;/strong&gt;. Functions get &lt;strong&gt;10 active Capacity-Hours, 400 waiting Capacity-Hours and one million monthly invocations&lt;/strong&gt;; Auth includes &lt;strong&gt;60,000 monthly active users&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Those figures cannot all be multiplied by 100. The &lt;a href="https://neon.com/docs/introduction/plans#functions" rel="noopener noreferrer"&gt;plan documentation&lt;/a&gt; applies the Functions limits across the account as well as independently per project. A project’s &lt;strong&gt;5 GB network allowance is shared&lt;/strong&gt; by Postgres, Object Storage and Functions.&lt;/p&gt;

&lt;p&gt;Active Capacity-Hours measure CPU work; waiting Capacity-Hours measure time waiting on I/O. Adding them together does not produce 410 hours of an arbitrary always-on server. AI Gateway also uses purchased prepaid credits, so the backend launch does not supply free model inference.&lt;/p&gt;

&lt;h2&gt;
  
  
  A compute credit may exclude the API you call
&lt;/h2&gt;

&lt;p&gt;For a GPU or model experiment, the product covered by the credit matters as much as its size.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Modal’s $30 is recurring compute credit.&lt;/strong&gt; Its &lt;a href="https://modal.com/pricing" rel="noopener noreferrer"&gt;current pricing FAQ&lt;/a&gt; covers CPU, GPU and memory usage in Functions, Sandboxes and Notebooks. Shared Endpoints tokens are excluded and billed from the first request, effective &lt;strong&gt;1 September 2026&lt;/strong&gt;. The &lt;a href="https://modal.com/docs/guide/budgets" rel="noopener noreferrer"&gt;Workspace spend limit&lt;/a&gt; is the cash cap; a usage budget measures usage before credits. A payment method is required, and extra compute can charge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lightning’s GPU credits are a starter balance.&lt;/strong&gt; The &lt;a href="https://lightning.ai/pricing" rel="noopener noreferrer"&gt;current offer&lt;/a&gt; gives five credits on registration and 25 more after adding a payment card. Unused credits expire after &lt;strong&gt;12 months&lt;/strong&gt;. GPU work spends that balance. One CPU Studio can remain free with a restart every four hours. September’s claim of 15 recurring monthly credits is not supported by today’s offer; no dated announcement was found to establish when it changed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mistral publishes a monthly API allowance.&lt;/strong&gt; Its &lt;a href="https://mistral.ai/pricing/" rel="noopener noreferrer"&gt;Free pricing&lt;/a&gt; now specifies &lt;strong&gt;$10 per month in API credits&lt;/strong&gt;. The &lt;a href="https://docs.mistral.ai/admin/billing-usage/subscriptions" rel="noopener noreferrer"&gt;subscription documentation&lt;/a&gt; shares usage across Studio, API and Vibe Code. With pay-as-you-go disabled, exhausted usage can stop until renewal; enabling it allows additional tokens to bill. Model access and rate limits still need checking for the account. This is a current pricing observation, with no established launch date.&lt;/p&gt;

&lt;h2&gt;
  
  
  Batch work can wait; VM promotions can end
&lt;/h2&gt;

&lt;p&gt;A prototype’s non-urgent processing has another dated option. Cloud Run added &lt;a href="https://docs.cloud.google.com/run/docs/release-notes" rel="noopener noreferrer"&gt;delayed job execution in Preview on 8 September&lt;/a&gt;: provisioning can wait up to &lt;strong&gt;12 hours&lt;/strong&gt; for a lower price, and execution can then run for up to 12 hours.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://cloud.google.com/run/pricing" rel="noopener noreferrer"&gt;pricing table&lt;/a&gt; gives a monthly free equivalent of &lt;strong&gt;342,857 vCPU-seconds and 642,857 GiB-seconds&lt;/strong&gt; at Tier 1 pricing. Google applies the free tier as a spending discount aggregated by billing account. The delayed-job, ordinary-job and request-based rows are not independent grants to add together. A billing-enabled account is involved, and builds, stored images, networking and excess usage can charge. The &lt;a href="https://docs.cloud.google.com/run/docs/delayed-jobs" rel="noopener noreferrer"&gt;scheduling documentation&lt;/a&gt; carries the execution limits.&lt;/p&gt;

&lt;p&gt;For a longer-running ARM VM, the &lt;a href="https://aws.amazon.com/ec2/instance-types/t4/" rel="noopener noreferrer"&gt;AWS T4g promotion&lt;/a&gt; still provides &lt;strong&gt;750 aggregate t4g.small hours per month&lt;/strong&gt;, ending &lt;strong&gt;31 December 2026&lt;/strong&gt;. Disks, public IPv4 and surplus CPU can charge. Its end date needs to travel with the instance-hours figure.&lt;/p&gt;

&lt;p&gt;The temporary database and SSH machine give a prototype somewhere to start. Their short claim windows remain relevant even while the resources still exist. Once work moves into an account, the covered product, shared meters and billing settings determine what the headline allowance actually buys.&lt;/p&gt;




&lt;p&gt;Want to talk about something I’ve written or built? &lt;a href="https://billiem.uk/contact/" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/free-compute-october-2026/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>webdev</category>
      <category>devops</category>
      <category>ai</category>
    </item>
    <item>
      <title>What GPT-6 Astra chose to build</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Sun, 13 Sep 2026 21:44:20 +0000</pubDate>
      <link>https://dev.to/billiem/what-gpt-6-astra-chose-to-build-5e9a</link>
      <guid>https://dev.to/billiem/what-gpt-6-astra-chose-to-build-5e9a</guid>
      <description>&lt;p&gt;The first prototype of Loose End had a string, three pegs and four bells. It also had arrow-key controls that ignored short presses. A coordinating agent found the input problem and sent it back to the builder before the game was finished.&lt;/p&gt;

&lt;p&gt;That repair is part of GPT-6 Astra's first &lt;a href="https://billiem.uk/choice/" rel="noopener noreferrer"&gt;Choice batch&lt;/a&gt;. I asked for three independent experiments with room to choose their subjects. Three Astra agents ran at max reasoning, assigned something to play, explore or use. They produced a string puzzle, a supply-chain model and a tool for making a printed booklet, published on 4 September.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://billiem.uk/posts/gpt-5-6-needed-a-more-ambitious-prompt/" rel="noopener noreferrer"&gt;existing brief&lt;/a&gt; required prototype review before polish. The &lt;a href="https://github.com/BillieM/billiemuk/tree/7c4dd81dcffaf70557d90ef45f5f5be756a5469d/surfaces/choice/batches/2026-09-04-astra" rel="noopener noreferrer"&gt;batch's curation record&lt;/a&gt; shows what the reviewer asked the builders to change. These were separate projects with corrections, not matched trials against another model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Short key taps did nothing
&lt;/h2&gt;

&lt;p&gt;In &lt;a href="https://billiem.uk/choice/experiments/2026-09-04-astra-play/outputs/loose-end" rel="noopener noreferrer"&gt;Loose End&lt;/a&gt;, you move the brass end of a finite string around the board. It catches around pegs, so a route that reaches a bell can leave less string available for the next one. Ringing all four bells is not enough: the string has to come home too.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbbhzfcdx0146bs756n9z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbbhzfcdx0146bs756n9z.png" alt="Bell I is marked as reached, and a red string bends around a peg between home and the brass end." width="768" height="505"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A played position in Loose End. The route taken around the pegs changes how much string remains; this capture does not show a completed game.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The input fix made brief keypresses work. The later checks went further than reaching one bell: they exercised two winning strategies and retraced a long move sequence back to its starting point. Those checks concern the game's route rules and unwinding. They do not turn its planar taut-string model into a simulation of friction or elastic rope. It remains one authored course.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put both order histories on the same scale
&lt;/h2&gt;

&lt;p&gt;The review request for &lt;a href="https://billiem.uk/choice/experiments/2026-09-04-astra-explore/outputs/the-order-echo" rel="noopener noreferrer"&gt;The Order Echo&lt;/a&gt; concerned how to see a comparison. Its shop, wholesaler, depot and factory could react to the same demand under different ordering rules, but the reviewer wanted their traces visible on a shared scale, with the ingredients of each order exposed.&lt;/p&gt;

&lt;p&gt;The finished interface lets you select a week and stage to inspect its arithmetic. One busy shopping week, twelve crates instead of four, drives the starting chain's factory orders to a peak of 76.8 crates in week twelve.&lt;/p&gt;

&lt;p&gt;Counting outstanding orders changes that result. The shop initially orders more, 14.5 rather than 12.5 crates: the higher demand forecast raises its desired amount on order, and the gap to that target adds a correction. As the simulation continues, the policy counts orders awaiting delivery, including supplier backlog. The factory peak drops to 38.1 in week nine.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9denbn29hvptvatxonb1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9denbn29hvptvatxonb1.png" alt="The Order Echo compares a factory peak of 38.1 in week 9 with 76.8 in week 12, with both chains on the same scale." width="768" height="496"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Discovery 02 counts crates already ordered. The dotted traces retain the starting policy; these values are results of this teaching model.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The dotted starting-policy traces remain beside the changed run. The result and its timing can be compared directly. All of these quantities belong to the model's chosen rules, continuous crate quantities and 48-week window, with unlimited production upstream of the factory. They are not business forecasts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The saved file and the visible sheet
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://billiem.uk/choice/experiments/2026-09-04-astra-use/outputs/pocket-press" rel="noopener noreferrer"&gt;Pocket Press&lt;/a&gt; takes a cover, manuscript and back cover and arranges an eight-page booklet on one A4 or US Letter sheet. It keeps a reading view alongside the print arrangement and can save printable output, an SVG sheet and editable source.&lt;/p&gt;

&lt;p&gt;The reviewer asked for discoverable save actions, folding instructions and a way to reopen that source. The recorded browser checks downloaded the actual files, read back their geometry and text, and restored the manuscript from its saved source. Overflow blocks printable exports while leaving source saving available.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx2o1byntr4sl0d48r3zb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx2o1byntr4sl0d48r3zb.png" alt="Pocket Press shows the sample manuscript beside an A4 sheet with pages 5, 4, 3, 2 upside down above pages 6, 7, 8, 1." width="768" height="762"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The supplied sample in Pocket Press's print-sheet view at a 900-pixel browser width. The unusual page order is for folding a single sheet into an eight-page book.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The print arrangement is deliberately unfamiliar: pages 5, 4, 3 and 2 upside down above 6, 7, 8 and 1. It is meant to become a book after cutting the central slit and folding the sheet.&lt;/p&gt;

&lt;p&gt;Capturing the article's images found a different problem. At a 1,280-pixel browser width on 9 September, the print sheet extended beyond its preview area and was cropped at the top and bottom. The 900-pixel view shown here fits. Reading back a correctly arranged export had not tested that wider preview.&lt;/p&gt;

&lt;p&gt;The paper result is still untested. All the checks described here were performed by agents or automation; no physical printer or fold was tried. Pocket Press supplies a numbered test sheet, ready for the part of the experiment that needs a sheet of paper.&lt;/p&gt;




&lt;p&gt;Want to talk about something I’ve written or built? &lt;a href="https://billiem.uk/contact/" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/gpt-6-astra-choice/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>testing</category>
      <category>showdev</category>
    </item>
    <item>
      <title>Local models are actually good now - playing with Qwen3.8-27B</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Wed, 09 Sep 2026 20:34:26 +0000</pubDate>
      <link>https://dev.to/billiem/local-models-are-actually-good-now-playing-with-qwen38-27b-3kce</link>
      <guid>https://dev.to/billiem/local-models-are-actually-good-now-playing-with-qwen38-27b-3kce</guid>
      <description>&lt;p&gt;The 8-bit model loaded. It generated for more than half an hour. It also failed before its agent session finished.&lt;/p&gt;

&lt;p&gt;The 4-bit version completed the same experiment, ran faster, and built the thing I liked most: a surprisingly complete cellular automata workbench called &lt;a href="https://billiem.uk/local-choice/runs/2026-09-04-qwen38-27b-4bit-pi-create-004/artifact/" rel="noopener noreferrer"&gt;Lattice&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That changed the question I ask about large local models. “Can my laptop load this?” is easy to answer and not especially useful. The better question is whether the model can stay alive through a real agentic loop and produce something worth keeping.&lt;/p&gt;

&lt;p&gt;On my 48 GB M4 Pro, 4-bit Qwen3.8-27B did.&lt;/p&gt;

&lt;h2&gt;
  
  
  I tested complete agent sessions, not isolated prompts
&lt;/h2&gt;

&lt;p&gt;My first attempt at this comparison was wrong.&lt;/p&gt;

&lt;p&gt;I asked the 4-bit, 6-bit, and 8-bit models to emit an entire web artifact in one long completion. That produced numbers, but it removed the part that made our earlier Local LLM Choice experiments interesting: the model was supposed to work through an agent harness. It needed to create a file, inspect it, run it, notice problems, and revise its own work.&lt;/p&gt;

&lt;p&gt;“I feel like we’ve not done this right” was the most useful conclusion from that first pass. We withdrew it and reran the experiment properly.&lt;/p&gt;

&lt;p&gt;For the replacement, every quantisation started Pi in a genuinely empty directory. Each received the same broad Billie-domain prompt, the same sampling settings, and up to 25 Pi iterations. There were no retries and no repairs after the run. The model had to choose its own project and carry it through the same working loop.&lt;/p&gt;

&lt;p&gt;You can &lt;a href="https://billiem.uk/local-choice/qwen38-quantization-comparison" rel="noopener noreferrer"&gt;open the public comparison&lt;/a&gt; to use all three preserved artifacts and inspect their run receipts and manual evaluations.&lt;/p&gt;

&lt;h2&gt;
  
  
  The useful result is whether the agent finishes
&lt;/h2&gt;

&lt;p&gt;Here is what happened:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Quantisation&lt;/th&gt;
&lt;th&gt;Model files&lt;/th&gt;
&lt;th&gt;Aggregate completion rate&lt;/th&gt;
&lt;th&gt;Wall time&lt;/th&gt;
&lt;th&gt;Peak MLX allocation&lt;/th&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;4-bit&lt;/td&gt;
&lt;td&gt;15.1 GB&lt;/td&gt;
&lt;td&gt;12.47 tok/s&lt;/td&gt;
&lt;td&gt;30m 20s&lt;/td&gt;
&lt;td&gt;33.98 GB&lt;/td&gt;
&lt;td&gt;Completed: Lattice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6-bit&lt;/td&gt;
&lt;td&gt;21.9 GB&lt;/td&gt;
&lt;td&gt;9.73 tok/s&lt;/td&gt;
&lt;td&gt;25m 32s&lt;/td&gt;
&lt;td&gt;36.10 GB&lt;/td&gt;
&lt;td&gt;Completed: The Long Now&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8-bit&lt;/td&gt;
&lt;td&gt;28.6 GB&lt;/td&gt;
&lt;td&gt;6.88 tok/s before failure&lt;/td&gt;
&lt;td&gt;32m 10s&lt;/td&gt;
&lt;td&gt;Not available&lt;/td&gt;
&lt;td&gt;Partial: Metal OOM&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are whole-session measurements, not raw decode benchmarks. Pi’s working history grew through the run, and the aggregate completion rate includes prompt processing. The figures are useful for comparing these three matched sessions, but they should not be treated as universal model speeds.&lt;/p&gt;

&lt;p&gt;Model-file size is not the same as memory use either. The 8-bit server exited before it could produce a comparable final MLX allocation receipt, so I have not invented one for the table.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four-bit was surprisingly good
&lt;/h2&gt;

&lt;p&gt;Lattice is not just a Game of Life grid with a play button. It has draw and erase tools, play and single-step controls, randomisation, undo, speed and grid settings, symmetry modes, named patterns, live population and generation figures, and a small history chart.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpw04c5a5vp1u6x7zj4zc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpw04c5a5vp1u6x7zj4zc.png" alt="Lattice running in a dark interface, with a turquoise cellular automata grid, transport and drawing controls, a pattern library, live census figures, and a population-history chart." width="800" height="556"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The 4-bit model chose the idea and built this complete cellular automata workbench through Pi on my laptop.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The broad prompt did not ask for a cellular automata tool. The model chose that idea, decided what the workbench should contain, built it, opened it, exercised parts of it, and made changes over several rounds.&lt;/p&gt;

&lt;p&gt;I’m ridiculously impressed by it.&lt;/p&gt;

&lt;p&gt;It is also real enough to have real defects. Some library patterns do not behave as described. A few die or remain static even though the model’s tests later called them correct. Four-way symmetry can clip a stamp near the edge of the rectangular board. The browser smoke test showed that the controls ran without throwing errors; it did not prove the automata were correct.&lt;/p&gt;

&lt;p&gt;That distinction matters. Lattice is impressive because it is a coherent piece of software I can use and criticise, not because it passed a flattering demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Six-bit won the scorecard; four-bit won me
&lt;/h2&gt;

&lt;p&gt;The 6-bit model built &lt;a href="https://billiem.uk/local-choice/runs/2026-09-04-qwen38-27b-6bit-pi-create-003/artifact/" rel="noopener noreferrer"&gt;The Long Now&lt;/a&gt;, an interactive 5,000-year timeline with logarithmic, hybrid, and linear views. During the run it measured its first scale mapping, found that the maths did not support its own explanation, and replaced it with a hybrid scale.&lt;/p&gt;

&lt;p&gt;That inspect-and-repair behaviour is exactly why using Pi mattered. The result earned 4/5 for interest, 3/5 for execution, and 4/5 for taste in the formal review. It also contains overlapping labels and several simplified or wrong historical claims.&lt;/p&gt;

&lt;p&gt;The 4-bit artifact scored 4/5 for interest, 3/5 for execution, and 3/5 for taste. I still preferred it. A structured scorecard can help describe an artifact without replacing the judgement of the person actually using it.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://billiem.uk/local-choice/runs/2026-09-04-qwen38-27b-8bit-pi-create-002/artifact/" rel="noopener noreferrer"&gt;8-bit artifact&lt;/a&gt; is a preserved partial result. It had already built a substantial timeline and repaired some behaviour when the fourteenth provider call failed with a Metal out-of-memory error. The page renders, but its date calculations have central bugs and its labels crowd together.&lt;/p&gt;

&lt;p&gt;All three models loaded. Only two completed the agent run.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would run on this machine
&lt;/h2&gt;

&lt;p&gt;I would start with 4-bit Qwen3.8-27B on this 48 GB Mac.&lt;/p&gt;

&lt;p&gt;In this experiment it was about 28% faster than 6-bit by the whole-session aggregate measure. It left more headroom for the harness and everything else on the laptop, and it made my favourite artifact. That is a much more useful combination than choosing the largest quantisation that will technically initialise.&lt;/p&gt;

&lt;p&gt;Six-bit is completely viable when I want to trade some speed and memory headroom for the more considered behaviour it showed in The Long Now. I would not start another long 8-bit Pi session on 48 GB without shortening the working history or narrowing the task.&lt;/p&gt;

&lt;p&gt;This was one run per quantisation, not a leaderboard. I did not measure battery use, temperature, or repeated-run variance, and another seed or project could reverse the quality order.&lt;/p&gt;

&lt;p&gt;But the line has moved. A local model at 4-bit chose its own idea, used an agent loop to build it on my laptop, and left me with software I wanted to keep playing with.&lt;/p&gt;

&lt;p&gt;It did this at 4-bit. Damn, yeah.&lt;/p&gt;




&lt;p&gt;Want to talk about something I’ve written or built? &lt;a href="https://billiem.uk/contact/" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/qwen3-8-27b-on-m4-pro/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>macos</category>
      <category>programming</category>
    </item>
    <item>
      <title>Free compute worth claiming in September 2026</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Fri, 04 Sep 2026 10:28:49 +0000</pubDate>
      <link>https://dev.to/billiem/free-compute-worth-claiming-in-september-2026-1hg0</link>
      <guid>https://dev.to/billiem/free-compute-worth-claiming-in-september-2026-1hg0</guid>
      <description>&lt;p&gt;Free compute is easiest to misunderstand when the headline number is technically true.&lt;/p&gt;

&lt;p&gt;A platform says it includes a million requests, five CPU-hours, or $200 of compute. That does not tell you whether usage stops at the limit, whether a card can be charged, whether the credit happens once, or whether the capacity will exist when you need it.&lt;/p&gt;

&lt;p&gt;I rechecked the free-compute list for September against current provider documentation. This month produced fewer dramatic closures than August, but it exposed several numbers that needed correcting and a useful new pattern: infrastructure an agent can create before you even make an account.&lt;/p&gt;

&lt;h2&gt;
  
  
  Correct the old numbers first
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://deno.com/deploy/pricing" rel="noopener noreferrer"&gt;Deno Deploy Free&lt;/a&gt; is smaller than it was in the August snapshot. The current plan lists 1 million requests, 20 GiB egress, 10 CPU-hours, 150 GiB-hours of memory, ten apps, no Sandbox or volume storage, and 1 GiB KV per month.&lt;/p&gt;

&lt;p&gt;That replaces the previous 15 CPU-hours, 350 GiB-hours, 20 apps, and 1 GiB volume allowance. Deno does not expose a date for the change. Its &lt;a href="https://docs.deno.com/deploy/changelog/" rel="noopener noreferrer"&gt;changelog&lt;/a&gt; also warns that unverified organisations can receive restricted limits until a payment method is linked.&lt;/p&gt;

&lt;p&gt;Four other corrections matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.freestyle.sh/pricing" rel="noopener noreferrer"&gt;Freestyle VMs&lt;/a&gt; provides 200 vCPU-hours, 400 GiB memory-hours, and 60,000 GiB storage-hours per month. Those are monthly figures, not daily ones.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.browserless.io/pricing" rel="noopener noreferrer"&gt;Browserless&lt;/a&gt; now documents two-minute sessions, alongside 1,000 monthly units and two concurrent browsers.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://firebase.google.com/docs/auth/limits" rel="noopener noreferrer"&gt;Firebase Authentication&lt;/a&gt; does not simply provide 50,000 MAU on Spark. Identity Platform on Spark is normally limited to 3,000 DAU, while the 50,000-MAU no-cost tier belongs to Identity Platform on the billing-enabled Blaze plan.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://axiom.co/docs/reference/limits" rel="noopener noreferrer"&gt;Axiom Personal&lt;/a&gt; currently includes three datasets rather than two, with 25 GB stored, 500 GB monthly ingest, and 10 GB-hours of query compute.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of those is a glamorous new product launch. They are still more useful than carrying a wrong number into another month's list.&lt;/p&gt;

&lt;h2&gt;
  
  
  Know how the free usage ends
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://developers.cloudflare.com/changelog/post/2026-09-01-d1-free-tier-limit-enforcement/" rel="noopener noreferrer"&gt;Cloudflare D1&lt;/a&gt; now has an unusually clear answer. Since 1 September, a Free account that reaches 5 million rows read or 100,000 rows written in a day gets query errors until midnight UTC. Stored data remains intact.&lt;/p&gt;

&lt;p&gt;That is a hard cap. It can break an application, but it cannot quietly become a query-overage bill.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://vercel.com/docs/sandbox/pricing" rel="noopener noreferrer"&gt;Vercel Sandbox&lt;/a&gt; works the same way at the account level. Hobby includes five CPU-hours, 420 GB-hours of memory, 5,000 creations, 20 GB transfer, 15 GB of lifetime snapshots, 45-minute sessions, and ten concurrent sandboxes. Creation pauses after the allowance is exhausted.&lt;/p&gt;

&lt;p&gt;Compare that with &lt;a href="https://cloud.google.com/run/pricing" rel="noopener noreferrer"&gt;Google Cloud Run&lt;/a&gt;. Its 2 million requests, 180,000 vCPU-seconds, and 360,000 GiB-seconds are recurring free usage inside a billing account. Build, registry, storage, network, and excess runtime can still cost money.&lt;/p&gt;

&lt;p&gt;Both are legitimately useful. They just fail differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Infrastructure before signup
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://blog.railway.com/p/deploy-without-account" rel="noopener noreferrer"&gt;Railway now allows an unauthenticated visitor&lt;/a&gt; to create a private site through Dev New or start a database. The build session lasts 60 minutes, the result remains claimable for 24 hours, anonymous workloads can be throttled, and agent building is capped at $3 of LLM credit.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://developers.cloudflare.com/workers/platform/claim-deployments/" rel="noopener noreferrer"&gt;Cloudflare temporary accounts&lt;/a&gt; provide a similar one-hour loop from the command line. An agent can deploy a Worker with supported resources, verify it, and return a claim link. If the claim is not completed within 60 minutes, the account and resources disappear.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://axiom.co/docs/console/intelligence/agent-created-orgs" rel="noopener noreferrer"&gt;Axiom's agent-created organisations&lt;/a&gt; bring temporary observability into the same pattern. An agent can create an unauthenticated US East organisation with 10 GB ingest and one GB-hour of query compute. It must be claimed within 24 hours.&lt;/p&gt;

&lt;p&gt;These offers are not permanent free hosting. They are disposable proof environments with a handoff step. Claim links should be treated like credentials, and every unclaimed environment is designed to vanish.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical shortlist
&lt;/h2&gt;

&lt;p&gt;For a small real VM, &lt;a href="https://docs.oracle.com/en-us/iaas/Content/FreeTier/freetier_topic-Always_Free_Resources.htm" rel="noopener noreferrer"&gt;OCI Always Free&lt;/a&gt; remains the standout if Ampere capacity exists in your home region. The supported allowance is 1,500 OCPU-hours and 9,000 GB-hours per month, equivalent to 2 OCPUs and 12 GB RAM, plus two AMD micro instances and 200 GB combined boot and block storage.&lt;/p&gt;

&lt;p&gt;For application code, &lt;a href="https://developers.cloudflare.com/workers/platform/limits/" rel="noopener noreferrer"&gt;Cloudflare Workers&lt;/a&gt; gives 100,000 requests per day with 10 ms CPU per HTTP request. Deno remains useful if its reduced allowance fits. Cloud Run is the better fit when you need a container and are willing to manage a billing-linked account.&lt;/p&gt;

&lt;p&gt;For data, &lt;a href="https://neon.com/pricing" rel="noopener noreferrer"&gt;Neon&lt;/a&gt; remains a strong intermittent Postgres option with 100 CU-hours per project and scale-to-zero. &lt;a href="https://turso.tech/pricing" rel="noopener noreferrer"&gt;Turso&lt;/a&gt; covers distributed SQLite with 5 GB storage, 500 million rows read, and 10 million rows written per month. D1 fits naturally beside Workers now that its stopping rule is explicit.&lt;/p&gt;

&lt;p&gt;For disposable code execution, start with Vercel Sandbox, &lt;a href="https://upstash.com/pricing/box" rel="noopener noreferrer"&gt;Upstash Box&lt;/a&gt;, or Freestyle. For browser execution, &lt;a href="https://developers.cloudflare.com/browser-run/pricing/" rel="noopener noreferrer"&gt;Cloudflare Browser Run&lt;/a&gt; provides ten minutes per day, &lt;a href="https://www.browserbase.com/pricing" rel="noopener noreferrer"&gt;Browserbase&lt;/a&gt; provides one hour per month, and Browserless provides short sessions from its unit pool.&lt;/p&gt;

&lt;p&gt;For burst compute, &lt;a href="https://modal.com/pricing" rel="noopener noreferrer"&gt;Modal&lt;/a&gt; still provides $30 of recurring credit across CPU, GPU, notebooks, schedules, sandboxes, and web functions. &lt;a href="https://docs.digitalocean.com/products/paperspace/pricing/" rel="noopener noreferrer"&gt;Paperspace&lt;/a&gt; still lists capacity-dependent C4 and M4000 notebooks. &lt;a href="https://huggingface.co/docs/hub/main/en/spaces-zerogpu" rel="noopener noreferrer"&gt;Hugging Face ZeroGPU&lt;/a&gt; is much narrower: eligible personal accounts receive five shared GPU minutes per day for public Gradio Spaces.&lt;/p&gt;

&lt;p&gt;Hosted model APIs remain harder to compare. &lt;a href="https://developers.cloudflare.com/workers-ai/platform/pricing/" rel="noopener noreferrer"&gt;Workers AI&lt;/a&gt; gives 10,000 Neurons per day, but some models require billing or prepaid credits. Gemini's free limits vary by model, project, account, and region. OpenRouter provides 50 free-model requests per day, rising to 1,000 after at least $10 of credit has been purchased.&lt;/p&gt;

&lt;p&gt;The useful check is not just “how much is free?” Ask what happens at the edge: stop, sleep, delete, throttle, expire, or bill. September's changes make several of those answers clearer.&lt;/p&gt;




&lt;p&gt;Want to talk about something I’ve written or built? &lt;a href="https://billiem.uk/contact/" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/free-compute-september-2026/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>devops</category>
      <category>serverless</category>
      <category>ai</category>
    </item>
    <item>
      <title>The home server I finally stopped turning off</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Thu, 03 Sep 2026 09:47:19 +0000</pubDate>
      <link>https://dev.to/billiem/the-home-server-i-finally-stopped-turning-off-36c6</link>
      <guid>https://dev.to/billiem/the-home-server-i-finally-stopped-turning-off-36c6</guid>
      <description>&lt;p&gt;The most useful thing my home server taught me was not how to install another Docker container. It was how quickly a problem stops belonging to one tidy layer.&lt;/p&gt;

&lt;p&gt;A service can be running while DNS is wrong. Plex can work while the machine doing the transcoding cannot reach the storage. A reverse proxy can be configured correctly while the network around it is a mess. When it is your own server and you actually want to use it, those boundaries become your problem.&lt;/p&gt;

&lt;p&gt;That is very different from the way many application-focused software-engineering jobs feel. You can spend years building applications without having to join Linux, storage, DNS, HTTPS and networking together yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiments that kept getting turned off
&lt;/h2&gt;

&lt;p&gt;Around the start of 2020, I got a Raspberry Pi and repeatedly installed Raspbian or Debian on it. I would add Sonarr, Radarr, maybe Prowlarr, a torrent client and Plex. Sometimes Pi-hole joined them.&lt;/p&gt;

&lt;p&gt;There was no reverse proxy and I was not putting my own domains behind it. It was primitive, and I learnt something each time, but it never stuck, right? I would decide to play with it and eventually turn it off again.&lt;/p&gt;

&lt;p&gt;The Pi proved that I could run these services. It did not give me infrastructure I depended on.&lt;/p&gt;

&lt;p&gt;That changed in summer 2024. I had an old i5 desktop lying around, knew it worked and could connect drives to it easily. Why the hell not?&lt;/p&gt;

&lt;p&gt;I installed &lt;a href="https://www.openmediavault.org/" rel="noopener noreferrer"&gt;OpenMediaVault&lt;/a&gt; and spent the next two or three months building the setup out. Docker-managed services were joined by &lt;a href="https://doc.traefik.io/traefik/" rel="noopener noreferrer"&gt;Traefik&lt;/a&gt; as a reverse proxy, &lt;a href="https://tailscale.com/docs/concepts/what-is-tailscale" rel="noopener noreferrer"&gt;Tailscale&lt;/a&gt;, proper DNS and network sharing.&lt;/p&gt;

&lt;p&gt;The useful result was a repeatable path for a new service. I could put it behind HTTPS and decide whether it should be public or only reachable inside my network. The machine was no longer an experiment waiting to be unplugged.&lt;/p&gt;

&lt;h2&gt;
  
  
  A second machine made the lessons real
&lt;/h2&gt;

&lt;p&gt;I also bought a separate OptiPlex with 4 GB of RAM and installed Debian. Its main job was Plex Pass transcoding, reading media over the network from storage attached to the OpenMediaVault machine.&lt;/p&gt;

&lt;p&gt;That small decision forced several pieces to work together. The OptiPlex needed the network share. Plex needed to read the right files. Deployment now covered two Linux machines rather than one. Running another Pi-hole there also gave me two DNS servers.&lt;/p&gt;

&lt;p&gt;None of this was exotic. That was precisely why it was useful. These were ordinary operational problems with an outcome I cared about: I wanted my own services to work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure that taught me to stop
&lt;/h2&gt;

&lt;p&gt;I did not keep every layer.&lt;/p&gt;

&lt;p&gt;That same summer, I tried turning an old laptop into my router with &lt;a href="https://openwrt.org/" rel="noopener noreferrer"&gt;OpenWrt&lt;/a&gt;. It kept dropping out, which meant I was repeatedly taking down my own internet while trying to learn the rest of the server.&lt;/p&gt;

&lt;p&gt;Eventually I bought a router.&lt;/p&gt;

&lt;p&gt;The lesson was not that I should keep going until I became better at OpenWrt. It was that you need to know when to just not do it yourself. I reached the point where I thought: “I just want something that works. I’m going to stop trying to overcomplicate it. I’m just going to make it work.”&lt;/p&gt;

&lt;p&gt;That judgement matters as much as knowing how to add another service. A home server gives you a place to learn across layers; it does not require you to own every layer.&lt;/p&gt;

&lt;p&gt;It is not free either. It uses energy. Plex Pass costs money, as can a VPN or another service. Sometimes the server is down when you just want to watch a film. Self-hosting will cause a fucking headache.&lt;/p&gt;

&lt;p&gt;What are you going to do, right? I still think it is worth it. I use the result all the time, and the problems have given me practical reasons to understand Docker, Linux, networking, storage and operations together.&lt;/p&gt;

&lt;p&gt;The next change was moving the setup onto a different desktop and Proxmox. That is another part of the story.&lt;/p&gt;




&lt;p&gt;Want to talk about something I’ve written or built? &lt;a href="https://billiem.uk/contact/" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/you-need-a-home-server/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>selfhosting</category>
      <category>linux</category>
      <category>docker</category>
      <category>homelab</category>
    </item>
    <item>
      <title>How I use scheduled Codex jobs to build daily reports</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Thu, 06 Aug 2026 09:14:22 +0000</pubDate>
      <link>https://dev.to/billiem/how-i-use-scheduled-codex-jobs-to-build-daily-reports-1n0m</link>
      <guid>https://dev.to/billiem/how-i-use-scheduled-codex-jobs-to-build-daily-reports-1n0m</guid>
      <description>&lt;p&gt;Most of my analytics job should not be done by an LLM.&lt;/p&gt;

&lt;p&gt;The source metrics, dates and repeatable transformations should stay exact. I used to think that meant the whole scheduled job should be left to deterministic code, and I still kind of agree.&lt;/p&gt;

&lt;p&gt;What changed my mind was not asking a model to calculate the report. It was putting a scheduled Codex job around reliable tools: run them, inspect their result, apply a small amount of bounded judgement, and leave behind an HTML artefact that explains what happened.&lt;/p&gt;

&lt;p&gt;I now use that shape for two daily reports. One is mostly analytics. The other has a more editorial middle. The useful boundary is similar in both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the hard boundary boring
&lt;/h2&gt;

&lt;p&gt;My billiem-analytics project collects website visitors, edge traffic, &lt;a href="https://support.google.com/webmasters/answer/7576553" rel="noopener noreferrer"&gt;Google Search Console&lt;/a&gt;, DEV and Hashnode into dated local snapshots.&lt;/p&gt;

&lt;p&gt;Those sources do not expose interchangeable numbers. Some are daily flows. DEV and Hashnode provide cumulative publishing counters. A real zero is different from an unavailable provider, missing configuration or failed collection.&lt;/p&gt;

&lt;p&gt;The deterministic code owns those distinctions. It rebuilds the private dashboard and keeps the reporting periods consistent. Codex does not get to smooth over a missing source or turn an unreliable metric into a reliable one.&lt;/p&gt;

&lt;p&gt;That is important because the final report is meant to be inspected, not merely produced. If one provider failed, I want that state to survive all the way to the page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give the model the awkward work around the edges
&lt;/h2&gt;

&lt;p&gt;The scheduled job handles the surrounding work: run the existing commands, inspect their output, follow a bounded recovery rule and decide how to present the result.&lt;/p&gt;

&lt;p&gt;That is a small role in the analytics workflow. Founder Brief gives the model more to do.&lt;/p&gt;

&lt;p&gt;There, a Python collector gathers and ranks candidate material into an editorial packet. The scheduled run reads the packet, checks the strongest sources and useful discussion branches, rejects weak or promotional material, and writes a short briefing. Candidate IDs and validation keep the result connected to the collected evidence.&lt;/p&gt;

&lt;p&gt;The judgement is real but limited. A fixed ranking can decide what to inspect first; it cannot comfortably decide whether an anecdote is useful, whether two observations form a pattern, or whether a source gap makes a conclusion too strong.&lt;/p&gt;

&lt;p&gt;Codex can help with those decisions against written editorial rules. It still cannot guarantee access to a site. Authentication, access controls, rate limits and format changes remain actual boundaries. Flexibility is not the same as a scraper that can never break.&lt;/p&gt;

&lt;h2&gt;
  
  
  The report is what made the automation useful
&lt;/h2&gt;

&lt;p&gt;Both jobs end in the same deliberately plain way: just chuck it on an HTML file.&lt;/p&gt;

&lt;p&gt;That gives me a stable artefact instead of a chat transcript or terminal session I need to reconstruct. Dates, source states, comparisons and caveats are together in a form I can reopen.&lt;/p&gt;

&lt;p&gt;Before the analytics report, checking the same picture meant opening several browser windows and lining up different periods myself. It was painful enough that I often did not bother. Now I get the nice graphs and the consistent dates in one private view.&lt;/p&gt;

&lt;p&gt;That consolidation made a quiet period without publishing visible across the sources. It was not a surprising new metric. The useful part was finally being able to see the existing information together.&lt;/p&gt;

&lt;p&gt;I also added the report to &lt;a href="https://manual.raycast.com/quicklinks" rel="noopener noreferrer"&gt;Raycast&lt;/a&gt;. Typing &lt;code&gt;analytics&lt;/code&gt; opens it in Chrome and regenerates it as appropriate. That tiny access path matters more than it sounds: I am actually opening the report rather than deciding that a manual Search Console session can wait.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recovery has to remain visible
&lt;/h2&gt;

&lt;p&gt;The analytics job includes one narrow recovery rule. A sandboxed DNS or Keychain failure should be retried with the correct local permissions before the provider is declared broken. After that retry, every provider still receives its own explicit state.&lt;/p&gt;

&lt;p&gt;This is useful orchestration, not a claim that the system heals itself. A local scheduled job can miss its time when the Mac is asleep or offline. Credentials expire. Providers change. Prompts can be wrong.&lt;/p&gt;

&lt;p&gt;The model can respond to a known operational wrinkle without hiding a genuine failure behind a confident summary. That is the line I care about.&lt;/p&gt;

&lt;p&gt;Tokens feel cheap enough to me that this extra layer is worth trying, and it does not need a frontier model. If a cheap model can wrap reliable code, handle a known failure path and occasionally surface something the fixed report would miss, why wouldn't I use it?&lt;/p&gt;

&lt;p&gt;I have not arrived at the conclusion that every scheduled task needs an LLM. Plenty still need a plain scheduler and a script. The pattern works for me when the exact work stays exact, the judgement stays bounded, and the final artefact makes the whole run inspectable.&lt;/p&gt;




&lt;p&gt;Want to talk about something I’ve written or built? &lt;a href="https://billiem.uk/contact/" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/scheduled-codex-jobs-daily-reports/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>codex</category>
      <category>automation</category>
      <category>analytics</category>
      <category>ai</category>
    </item>
    <item>
      <title>The LLM was better at building a solver than playing the game</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Tue, 04 Aug 2026 09:47:21 +0000</pubDate>
      <link>https://dev.to/billiem/the-llm-was-better-at-building-a-solver-than-playing-the-game-5ck0</link>
      <guid>https://dev.to/billiem/the-llm-was-better-at-building-a-solver-than-playing-the-game-5ck0</guid>
      <description>&lt;p&gt;I started this project because an LLM annoyed me.&lt;/p&gt;

&lt;p&gt;I gave a very strong model &lt;a href="https://322-0.app/" rel="noopener noreferrer"&gt;322&lt;/a&gt;, a small Dota 2 drafting game. The choices looked like the kind of work a computer should enjoy: repeated packs of players and heroes, visible ratings, familiarity scores, chemistry, rerolls and a simulated tournament at the end.&lt;/p&gt;

&lt;p&gt;I was disappointed by how well the LLM did. I am not a Dota expert, and I had only started watching it occasionally again during the previous six months or year. I still seemed to be doing better.&lt;/p&gt;

&lt;p&gt;The interesting engineering question was not how to write a longer prompt. It was how to replace the card-by-card language-model judgement with a deterministic policy, then test that policy without confusing improvement with luck.&lt;/p&gt;

&lt;h2&gt;
  
  
  A stochastic benchmark needs shared randomness
&lt;/h2&gt;

&lt;p&gt;The browser history gave us a useful irritation and almost no reliable comparison.&lt;/p&gt;

&lt;p&gt;My earlier manual record contained 50 runs with a 14% title rate. The LLM won once in nine attempts. Putting 14% beside 11% looks temptingly quantitative, but the random offers, rejected packs and opponent fields were not preserved. The samples were small, unpaired and produced under different choices.&lt;/p&gt;

&lt;p&gt;That is not a model benchmark. It is a reason to build one.&lt;/p&gt;

&lt;p&gt;The offline solver generated every random choice from indexed tapes. Policy A and policy B received the same player offers, hero samples, field candidates and tournament randomness for a given episode. We could then compare the paired result: did the new policy win this exact episode where the old policy lost it?&lt;/p&gt;

&lt;p&gt;This is the common-random-numbers idea in a practical form. Sharing the luck removes a large amount of noise that has nothing to do with the policy change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the simulator separate from the policy
&lt;/h2&gt;

&lt;p&gt;Before evaluating a strategy, we reproduced the game.&lt;/p&gt;

&lt;p&gt;The public client and seven data files were frozen with SHA-256 hashes. Draft legality, automatic hero allocation, chemistry, scoring and the tournament were ported into a deterministic Python engine. Automatic allocation evaluates all 120 player-to-hero permutations, so reproducing that detail mattered.&lt;/p&gt;

&lt;p&gt;The engine also kept an important semantic boundary. A finished roster has an exact score under the copied rules. An unfinished draft does not. A policy can estimate the future value of a player, hero or reroll, but that partial-state estimate is not made exact by giving it several decimal places.&lt;/p&gt;

&lt;p&gt;Keeping the simulator authoritative meant we could change policies without changing what success meant.&lt;/p&gt;

&lt;h2&gt;
  
  
  The redraw budget is part of the experiment
&lt;/h2&gt;

&lt;p&gt;Opponent fields introduced a slightly odd control problem. If a solver can redraw for free forever, “find an easier field” eventually dominates every other decision.&lt;/p&gt;

&lt;p&gt;We declared three finite budgets instead:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;K=1&lt;/code&gt; accepts the first opponent field.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;K=4&lt;/code&gt; selects from four fields.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;K=16&lt;/code&gt; selects from sixteen fields.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The budgets were nested. The first field in &lt;code&gt;K=1&lt;/code&gt; was also the first field in &lt;code&gt;K=4&lt;/code&gt; and &lt;code&gt;K=16&lt;/code&gt;; the first four were shared too. Policies could be compared at each budget without quietly receiving different field draws.&lt;/p&gt;

&lt;p&gt;This was a small design choice with a large effect on the claim. A title rate without the field budget would not describe a reproducible policy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the final test expensive to touch
&lt;/h2&gt;

&lt;p&gt;The solver could be improved and improved and improved. That was satisfying, but it created the usual risk: every time a result influences the next change, that result has become part of training.&lt;/p&gt;

&lt;p&gt;We separated train, validation and test tapes. Candidate policies first went through successive halving on shared training episodes. Weak candidates stopped early; survivors received more budget. The selected candidate then needed to pass a fresh validation gate before the test partition could be opened.&lt;/p&gt;

&lt;p&gt;The gate required positive title-rate results, no meaningful top-four or top-eight regression, zero failures, deterministic replay and acceptable runtime. Once the final test started, policy code, features, model bytes, parameters, seeds and episode counts were frozen.&lt;/p&gt;

&lt;p&gt;The controls themselves still needed scrutiny. An independent audit found that an early replay check compared duplicate in-memory runs rather than the public replay path, and that the first runtime gate reused warm policy caches. Both were repaired before the production claim. The audit separately recomputed the final statistics from the raw logs.&lt;/p&gt;

&lt;p&gt;The inconvenience was deliberate. It stopped “one more tweak” from leaking the answer back into the candidate.&lt;/p&gt;

&lt;h2&gt;
  
  
  A small deterministic model was enough
&lt;/h2&gt;

&lt;p&gt;The final test compared four generations of policy on 10,000 paired episodes at each field budget. Across all policy and budget combinations, that was 120,000 outcomes with zero failures.&lt;/p&gt;

&lt;p&gt;At &lt;code&gt;K=16&lt;/code&gt; the progression was:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Policy&lt;/th&gt;
&lt;th&gt;Title rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pick the highest player rating&lt;/td&gt;
&lt;td&gt;2.30%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hand-written heuristic&lt;/td&gt;
&lt;td&gt;14.77%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assignment and chemistry components&lt;/td&gt;
&lt;td&gt;22.33%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frozen 37-feature value model&lt;/td&gt;
&lt;td&gt;27.45%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The final value model was ridge-linear, deterministic and trained offline on 8,511 action rows. It used features for role-feasible future strength, partial hero allocation, current and potential chemistry, familiarity, draft stage and remaining rerolls. It did not call an LLM while playing.&lt;/p&gt;

&lt;p&gt;Against the previous component policy, its &lt;code&gt;K=16&lt;/code&gt; title improvement was +5.12 percentage points with a paired 95% interval from +4.24 to +6.00. Top-four and top-eight results improved at every budget too.&lt;/p&gt;

&lt;p&gt;The complete draft loop took a median 47.1 milliseconds with a cold model build. That was roughly 1.8 times the component policy's latency, and the runtime policy made no large-model request for any offer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put the LLM around the decision loop
&lt;/h2&gt;

&lt;p&gt;We try to solve everything with LLMs now, and we do not necessarily need to.&lt;/p&gt;

&lt;p&gt;For 322, the runtime decision-maker became ordinary deterministic code. The LLM was much more useful around it: inspecting the browser game, reconstructing mechanics, implementing the engine, proposing policies, running controlled experiments, auditing weak controls and producing visualisations that made the improvement legible.&lt;/p&gt;

&lt;p&gt;The result was quite a cool little statistical project. The useful boundary was not “LLMs are bad at games.” It was more specific: repeated numeric decisions with a reproducible state and measurable outcome did not need an LLM in the hot path.&lt;/p&gt;

&lt;p&gt;If I continued the project, I would use the LLM to choose and build the next experiment. I would leave the next card choice to the solver.&lt;/p&gt;




&lt;p&gt;Want to talk about something I’ve written or built? &lt;a href="https://billiem.uk/contact/" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/llm-better-building-322-solver/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>datascience</category>
      <category>gamedev</category>
    </item>
    <item>
      <title>Free compute worth claiming in August 2026</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Mon, 03 Aug 2026 11:26:44 +0000</pubDate>
      <link>https://dev.to/billiem/free-compute-worth-claiming-in-august-2026-154d</link>
      <guid>https://dev.to/billiem/free-compute-worth-claiming-in-august-2026-154d</guid>
      <description>&lt;p&gt;A monthly free-compute list is only useful if it is willing to delete things.&lt;/p&gt;

&lt;p&gt;Between my &lt;a href="https://billiem.uk/posts/free-compute-july-2026/" rel="noopener noreferrer"&gt;July snapshot&lt;/a&gt; and 3 August, GitHub Models disappeared, SageMaker Studio Lab closed to new users, and the current Hugging Face documentation stopped supporting the broad free CPU Spaces claim I had made. Paperspace, meanwhile, turned up with documented free M4000 notebook access. Neon put Functions and Object Storage into a free beta. Managed agent runtime and browser automation became harder to treat as side notes.&lt;/p&gt;

&lt;p&gt;I also had to correct my own July guide. GitHub had already announced that Models would retire on 30 July, but I still described it as a recurring public preview on 10 July. That was wrong.&lt;/p&gt;

&lt;p&gt;This is the useful August shortlist: what to remove, what changed, and which kind of free allowance fits a particular side project. It is documentation-backed rather than a claim that I opened and tested every account. Capacity, regions, billing checks, data-use terms, and unpublished guardrails still matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Remove these from the old shortlist
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.blog/changelog/2026-07-01-github-models-is-being-fully-retired-on-july-30-2026/" rel="noopener noreferrer"&gt;GitHub Models retired completely&lt;/a&gt;, including its playground, catalogue, inference API, and BYOK endpoints. It is no longer a free model API.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/sagemaker/latest/dg/studio-lab-overview.html" rel="noopener noreferrer"&gt;SageMaker Studio Lab&lt;/a&gt; closed to new customers on 30 July. Existing users retain free sessions with 15 GB storage, 16 GB RAM, up to eight CPU hours or four GPU hours per day, subject to capacity. A new reader cannot claim it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://huggingface.co/docs/hub/spaces-overview" rel="noopener noreferrer"&gt;Hugging Face Spaces&lt;/a&gt; needs a narrower description. Normal Gradio and Docker compute Spaces are gated behind paid plans in the current documentation. A free personal account can still host up to two &lt;a href="https://huggingface.co/docs/hub/main/en/spaces-zerogpu" rel="noopener noreferrer"&gt;Gradio ZeroGPU Spaces&lt;/a&gt;, but that is queued public-demo GPU capacity rather than a free Linux box.&lt;/p&gt;

&lt;p&gt;One more correction is smaller but worth keeping: &lt;a href="https://axiom.co/docs/reference/limits" rel="noopener noreferrer"&gt;Axiom Personal&lt;/a&gt; currently includes two datasets, not the three I listed in July.&lt;/p&gt;

&lt;h2&gt;
  
  
  The August additions worth knowing about
&lt;/h2&gt;

&lt;h3&gt;
  
  
  A genuinely free GPU notebook appeared
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://docs.digitalocean.com/products/paperspace/pricing/" rel="noopener noreferrer"&gt;Paperspace Notebooks&lt;/a&gt; now documents C4 CPU and M4000 GPU machines on its Free Gradient plan. It includes 5 GB storage and permits &lt;a href="https://docs.digitalocean.com/products/paperspace/notebooks/getting-started/run-example-notebooks/" rel="noopener noreferrer"&gt;one running notebook at a time&lt;/a&gt;. A &lt;a href="https://docs.digitalocean.com/products/paperspace/notebooks/details/features/" rel="noopener noreferrer"&gt;free-machine session can last up to six hours&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That does not mean guaranteed GPU hours. Capacity may queue, and there is no published monthly-hour entitlement. It is useful shared notebook access, not a persistent GPU server.&lt;/p&gt;

&lt;h3&gt;
  
  
  Neon temporarily became more than Postgres
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://neon.com/blog/neon-backend-is-beta" rel="noopener noreferrer"&gt;Neon Functions and Object Storage&lt;/a&gt; are available on the Free plan during beta. The current offer is restricted to &lt;code&gt;us-east-2&lt;/code&gt;, retains logs for three days, and uses unpublished rate and usage guardrails. AI Gateway is not included on Free.&lt;/p&gt;

&lt;p&gt;That makes Neon an interesting temporary full-stack backend, but a beta is not a promise. The underlying &lt;a href="https://neon.com/pricing" rel="noopener noreferrer"&gt;free Postgres plan&lt;/a&gt; remains easier to reason about: up to 100 projects with 100 CU-hours and 0.5 GB storage per project, plus scale-to-zero after five minutes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Agent compute now has several shapes
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://docs.cloud.google.com/free/docs/free-cloud-features" rel="noopener noreferrer"&gt;Google Agent Engine&lt;/a&gt; is listed with 180,000 vCPU-seconds and 360,000 GiB-seconds per month. Billing must be enabled. The table does not establish that Code Execution sandboxes share this allowance, so I would use the claim narrowly: managed agent runtime.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://upstash.com/pricing/box" rel="noopener noreferrer"&gt;Upstash Box&lt;/a&gt; provides ten concurrent boxes and five active CPU-hours each month. A box has 2 vCPU, 4 GB RAM, and 5 GB disk, freezes after an idle hour, and currently runs only in &lt;code&gt;us-east-1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://developers.cloudflare.com/workers/platform/claim-deployments/" rel="noopener noreferrer"&gt;Cloudflare temporary accounts&lt;/a&gt; solve a different problem. An unauthenticated agent can deploy a Worker and supported resources with &lt;code&gt;wrangler deploy --temporary&lt;/code&gt;, verify the result, and return a claim URL. The account disappears after 60 minutes if it is not claimed. This is throwaway deployment compute, not a permanent anonymous account.&lt;/p&gt;

&lt;h3&gt;
  
  
  Browser automation is now its own free-compute category
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://developers.cloudflare.com/browser-run/pricing/" rel="noopener noreferrer"&gt;Cloudflare Browser Run&lt;/a&gt; gives ten browser minutes per day and three concurrent browsers. Cloudflare's general-purpose Sandbox SDK is separate and has no Workers Free allocation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.browserbase.com/pricing" rel="noopener noreferrer"&gt;Browserbase&lt;/a&gt; includes one browser hour, three concurrent browsers, three agent runs, and 1,000 each of Search and Fetch uses per month. Sessions stop after 15 minutes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://cloud.browserless.io/pricing" rel="noopener noreferrer"&gt;Browserless&lt;/a&gt; uses a 1,000-unit monthly pool with two concurrent browsers and one-minute sessions. One unit covers up to 30 seconds, while proxy and CAPTCHA work can consume more. It does not translate cleanly into a neat number of hours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pick by the way the allowance can fail
&lt;/h2&gt;

&lt;p&gt;The word “free” hides several different products. The failure mode matters more than the headline number.&lt;/p&gt;

&lt;h3&gt;
  
  
  For a persistent Linux VM
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://docs.oracle.com/en-us/iaas/Content/FreeTier/freetier_topic-Always_Free_Resources.htm" rel="noopener noreferrer"&gt;OCI Always Free&lt;/a&gt; is still the first claim to try. The current supported allocation is 1,500 Ampere A1 OCPU-hours and 9,000 GB-hours per month, equivalent to 2 OCPUs and 12 GB RAM. The tenancy also gets 200 GB combined boot and block storage plus 10 TB monthly outbound transfer.&lt;/p&gt;

&lt;p&gt;The catches are practical: home-region capacity can be unavailable, and idle instances can be reclaimed. &lt;a href="https://aws.amazon.com/ec2/instance-types/t4/" rel="noopener noreferrer"&gt;AWS EC2 T4g&lt;/a&gt; is useful during 2026, but its 750 aggregate &lt;code&gt;t4g.small&lt;/code&gt; hours per month are a promotion ending on 31 December rather than durable infrastructure.&lt;/p&gt;

&lt;h3&gt;
  
  
  For functions and containers
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://developers.cloudflare.com/workers/platform/limits/" rel="noopener noreferrer"&gt;Cloudflare Workers&lt;/a&gt; has a recurring hard cap of 100,000 requests per day, 10 ms CPU per HTTP request, and 128 MB memory. Static assets do not consume the dynamic request allowance.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://deno.com/deploy/pricing" rel="noopener noreferrer"&gt;Deno Deploy&lt;/a&gt; publishes a broader monthly cap: 1 million requests, 20 GB egress, 15 CPU-hours, 350 GB-hours memory, 20 apps, 1 GiB volume storage, and 1 GiB KV storage. The old Deploy Classic migration deadline has passed, and projects did not transfer automatically.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://cloud.google.com/run/pricing" rel="noopener noreferrer"&gt;Google Cloud Run&lt;/a&gt; gives request-based services 2 million requests, 180,000 vCPU-seconds, and 360,000 GiB-seconds monthly. It is attached to billing, so registry, build, network, storage, or overage usage can still charge.&lt;/p&gt;

&lt;h3&gt;
  
  
  For data and product plumbing
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://turso.tech/pricing" rel="noopener noreferrer"&gt;Turso&lt;/a&gt; is the cleaner hard-capped SQLite-shaped option: 100 databases, 5 GB storage, 500 million rows read, and 10 million rows written per month. Requests stop at quota rather than becoming an invoice.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://developers.cloudflare.com/d1/platform/pricing/" rel="noopener noreferrer"&gt;Cloudflare D1&lt;/a&gt; fits naturally when the application already lives on Workers. The free allocation is ten databases, 500 MB each, 5 GB total, 5 million rows read and 100,000 rows written per day.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://upstash.com/pricing" rel="noopener noreferrer"&gt;Upstash&lt;/a&gt; covers Redis, QStash, Workflow, Vector, and preview Search under separate free quotas. &lt;a href="https://developers.cloudflare.com/r2/pricing/" rel="noopener noreferrer"&gt;Cloudflare R2&lt;/a&gt; provides 10 GB-month Standard storage, 1 million Class A operations, 10 million Class B operations, and free egress, but enabling it requires account checkout.&lt;/p&gt;

&lt;h3&gt;
  
  
  For GPU and hosted inference
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://modal.com/pricing" rel="noopener noreferrer"&gt;Modal Starter&lt;/a&gt; gives $30 of recurring compute credit for serverless CPU, GPU, notebooks, schedules, sandboxes, and web functions. &lt;a href="https://lightning.ai/pricing/" rel="noopener noreferrer"&gt;Lightning AI Free&lt;/a&gt; gives 15 credits per month plus one active CPU Studio that restarts every four hours. The GPU time those credits buy varies with the selected hardware and current rates.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://research.google.com/colaboratory/faq.html" rel="noopener noreferrer"&gt;Google Colab&lt;/a&gt; remains shared and dynamic, with no stable hardware or quota promise. Sessions can run for up to 12 hours depending on usage and availability. Paperspace is now the more concrete free M4000 notebook claim, while ZeroGPU is for public Gradio demos.&lt;/p&gt;

&lt;p&gt;Hosted APIs are another category again. &lt;a href="https://developers.cloudflare.com/workers-ai/platform/pricing/" rel="noopener noreferrer"&gt;Cloudflare Workers AI&lt;/a&gt; stops a Free account after 10,000 Neurons per day. &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Gemini&lt;/a&gt; has model-dependent free pricing and may use free-tier content to improve Google products. &lt;a href="https://openrouter.ai/docs/api/reference/limits" rel="noopener noreferrer"&gt;OpenRouter Free&lt;/a&gt; gives 50 free-model requests per day unless the account has purchased credit.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://freeinference.org/" rel="noopener noreferrer"&gt;FreeInference&lt;/a&gt; is useful only with its privacy boundary visible: its &lt;a href="https://freeinference.org/terms" rel="noopener noreferrer"&gt;terms&lt;/a&gt; allow prompts and responses to be logged and anonymised derived data to be released. Never send it secrets, private code, or personal data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three stacks I would still assemble from the list
&lt;/h2&gt;

&lt;p&gt;For a small real server: OCI for the VM, Cloudflare for DNS, CDN, and TLS, R2 or Tigris for objects, Neon or Aiven for Postgres, Upstash for queues and Redis, Resend for mail, and Grafana Cloud plus Sentry for visibility.&lt;/p&gt;

&lt;p&gt;For a no-server product: Cloudflare Workers and static assets, D1 or Neon, R2, Queues or Workflows, Turnstile, Clerk or AuthKit, and Workers AI.&lt;/p&gt;

&lt;p&gt;For an AI demo: Paperspace or Colab for notebook work, Modal for repeatable bursts, ZeroGPU for a public Gradio surface, and Workers AI, Gemini, Groq, or OpenRouter for hosted inference.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://billiem.uk/posts/free-compute-august-2026/" rel="noopener noreferrer"&gt;full August guide on billiem.uk&lt;/a&gt; keeps the broader directory, including email, auth, observability, CI, storage, startup programmes, and academic credits.&lt;/p&gt;

&lt;p&gt;In 24 days, one model API disappeared, one notebook closed to new users, and a supposed free hosting offer became much narrower. The next version may remove as much as it adds. That is fine. A dated guide is more useful when it is willing to get shorter.&lt;/p&gt;




&lt;p&gt;Want to talk about something I’ve written or built? &lt;a href="https://billiem.uk/contact/" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/free-compute-august-2026/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>webdev</category>
      <category>devops</category>
      <category>ai</category>
    </item>
    <item>
      <title>I let Codex build and test my first native Mac app</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Mon, 03 Aug 2026 11:19:01 +0000</pubDate>
      <link>https://dev.to/billiem/i-let-codex-build-and-test-my-first-native-mac-app-4ik8</link>
      <guid>https://dev.to/billiem/i-let-codex-build-and-test-my-first-native-mac-app-4ik8</guid>
      <description>&lt;p&gt;The first public Billie Flow build launched, stayed alive, passed its signature check and could find its source and models. A replacement install still left a new user with no Settings window, which meant there was no visible way to install the local worker.&lt;/p&gt;

&lt;p&gt;The process was healthy. The app was unusable.&lt;/p&gt;

&lt;p&gt;That gap is the part of my local dictation experiment that I find most useful. I had let Codex work well beyond a repository: it built my first native Mac app, installed it, operated the macOS permission surfaces, exercised physical audio and inspected the clipboard. The interesting failures only appeared because it kept going after the code checks were green.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing the installed app changed the work
&lt;/h2&gt;

&lt;p&gt;I started Billie Flow because I use &lt;a href="https://wisprflow.ai/" rel="noopener noreferrer"&gt;Wispr Flow&lt;/a&gt; and wondered whether I could make the useful loop run locally. What led me to it was just curiosity. Could I really?&lt;/p&gt;

&lt;p&gt;It was never meant to become a replacement. Wispr Flow gives me enough value that I still use it. I wanted to see whether a personal version could be made, and then I became much more interested in how involved I let Codex become.&lt;/p&gt;

&lt;p&gt;In 5.6 it could literally write the thing, build the thing, install the dependencies for the thing, test the thing. For this project, "test" eventually meant opening the installed Swift app, working through microphone permissions with Computer Use, holding the global shortcut and checking whether the resulting text reached the clipboard.&lt;/p&gt;

&lt;p&gt;At one point I was lying there waiting for it to work away when speech started coming out of my speakers. Billie Flow recorded it and transcribed it. I do not know what produced that speech, so I am not going to turn the observation into a tidier technical claim. I remember thinking that 5.5 would not have taken the same initiative, but that is an impression rather than a controlled comparison.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fapy3qtpscwxs57beotro.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fapy3qtpscwxs57beotro.png" alt="A compact Billie Flow HUD reads Recording, 0:01, release to finish, beside a five-bar level meter." width="672" height="236"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The global shortcut records only while held; releasing it submits the temporary audio for local processing.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The physical run eventually finished with non-empty text on the clipboard, a healthy app and worker, no new crash and no temporary recording left behind.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4f4lskf2xz4spzmv56oy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4f4lskf2xz4spzmv56oy.png" alt="Billie Flow's HUD reads Copied, Light cleanup, and Ready on the clipboard beside a document icon." width="672" height="236"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is the visible end state from the installed app after local recognition and light cleanup completed.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The expected model was not the useful one
&lt;/h2&gt;

&lt;p&gt;The app work started with a model decision. I had an audio-capable Gemma 12B route in my head, but I ran the same 35.3-second voice memo through several local recognition paths before building around that assumption.&lt;/p&gt;

&lt;p&gt;I also kept speech recognition separate from cleanup. Otherwise a cleanup model could make a transcript sound polished while preserving the important names the recogniser had already got wrong.&lt;/p&gt;

&lt;p&gt;Gemma completed the memo in 258.79 seconds and drifted around an overlapping chunk. MLX Whisper large-v3-turbo produced the most useful recognition in about 3.68 seconds, then Qwen2.5 1.5B ran the selected light-cleanup pass in about 0.62 seconds.&lt;/p&gt;

&lt;p&gt;The model I first had in my head was not the right answer. Two smaller models produced the quicker and more useful result for this app. That is deliberately narrow: it was one memo on one machine, and every recognition path still missed at least one important project term. The &lt;a href="https://billiem.uk/reports/billie-flow-model-analysis/" rel="noopener noreferrer"&gt;full Billie Flow model analysis&lt;/a&gt; contains the other branches, timings and vocabulary failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three different kinds of green failure
&lt;/h2&gt;

&lt;p&gt;The absent Settings window was not the only thing the replacement-install pass found.&lt;/p&gt;

&lt;p&gt;Cleanup was silently falling back to raw recognition because the pinned MLX library was receiving an obsolete argument. The worker request returned success, but the UI was claiming cleanup that had not happened. A final setup check could also exit before the app registered completion and leave setup stuck at verification.&lt;/p&gt;

&lt;p&gt;Earlier microphone tests had exposed an actor-isolation crash on the first audio buffer, then a format mismatch while writing the converted WAV. These were not five versions of the same bug. They crossed Swift concurrency, Core Audio, a pinned Python dependency, process lifecycle and visible macOS presentation.&lt;/p&gt;

&lt;p&gt;Codex kept following each failure into the layer that produced it. The release gate now checks for an actual first-launch window and warning-free cleanup instead of treating a living process or successful response as sufficient evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  A public proof of concept, with a blunt boundary
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://github.com/BillieM/billie-flow" rel="noopener noreferrer"&gt;source&lt;/a&gt; and &lt;a href="https://github.com/BillieM/billie-flow/releases/tag/v0.2.1" rel="noopener noreferrer"&gt;install-tested v0.2.1 build&lt;/a&gt; are public. I want Billie Flow to be usable should you want to, but I do not really care if anyone does.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgd5flibsqfbzjr6b0opd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgd5flibsqfbzjr6b0opd.png" alt="Billie Flow's Install local speech models dialog says it will download about 3.5 GB from Hugging Face and requires Apple Silicon and macOS 26." width="800" height="885"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Nothing large starts until Install is chosen; the disclosure also limits the proof of concept to English speech on Apple Silicon and macOS 26.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The boundary is explicit: Apple Silicon, macOS 26, English recognition and about 3.5 GB of consented runtime and model setup. Inference is local after setup, but setup still downloads Python dependencies and fixed models. The app is ad-hoc signed and unnotarised, with no updater, App Store release, compatibility promise or support plan. I am definitely not getting an Apple Developer Programme membership for it at this point.&lt;/p&gt;

&lt;p&gt;Billie Flow does not establish that everybody should rebuild the software they already use. It shows that a surprisingly good personal version of this loop can now be made quickly. What I am going to remember is still the oddest part: lying there while Codex decided it needed to test the microphone and speech started coming out of my speakers. It was really really impressive.&lt;/p&gt;




&lt;p&gt;Want to talk about something I’ve written or built? &lt;a href="https://billiem.uk/contact/" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/codex-built-local-wispr-flow/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>macos</category>
      <category>machinelearning</category>
      <category>programming</category>
    </item>
    <item>
      <title>A rotation changed my embedding plot, not its neighbours</title>
      <dc:creator>Billie M</dc:creator>
      <pubDate>Tue, 14 Jul 2026 15:18:33 +0000</pubDate>
      <link>https://dev.to/billiem/a-rotation-changed-my-embedding-plot-not-its-neighbours-53f3</link>
      <guid>https://dev.to/billiem/a-rotation-changed-my-embedding-plot-not-its-neighbours-53f3</guid>
      <description>&lt;p&gt;One switch in my embedding animation makes the plot look substantially different while preserving every top-five cosine neighbour.&lt;/p&gt;

&lt;p&gt;I did not start with that as a teaching point. I had been playing with embeddings while learning about AI, first by using them to modulate sound and then by embedding code and looking for similar code. PCA and UMAP projections seemed pretty cool. At some point the next thought just popped into my head: “Huh, it'd be cool to try and animate embeddings rather than just flatten them.”&lt;/p&gt;

&lt;p&gt;I also thought, “Oh, this is just not going to work.”&lt;/p&gt;

&lt;p&gt;It worked, which surprised me.&lt;/p&gt;

&lt;h2&gt;
  
  
  The changing picture
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://billiem.uk/posts/embedding-tours/interactives/embedding-tours/" rel="noopener noreferrer"&gt;working Embedding Tours instrument&lt;/a&gt; has a deterministic eight-dimensional synthetic dataset and 180 short phrases embedded into 384 dimensions with MiniLM. You can move through raw coordinate pairs, pairs of principal components, or a seeded grand tour through projection planes that are not aligned to the raw axes.&lt;/p&gt;

&lt;p&gt;Those are not three names for the same thing. The raw tour cycles through coordinate pairs. PCA puts the highest-variance directions first, although that does not make them the true semantic axes. The grand tour moves through more general planes.&lt;/p&gt;

&lt;p&gt;Every frame is still a two-dimensional projection. Motion gives me more partial views; it does not recover all the information lost when hundreds of dimensions are put on a flat screen. It is also just really satisfying to watch the structure move.&lt;/p&gt;

&lt;p&gt;The raw-coordinate mode adds a more pointed comparison. Its basis switch applies the same orthogonal transformation to the whole dataset. Individual coordinates change and the raw plot can look very different. Dot products, lengths, distances and cosines do not change under that shared rotation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9fqq3ydquqpfln8kyam4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9fqq3ydquqpfln8kyam4.png" alt="Embedding Tours in the orthogonally rotated basis, with the plot changed and readouts showing 5.55e-16 similarity drift and 100% top-five neighbours retained." width="799" height="521"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The raw projection changes after the basis switch; cosine-neighbour geometry remains fixed to floating-point precision.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In this run, the measured similarity drift is &lt;code&gt;5.55e-16&lt;/code&gt; and all top-five neighbours are retained. That is visible evidence for the narrow claim I wanted the switch to make. It does not prove that every transformation preserves an embedding, and it does not mean individual axes can never be useful. It shows that the cosine-neighbour structure does not uniquely privilege the raw coordinates I happened to receive from the model.&lt;/p&gt;

&lt;p&gt;That limitation matters when interpreting embedding coordinates. A very different-looking coordinate tour can represent the same neighbour geometry.&lt;/p&gt;

&lt;h2&gt;
  
  
  I had walked into an existing idea
&lt;/h2&gt;

&lt;p&gt;My first version just moved from dimensions 1 and 2 towards dimensions 2 and 3, then kept going. I later learned that this belongs to an established family of visualisations called tours.&lt;/p&gt;

&lt;p&gt;Daniel Asimov described the &lt;a href="https://doi.org/10.1137/0906011" rel="noopener noreferrer"&gt;grand tour in 1985&lt;/a&gt;. The &lt;a href="https://ggobi.github.io/tourr/articles/intro.html" rel="noopener noreferrer"&gt;&lt;code&gt;tourr&lt;/code&gt; project&lt;/a&gt; gives the useful distinction: cycling through axis-aligned views is a little tour, while a grand tour moves through general projection planes. &lt;a href="https://distill.pub/2020/grand-tour/" rel="noopener noreferrer"&gt;Distill used a grand tour for neural-network activations&lt;/a&gt;, and the recent &lt;a href="https://arxiv.org/abs/2605.04306" rel="noopener noreferrer"&gt;dtour project&lt;/a&gt; has a much more complete browser interface for steering through high-dimensional data.&lt;/p&gt;

&lt;p&gt;I had not invented a new visualisation technique. I had kind of independently arrived at a known idea, or at least the entrance to one. That was probably the coolest part.&lt;/p&gt;

&lt;p&gt;The instrument stayed brief: two datasets, three projection modes, the basis comparison, selection and playback controls. I do not expect a researcher to discover a new technique in it. Maybe it helps someone like me who is still learning and finds embeddings confusing. Maybe it does not. Maybe it is just a cool experiment.&lt;/p&gt;




&lt;p&gt;This article was adapted with AI assistance from &lt;a href="https://billiem.uk/posts/embedding-tours/" rel="noopener noreferrer"&gt;an original article on billiem.uk&lt;/a&gt;. The original article was reviewed before publication.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>datavisualization</category>
      <category>embeddings</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
