<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Induwara Ashinsana</title>
    <description>The latest articles on DEV Community by Induwara Ashinsana (@induwara_ashinsana_9e4d5b).</description>
    <link>https://dev.to/induwara_ashinsana_9e4d5b</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3958655%2F6fa1c062-e3e8-4949-affc-f60cccfc2dfb.jpg</url>
      <title>DEV Community: Induwara Ashinsana</title>
      <link>https://dev.to/induwara_ashinsana_9e4d5b</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/induwara_ashinsana_9e4d5b"/>
    <language>en</language>
    <item>
      <title>OpenAI's data center chief left. Why builders should care</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Wed, 26 Aug 2026 08:36:04 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/openais-data-center-chief-left-why-builders-should-care-3l3n</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/openais-data-center-chief-left-why-builders-should-care-3l3n</guid>
      <description>&lt;p&gt;OpenAI lost its top data center exec last week, and unlike most executive departures at a frontier lab, this one sits close to something you actually depend on: capacity. &lt;strong&gt;Chris Malone&lt;/strong&gt;, who ran data centers there for roughly 16 months, is out, according to &lt;a href="https://techcrunch.com/2026/08/25/openai-loses-a-top-data-center-exec-as-stream-of-high-profile-departures-continues/" rel="noopener noreferrer"&gt;TechCrunch's August 25 report&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I don't care much about who sits where on OpenAI's org chart. I care that the people responsible for pouring concrete are churning while every AI product I build assumes that concrete gets poured on schedule.&lt;/p&gt;




&lt;h2&gt;
  
  
  🏗️ The bottleneck stopped being the model a while ago
&lt;/h2&gt;

&lt;p&gt;Malone's background says a lot about what the job actually is. Per TechCrunch, he spent nearly five years at &lt;strong&gt;Meta&lt;/strong&gt; and over ten at &lt;strong&gt;Google&lt;/strong&gt; before joining OpenAI in March 2025. That is a career in physical plant: power, land, cooling, supply chain. It is not a research CV.&lt;/p&gt;

&lt;p&gt;Look at who else is named in that infrastructure group:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Person&lt;/th&gt;
&lt;th&gt;Reported role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Sachin Katti&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;VP now leading the infrastructure group&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Uday Ruddarraju&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Data center team lead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Brent Mayo&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Build and delivery program lead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spas Lazarov&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Data center engineering lead&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Before Malone left, his reporting line was moved off &lt;strong&gt;President Greg Brockman&lt;/strong&gt; and onto Katti. That is a reorg of the group that builds the buildings, at a company whose headline constraint is how many buildings it has.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; The limiting factor on your AI product in 2026 is not model quality. It is whether someone finished a substation in Texas on time.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  📉 Thirteen departures is a schedule signal, not gossip
&lt;/h2&gt;

&lt;p&gt;TechCrunch counts &lt;strong&gt;13 senior departures at OpenAI in 2026&lt;/strong&gt;. The pattern is what interests me, not any single exit:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Who&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;When they left&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Chris Malone&lt;/td&gt;
&lt;td&gt;Head of data centers&lt;/td&gt;
&lt;td&gt;August 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Denise Dresser&lt;/td&gt;
&lt;td&gt;Chief revenue officer&lt;/td&gt;
&lt;td&gt;August 2026, after 8 months&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Brad Lightcap&lt;/td&gt;
&lt;td&gt;Chief operating officer&lt;/td&gt;
&lt;td&gt;Early August 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fidji Simo&lt;/td&gt;
&lt;td&gt;Product and business chief&lt;/td&gt;
&lt;td&gt;July 2026, stayed on as advisor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chloé Bakalar&lt;/td&gt;
&lt;td&gt;Head of ethics&lt;/td&gt;
&lt;td&gt;July 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bill Peebles&lt;/td&gt;
&lt;td&gt;Head of Sora&lt;/td&gt;
&lt;td&gt;April 2026, on shutdown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kate Rouch&lt;/td&gt;
&lt;td&gt;Chief marketing officer&lt;/td&gt;
&lt;td&gt;April 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two more data points from the same report: the &lt;strong&gt;preparedness team&lt;/strong&gt; that assessed catastrophic risk was disbanded, and the &lt;strong&gt;IPO slipped from 2026 to 2027&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;You cannot read intent out of any of this, and I won't pretend to. What you can read is variance. Revenue leadership, operations leadership, and infrastructure leadership all turned over inside a few months at the company whose roadmap half the industry has quietly built its 2027 plans around, including the multi-partner &lt;strong&gt;Stargate&lt;/strong&gt; build-out with Oracle, Nvidia, SoftBank and Microsoft.&lt;/p&gt;

&lt;p&gt;When leadership churns, plans get re-litigated. Re-litigated plans slip. Slipped capacity plans show up in your app as rate limits, waitlists, region gaps, and price changes.&lt;/p&gt;




&lt;h2&gt;
  
  
  🌐 What that actually feels like from Colombo
&lt;/h2&gt;

&lt;p&gt;If you build from Sri Lanka, you are already at the thin end of every capacity decision. You feel provider strain earlier and harder than a team in San Francisco does, in four specific ways:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Rate limits tighten from the bottom up.&lt;/strong&gt; Low-spend accounts get squeezed first. If your org is on a starter tier, you are the shock absorber.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;New models land regionally.&lt;/strong&gt; Preview access, cheaper tiers, and batch endpoints often reach some regions late or not at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency has no floor you control.&lt;/strong&gt; Round trips already cost you a few hundred milliseconds. Congestion in a strained region adds more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Price changes hit harder in LKR.&lt;/strong&gt; A 20% list-price rise is a 20% rise plus whatever the rupee did that quarter. Card limits and forex controls do the rest.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;None of this requires OpenAI to have a bad year. It only requires them to have a &lt;em&gt;busy&lt;/em&gt; one.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🛠️ Build so an org chart can't break your product
&lt;/h2&gt;

&lt;p&gt;I use one rule for this: &lt;strong&gt;no provider-specific code above the adapter layer.&lt;/strong&gt; Everything else follows from it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Route through one interface.&lt;/strong&gt; One &lt;code&gt;chat()&lt;/code&gt; function in your codebase, provider chosen by env var. If swapping providers means touching 40 files, you don't have a fallback, you have a hope.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep a real eval set.&lt;/strong&gt; Twenty to fifty prompts with expected outputs, specific to your product. Without it you cannot answer "is the cheaper model good enough for us" in under a week.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail over on 429 and 5xx, not just on outages.&lt;/strong&gt; Rate limiting is the failure mode you will actually hit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cache aggressively.&lt;/strong&gt; Identical prompts in a Sinhala FAQ bot should hit your database, not a GPU in another hemisphere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Know your unit cost before you need to.&lt;/strong&gt; Cost per user, per conversation, per document.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// The whole portability story, roughly&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;providers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;gemini&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Msg&lt;/span&gt;&lt;span class="p"&gt;[])&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;providers&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;isRetryable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;all providers exhausted&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three tools on this site exist for exactly this planning work, and I built them because I needed them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://induwara.lk/tools/ai-llm-api-price-comparison" rel="noopener noreferrer"&gt;AI API price comparison&lt;/a&gt; — per-million-token list prices side by side, so a swap is a number, not a guess.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://induwara.lk/tools/ai-model-deprecation-tracker" rel="noopener noreferrer"&gt;Model deprecation tracker&lt;/a&gt; — which model IDs are on the retirement list before your build breaks.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://induwara.lk/tools/ai-openai-to-anthropic-converter" rel="noopener noreferrer"&gt;OpenAI to Anthropic converter&lt;/a&gt; — request-shape translation when you need a second provider working this afternoon.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;One executive leaving one company is a small event. The reason I wrote about it is that it lands on the exact seam where AI stops being software and starts being civil engineering, and that seam is where small teams get hurt without ever seeing the cause.&lt;/p&gt;

&lt;p&gt;Practical version, three things to do this month:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Add a second provider.&lt;/strong&gt; Not a migration. A tested fallback path and an API key that works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write down your worst-case cost.&lt;/strong&gt; What happens to your margin if inference prices rise 30%? If you cannot answer in ten minutes, model it once and keep the sheet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stop treating roadmaps as commitments.&lt;/strong&gt; Announced capacity, announced pricing tiers, announced launch quarters. Build for what ships, not what is promised.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; You cannot control who runs OpenAI's data centers. You can control how many of your assumptions depend on them.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The teams that will be fine through the next two years of AI infrastructure turbulence are not the ones who picked the right provider. They are the ones who stayed cheap to switch.&lt;/p&gt;

</description>
      <category>openai</category>
      <category>aiinfrastructure</category>
      <category>developerstrategy</category>
    </item>
    <item>
      <title>Thomson Reuters built a frontier model for $40M</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Tue, 25 Aug 2026 12:31:27 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/thomson-reuters-built-a-frontier-model-for-40m-4j95</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/thomson-reuters-built-a-frontier-model-for-40m-4j95</guid>
      <description>&lt;p&gt;The &lt;strong&gt;Thomson Reuters frontier model&lt;/strong&gt; announcement landed on 24 August 2026, and the number that matters is not a benchmark score. It's &lt;strong&gt;$40 million&lt;/strong&gt; — talent and compute combined — to take an open-source base model and specialise it until, by the company's own evaluation, it sits alongside the latest frontier models on the work it was built for.&lt;/p&gt;

&lt;p&gt;What makes this one worth reading is the admission in &lt;a href="https://www.thomsonreuters.com/en/press-releases/2026/august/thomson-reuters-leverages-its-world-class-data-assets-to-launch-its-own-frontier-model" rel="noopener noreferrer"&gt;Thomson Reuters' press release&lt;/a&gt;: they did not start from scratch.&lt;/p&gt;




&lt;h2&gt;
  
  
  💰 $40 million is the headline, and it's a small number
&lt;/h2&gt;

&lt;p&gt;Frontier labs spend billions. Thomson Reuters spent $40 million and shipped a model called &lt;strong&gt;Thomson&lt;/strong&gt; that it fully owns and runs at what it describes as a fraction of the inference cost of comparable frontier models.&lt;/p&gt;

&lt;p&gt;Here's what the release actually commits to, versus what it carefully doesn't:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Claim&lt;/th&gt;
&lt;th&gt;Stated in the release?&lt;/th&gt;
&lt;th&gt;My read&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;$40M total, covering talent and compute&lt;/td&gt;
&lt;td&gt;✅ Explicit&lt;/td&gt;
&lt;td&gt;Real, and small by frontier standards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Built on "a strong open-source foundation"&lt;/td&gt;
&lt;td&gt;✅ Explicit&lt;/td&gt;
&lt;td&gt;The base model is &lt;strong&gt;never named&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"On par with the latest frontier models"&lt;/td&gt;
&lt;td&gt;✅ CEO quote&lt;/td&gt;
&lt;td&gt;Their own early evals, on their own tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trained on &lt;strong&gt;less than 10%&lt;/strong&gt; of TR's content&lt;/td&gt;
&lt;td&gt;✅ Explicit&lt;/td&gt;
&lt;td&gt;So the ceiling is higher than what shipped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Specific benchmark numbers&lt;/td&gt;
&lt;td&gt;❌ Absent&lt;/td&gt;
&lt;td&gt;Pointed to a technical report instead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Parameter count, pricing, HF repo name&lt;/td&gt;
&lt;td&gt;❌ Absent&lt;/td&gt;
&lt;td&gt;Don't assume any of it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; the moat moved. It is no longer "who can afford the biggest training run." It's "who owns data nobody else can legally train on, and who can afford the experts to shape it."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The release says hundreds of subject matter experts were involved, from the design of the training objectives through to the final evaluations. That is an editorial payroll, not a GPU bill.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 The claim that should interest you most
&lt;/h2&gt;

&lt;p&gt;Buried in the middle is an argument aimed straight at everyone building RAG systems:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The domain-specific gain challenges a common assumption, that the most capable general-purpose models only need access to the right content to perform at an expert level."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In plain terms: &lt;strong&gt;giving GPT-class models your documents at inference time is not the same as training on them.&lt;/strong&gt; Thomson Reuters says its post-training produced gains that content access alone did not.&lt;/p&gt;

&lt;p&gt;Treat that as a hypothesis, not a result. It comes from the vendor selling the specialised model, and no head-to-head numbers are published. But it is cheap to test at small scale:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Build your baseline: a strong general model + retrieval over your corpus.&lt;/li&gt;
&lt;li&gt;Score it on 50–100 hand-written questions from your actual domain, graded by someone who knows the answers.&lt;/li&gt;
&lt;li&gt;Fine-tune a small open model on the same corpus.&lt;/li&gt;
&lt;li&gt;Score again with the same rubric, same grader.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If step 4 beats step 2, you've reproduced their argument for the price of a weekend. Our &lt;a href="https://induwara.lk/tools/ai-fine-tuning-cost-calculator" rel="noopener noreferrer"&gt;AI fine-tuning cost calculator&lt;/a&gt; gives you the token-and-epoch maths before you rent a GPU.&lt;/p&gt;




&lt;h2&gt;
  
  
  🌐 What this pattern looks like from Sri Lanka
&lt;/h2&gt;

&lt;p&gt;$40 million is not a Sri Lankan budget. The &lt;em&gt;pattern&lt;/em&gt; scales down, and the pattern is what's copyable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Start from open weights.&lt;/strong&gt; Nobody at the frontier is pretraining from zero for a vertical any more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Own data that isn't on the open web.&lt;/strong&gt; This is the real barrier to entry.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pay for domain experts, not just GPUs.&lt;/strong&gt; Their spend was talent &lt;em&gt;and&lt;/em&gt; compute, in that order in the sentence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deploy narrowly first.&lt;/strong&gt; Thomson's first production surface is one feature — Tabular Analysis inside CoCounsel Legal — not a general chatbot.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Sri Lanka has exactly the kind of data this argument favours, and almost none of it is well-represented in a general model's training set:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Corpus&lt;/th&gt;
&lt;th&gt;Why a general model is weak on it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;IRD rulings, circulars, gazette notices&lt;/td&gt;
&lt;td&gt;Thin on the open web, changes yearly, PDF-locked&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sri Lankan case law and Supreme Court judgments&lt;/td&gt;
&lt;td&gt;Poorly digitised, rarely scraped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CBSL circulars and banking directions&lt;/td&gt;
&lt;td&gt;Scattered across PDFs, no clean corpus&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sinhala and Tamil professional writing&lt;/td&gt;
&lt;td&gt;Chronically under-represented in pretraining data&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A model that genuinely understands Sri Lankan tax practice is not a $40 million project. It's a curated dataset, a small open base, and accountants who will tell you when the output is wrong. Their time is the expensive ingredient, and that is the honest lesson here.&lt;/p&gt;




&lt;h2&gt;
  
  
  ⚠️ The licence trap in the free version
&lt;/h2&gt;

&lt;p&gt;Thomson Reuters is releasing a &lt;strong&gt;"small" version of Thomson as an open-weight model on Hugging Face&lt;/strong&gt; — and this is where I'd slow down before getting excited.&lt;/p&gt;

&lt;p&gt;The release says it is for &lt;strong&gt;academic and non-commercial use&lt;/strong&gt;, published to help external parties validate the model. That is not the same as open source.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning:&lt;/strong&gt; "open-weight" and "you may ship this in your product" are unrelated statements. A non-commercial licence means a student project is fine and a paid client deliverable is not.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Two things the release does &lt;em&gt;not&lt;/em&gt; tell you, and you should not guess at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Which open-source model it was built on.&lt;/strong&gt; Only "a strong open-source foundation" is stated. If the base carries its own licence terms, those may travel downstream.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The exact licence text on the Hugging Face release.&lt;/strong&gt; Read it on the model card, not in a blog post.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're unsure how a given model licence maps to commercial use, our &lt;a href="https://induwara.lk/tools/ai-llm-license-checker" rel="noopener noreferrer"&gt;LLM licence checker&lt;/a&gt; covers the common families, and the &lt;a href="https://induwara.lk/tools/ai-self-hosting-cost-calculator" rel="noopener noreferrer"&gt;self-hosting cost calculator&lt;/a&gt; tells you whether running weights yourself is cheaper than an API in the first place.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;If you're a student or a small-team builder here, three concrete things come out of this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Stop treating "train a model" as a billionaire's activity.&lt;/strong&gt; The gap between "call an API" and "own a model" now has a middle: post-training an open base on data you control. $40 million was the enterprise price. The technique is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit what data you already have access to.&lt;/strong&gt; A university lab, a law firm, a hospital, a small accounting practice — each is sitting on a corpus with no public equivalent. That, not compute, is the thing that's hard to buy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify before repeating.&lt;/strong&gt; No benchmark numbers accompanied this launch, only a pointer to a technical report and two named academic testers. Both were positive; one described citation quality as "generally competitive" with leading models, which is a careful phrase, not a victory lap.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The honest summary: a company with decades of proprietary content and its own editorial staff proved that owning the data and the experts can substitute for owning the largest training cluster. That's good news for anyone building in a small market with deep local knowledge, which describes most useful software work in Sri Lanka.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Commentary based on Thomson Reuters' press release of 24 August 2026. All figures and quotes are from that source; nothing here is independently verified.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>opensourceai</category>
      <category>domainspecificai</category>
    </item>
    <item>
      <title>Training AI on Copyrighted Books: It's the Download That Bites</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Mon, 24 Aug 2026 16:16:44 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/training-ai-on-copyrighted-books-its-the-download-that-bites-5e4b</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/training-ai-on-copyrighted-books-its-the-download-that-bites-5e4b</guid>
      <description>&lt;p&gt;Is it legal to train AI models on copyrighted books? After the Anthropic case, the honest answer is that the &lt;strong&gt;training&lt;/strong&gt; was ruled lawful and the &lt;strong&gt;downloading&lt;/strong&gt; was what cost $1.5 billion. That split is the entire story, and most coverage flattened it into "AI companies lose."&lt;/p&gt;

&lt;p&gt;TechCrunch's piece &lt;a href="https://techcrunch.com/2026/08/23/is-it-legal-to-train-ai-models-on-copyrighted-books-its-complicated/" rel="noopener noreferrer"&gt;Is it legal to train AI models on copyrighted books? It's complicated&lt;/a&gt; lays out the cases. I want to talk about what they mean if you are building something small here with a laptop, a free tier, and a scraper.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 The ruling everyone read backwards
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Judge William Alsup&lt;/strong&gt; ordered Anthropic to pay a $1.5 billion copyright settlement to writers whose books were used in training. The part that got lost: Alsup ruled the training itself was &lt;em&gt;lawful&lt;/em&gt;. He compared an LLM learning from books to a writer studying literature. The penalty attached to something else entirely — Anthropic sourced books from illegal shadow libraries.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Case&lt;/th&gt;
&lt;th&gt;Decided by&lt;/th&gt;
&lt;th&gt;What actually happened&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic copyright settlement&lt;/td&gt;
&lt;td&gt;Judge William Alsup&lt;/td&gt;
&lt;td&gt;$1.5B settlement, but &lt;strong&gt;training ruled lawful&lt;/strong&gt;; the liability came from pirated sourcing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thomson Reuters v. Ross Intelligence&lt;/td&gt;
&lt;td&gt;Judge Stephanos Bibas&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Not fair use&lt;/strong&gt; — Ross trained on Reuters content to build a directly competing legal platform&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thaler v. Perlmutter&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Works that are &lt;strong&gt;100% AI-generated are not copyrightable&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; Your legal exposure is not in the model architecture. It's in your data supply chain — where the bytes came from, and whether you can prove it.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  ⚖️ What a court is actually weighing
&lt;/h2&gt;

&lt;p&gt;Copyright law in the US hasn't been meaningfully updated since &lt;strong&gt;1976&lt;/strong&gt;, which is why every one of these questions ends up decided case by case. The factors judges weigh, per the TechCrunch piece, include the purpose and nature of the work, the amount used, and the effect on the market — with market impact carrying the most weight.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Factor&lt;/th&gt;
&lt;th&gt;The question it asks&lt;/th&gt;
&lt;th&gt;Where small teams get caught&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Purpose and nature&lt;/td&gt;
&lt;td&gt;Is the use transformative, or just a repackage?&lt;/td&gt;
&lt;td&gt;Wrapping someone's content in a thin UI is not transformative&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amount used&lt;/td&gt;
&lt;td&gt;How much of the original was taken?&lt;/td&gt;
&lt;td&gt;Full-corpus scrapes are hard to argue down&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Market impact&lt;/td&gt;
&lt;td&gt;Does it substitute for the original?&lt;/td&gt;
&lt;td&gt;Building the thing your data source sells&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;IP attorney &lt;strong&gt;Cathy Gellis&lt;/strong&gt; put the mechanism plainly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Copyright law hinges on copying, but it doesn't hinge on using the work or experiencing the work, consuming the work, reading the work."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is why Alsup could bless the training and still punish the acquisition. Reading is not the infringement. Obtaining an illegal copy is.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ The provenance habit that costs you nothing today
&lt;/h2&gt;

&lt;p&gt;The practical lesson for anyone fine-tuning a model, building a RAG index, or shipping a dataset: &lt;strong&gt;log where every file came from, at the moment you get it.&lt;/strong&gt; Reconstructing that two years later, under pressure, is impossible.&lt;/p&gt;

&lt;p&gt;A manifest entry per source costs about thirty seconds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.gov.lk/reports/2025-annual.pdf"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"obtained"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-08-24"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"method"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"direct download, public URL"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"licence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Sri Lanka government publication, no stated restriction"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"sha256"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"a3f1..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"robots_txt_checked"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four rules I'd hold anyone to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Never pull from a shadow library.&lt;/strong&gt; This is the single fact pattern that produced a nine-figure bill. Free is not the same as legal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Record a checksum per file&lt;/strong&gt; so you can prove which version you trained on. Our &lt;a href="https://induwara.lk/tools/hash-generator" rel="noopener noreferrer"&gt;hash generator&lt;/a&gt; does SHA-256 in the browser — nothing gets uploaded.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prefer licensed and public-domain corpora.&lt;/strong&gt; Project Gutenberg, government publications, Creative Commons, and datasets with an explicit licence field.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check &lt;code&gt;robots.txt&lt;/code&gt; and terms before scraping&lt;/strong&gt;, and write down the date you checked.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning:&lt;/strong&gt; "I got it from a HuggingFace mirror" is not provenance. If the upstream dataset was assembled from pirated books, inheriting it does not launder it.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  📄 The other half nobody plans for: what you can sell
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Thaler v. Perlmutter&lt;/strong&gt; held that fully AI-generated works are not copyrightable. If you're freelancing — logos, marketing copy, landing pages, boilerplate code — that's not academic. It's the question of what your client is actually paying for.&lt;/p&gt;

&lt;p&gt;The awkward follow-on, which the TechCrunch piece flags directly: nobody has settled how you'd &lt;em&gt;prove&lt;/em&gt; how much AI was involved, or what percentage of human input makes a work protectable.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deliverable&lt;/th&gt;
&lt;th&gt;Human contribution&lt;/th&gt;
&lt;th&gt;Practical risk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Prompt in, image out, shipped as-is&lt;/td&gt;
&lt;td&gt;Near zero&lt;/td&gt;
&lt;td&gt;Client may own nothing enforceable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI draft, substantially edited and art-directed by you&lt;/td&gt;
&lt;td&gt;Meaningful&lt;/td&gt;
&lt;td&gt;Much stronger position&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI-assisted code you architected, reviewed, and tested&lt;/td&gt;
&lt;td&gt;Meaningful&lt;/td&gt;
&lt;td&gt;Standard practice, low concern&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;My read: keep your working files. Drafts, revision history, the notes where you rejected three versions. That record is the evidence of human authorship, and it's free to keep.&lt;/p&gt;




&lt;h2&gt;
  
  
  🌐 Why US rulings matter from Colombo
&lt;/h2&gt;

&lt;p&gt;None of these decisions bind a Sri Lankan court. I'm an engineer, not a lawyer, and I'm not going to pretend otherwise. But they matter here anyway, for reasons that have nothing to do with jurisdiction:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your clients are often abroad.&lt;/strong&gt; A US or EU company commissioning work will push provenance and IP-warranty terms down to you in the contract.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The vendors you build on are US companies.&lt;/strong&gt; What Alsup and Bibas decide shapes what OpenAI, Anthropic and Google will and won't ship, and what their terms of service allow you to do downstream.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Licensed data is becoming a paid market.&lt;/strong&gt; Attorney &lt;strong&gt;Jason Henderson&lt;/strong&gt; framed it as: &lt;em&gt;"Copyright is always about protecting and growing the market."&lt;/em&gt; Expect more datasets to carry a price tag rather than a takedown notice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Competing with your source is the risky shape.&lt;/strong&gt; Ross lost because it trained on Reuters content to build a Reuters competitor. If you fine-tune on a company's data to sell against that company, fair use gets much harder to argue.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;If you're building on a learning budget, the good news is that the expensive mistake in this story is also the easiest one to avoid. Alsup didn't say training on books is illegal. He said helping yourself to pirated copies is.&lt;/p&gt;

&lt;p&gt;So: keep a manifest, checksum your sources, prefer licensed and public-domain corpora, and don't fine-tune a model to compete head-on with the people whose data you used. Keep your drafts so you can show human authorship in what you deliver.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; Train on what you're allowed to have, and keep the receipt.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;None of that requires a lawyer or a budget. It requires a habit, and the cheapest time to start it is on the project you haven't collected data for yet.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>copyright</category>
      <category>developers</category>
    </item>
    <item>
      <title>Claude Code effort levels: what the A/B test really shows</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Sun, 23 Aug 2026 19:52:31 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/claude-code-effort-levels-what-the-ab-test-really-shows-3o48</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/claude-code-effort-levels-what-the-ab-test-really-shows-3o48</guid>
      <description>&lt;p&gt;&lt;strong&gt;Claude Code effort levels&lt;/strong&gt; are supposed to be a dial you control, so it matters that people have started asking whether the level you pick is the level you actually get. A &lt;a href="https://x.com/argofowl/status/2091150597374537729" rel="noopener noreferrer"&gt;post by &lt;strong&gt;@argofowl&lt;/strong&gt; on X&lt;/a&gt;, &lt;a href="https://news.ycombinator.com/item?id=49401549" rel="noopener noreferrer"&gt;surfaced on Hacker News&lt;/a&gt;, claimed Anthropic was quietly reducing selected effort levels while showing users misleading numbers.&lt;/p&gt;

&lt;p&gt;Anthropic replied. The reply is more interesting than the accusation.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 What Anthropic actually said
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Thariq&lt;/strong&gt;, from the Claude Code team, posted the same response on X and in the HN thread:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently. That's why Claude may tell some of you it's at '10' on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Read that carefully. It is not a denial of A/B testing. It is a confirmation of it, plus a claim that this particular test is behaviour-neutral. Both things can be true at once.&lt;/p&gt;

&lt;p&gt;The original evidence was weak, and the thread said so. As commenter &lt;strong&gt;Wowfunhappy&lt;/strong&gt; put it: the proof was that someone &lt;em&gt;asked Claude what effort level it was set to&lt;/em&gt;, and "how would the model even know that?" Others argued effort is injected through the system prompt, so the model might genuinely see it. Nobody in the thread could settle it, which is exactly the problem.&lt;/p&gt;




&lt;h2&gt;
  
  
  📊 The real issue is unverifiable configuration
&lt;/h2&gt;

&lt;p&gt;Strip out the outrage and what's left is a supply-chain question. You are paying for a product whose behaviour is set by three layers, and you can only see one of them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Who controls it&lt;/th&gt;
&lt;th&gt;Can you inspect it?&lt;/th&gt;
&lt;th&gt;Is it versioned?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model weights&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes, by model name&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Harness / system prompt&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;Partially&lt;/td&gt;
&lt;td&gt;Not publicly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;API serving config&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Your prompt, files, CLAUDE.md&lt;/td&gt;
&lt;td&gt;You&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes, in git&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The third row is where this story lives. A serving config can change on a Tuesday afternoon with no release note, no version bump, and no way for you to pin the old one. Commenter &lt;strong&gt;cube00&lt;/strong&gt; asked the sharpest question in the thread: "Why is it considered acceptable to test on paying customers without letting them know or giving them a way to opt out?"&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; "Claude Opus 5, high effort" is not a specification. It's a label pointing at a moving target. Treat model behaviour as an external dependency with no lockfile, because that's what it is.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  💰 Why over-thinking is a budget problem here
&lt;/h2&gt;

&lt;p&gt;The complaint underneath the effort-level story is not that the model got dumber. It's that it got &lt;strong&gt;longer&lt;/strong&gt;, and length costs money.&lt;/p&gt;

&lt;p&gt;One commenter described asking for a config file update: on the older model it took under 2 minutes to read, parse and patch. On the newer one it ran &lt;strong&gt;43 minutes&lt;/strong&gt;, pulling containers, spinning up sandboxes, and writing test suites for the whole repo. Same single-file change at the end.&lt;/p&gt;

&lt;p&gt;That's funny when your employer pays. It's not funny when you're a student or a two-person shop converting a USD subscription into rupees every month. A few things follow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Verbosity is the main cost driver, not model choice.&lt;/strong&gt; One user (&lt;code&gt;perching_aix&lt;/code&gt;) reported that comparing like-for-like, the majority of their cost overhead came from the bigger model simply being chattier, and that dropping to &lt;em&gt;low&lt;/em&gt; reasoning beat falling back to a smaller model on quality-per-rupee.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Higher effort is not free accuracy.&lt;/strong&gt; Several developers in the thread independently landed on the same policy: medium by default, high only for genuinely hard architectural work, never the top tiers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Subscription tiers don't shield you.&lt;/strong&gt; Limits are real. Burn your weekly quota on a task that needed two minutes of thinking and you've paid for it in lost days, not just tokens.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you've never actually measured your prompt sizes, our &lt;a href="https://induwara.lk/tools/ai-token-counter" rel="noopener noreferrer"&gt;AI token counter&lt;/a&gt; will show you what a given context costs against Claude, GPT and Gemini tokenizers before you send it. And if you're budgeting a USD subscription in rupees, the &lt;a href="https://induwara.lk/tools/lkr-exchange-rate" rel="noopener noreferrer"&gt;live LKR exchange rate&lt;/a&gt; is the number that actually decides whether the $200 tier makes sense for you.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ Run your own eval instead of arguing about vibes
&lt;/h2&gt;

&lt;p&gt;Here's my honest position. I run agents against this codebase every day, so I have opinions about regressions, and I've noticed that almost all of them evaporate when I write the test down. Human memory of "it felt better last month" is worthless as evidence. So build a tiny personal benchmark. It takes an afternoon.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pick 5 real tasks from your own repo&lt;/strong&gt; that you already know the correct answer to. Not puzzles. Actual work: a bug fix, a refactor, a migration, a config change, a test.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Freeze the inputs.&lt;/strong&gt; Same branch, same files, same prompt text, committed to git.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log the outputs.&lt;/strong&gt; Wall-clock time, tokens in and out, and a manual pass/fail you write yourself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-run monthly&lt;/strong&gt;, and any time you feel a regression.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use the vendor's channel when you have data.&lt;/strong&gt; Thariq explicitly asked users seeing a clear regression to file &lt;code&gt;/feedback&lt;/code&gt; with the session ID, and offered credits. A session ID beats a complaint.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# crude but sufficient: one row per run, append-only&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; &lt;span class="nt"&gt;-Iseconds&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;,opus5,high,&lt;/span&gt;&lt;span class="nv"&gt;$TASK&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="nv"&gt;$SECONDS&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="nv"&gt;$IN_TOK&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="nv"&gt;$OUT_TOK&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="nv"&gt;$VERDICT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; ~/agent-evals.csv
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Five tasks, five minutes of logging each. That CSV is worth more than every "models are getting dumber" thread combined, because it's about &lt;em&gt;your&lt;/em&gt; code.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;Anthropic's answer here was fast, named, and specific, which is better than the silence you'd get from most vendors. I don't think this particular test is evidence of deliberate degradation, and the accusation that triggered it rested on asking a model to introspect on its own settings, which is not a measurement.&lt;/p&gt;

&lt;p&gt;But the structural point survives the debunking. &lt;strong&gt;The behaviour of a hosted coding agent is a dependency you cannot pin, cannot diff, and cannot roll back.&lt;/strong&gt; That's a genuinely new category of risk for anyone shipping software, and it lands hardest on small teams without the budget to keep a second provider warm.&lt;/p&gt;

&lt;p&gt;So: default to lower effort than you think you need, keep your prompts short because verbosity is where the money goes, hold a small eval suite you actually re-run, and keep one alternative agent configured so switching is a decision and not an emergency. None of that requires believing anyone is lying to you. It just requires treating the dial on your screen as a hint rather than a contract.&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>aicodingagents</category>
      <category>developertooling</category>
    </item>
    <item>
      <title>Engineering maturity isn't years served: 3 belief updates</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Sat, 22 Aug 2026 23:32:29 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/engineering-maturity-isnt-years-served-3-belief-updates-3b8</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/engineering-maturity-isnt-years-served-3-belief-updates-3b8</guid>
      <description>&lt;p&gt;&lt;strong&gt;Engineering maturity&lt;/strong&gt; in Sri Lanka is usually measured in years served and job titles collected. Five years, "Senior". Eight years, "Tech Lead". Nobody checks whether anything in your head actually changed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thomas Dullien&lt;/strong&gt; (&lt;strong&gt;Halvar Flake&lt;/strong&gt;) published &lt;a href="https://thomasdullien.github.io/posts/2026-08-21-three-important-steps-in-my-maturation-process/" rel="noopener noreferrer"&gt;Three important steps in my maturation process&lt;/a&gt; on 21 August 2026. It proposes a different measure: three specific beliefs you had to update. Two of the three land harder here than in Mountain View.&lt;/p&gt;




&lt;h2&gt;
  
  
  🧠 Step 1: your incentive structure is writing your opinions
&lt;/h2&gt;

&lt;p&gt;Dullien's first realisation comes from his 0day years: he agonised over whether holding an exploit nobody else had made him responsible for what followed. His conclusion was that the agonising was partly vanity.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Anxiety about the impact of your work is self-flattering, and you have to recognize it as such."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;He pairs it with von Neumann's line that "some people profess guilt to claim credit for sin." The instruction is blunt: &lt;strong&gt;do not believe everything you think&lt;/strong&gt;, and ask "how might I be the villain in this story?" Translate that to a Colombo engineering career and it stops being philosophy:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;The belief you hold&lt;/th&gt;
&lt;th&gt;The incentive quietly holding it up&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"Our stack is fine, rewriting is a waste"&lt;/td&gt;
&lt;td&gt;You are the only person who knows the stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Remote USD work is the only sensible path"&lt;/td&gt;
&lt;td&gt;Your last local interview went badly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Local companies never pay properly"&lt;/td&gt;
&lt;td&gt;You've never actually negotiated hard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"I should stay 3 more years for the title"&lt;/td&gt;
&lt;td&gt;Leaving means admitting the title was the point&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None of those beliefs are automatically wrong. The point is that you acquired them in an order that flattered you, and never tried the other side.&lt;/p&gt;

&lt;p&gt;The cheapest version of this discipline is arithmetic. Before deciding a job or a rate is unfair, work out the real number instead of the felt number: our &lt;a href="https://induwara.lk/tools" rel="noopener noreferrer"&gt;salary and rate tools&lt;/a&gt; do net-to-gross and freelance hourly maths in seconds. A belief that survives a spreadsheet beats one that only survives your friends.&lt;/p&gt;




&lt;h2&gt;
  
  
  ⚡ Step 2: determinism is a luxury our infrastructure does not grant
&lt;/h2&gt;

&lt;p&gt;This is the section I'd hand to every CS undergraduate in the country. Dullien argues that the &lt;strong&gt;monocausal determinism&lt;/strong&gt; young programmers get used to is an illusion maintained by generations of electrical and process engineers: computing machines are physical devices subject to wear, unit-to-unit variation, and "probabilistically deterministic behavior." Push on temperature, voltage, electromagnetic fields, or repeated accesses to adjacent DRAM rows, he writes, and it collapses.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The real world is one where very few things that happen have a single reason, and very few truly deterministic transmission mechanisms. Everything is probabilistic, and everything is multicausal."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here's my angle. In a well-funded environment that illusion holds for years, because money buys it: dedicated instances, redundant power, generous headroom. On a shared &lt;strong&gt;4-core VPS&lt;/strong&gt; with a few gigabytes of usable RAM, which is what most Sri Lankan side projects run on, the illusion never gets a chance to form.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;The deterministic model&lt;/th&gt;
&lt;th&gt;What actually happens on shared, budget infra&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Build passes or fails on code&lt;/td&gt;
&lt;td&gt;Build gets OOM-killed because something else spiked at the same moment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failed test = real bug&lt;/td&gt;
&lt;td&gt;Failed test = contention, timeout, or a cold cache&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lab performance score = user experience&lt;/td&gt;
&lt;td&gt;Lab score looks good; real field data still fails Core Web Vitals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The deploy caused the outage&lt;/td&gt;
&lt;td&gt;The deploy, plus a health check, plus a redirect loop, together caused it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I've been on the wrong side of every row in that table. The costliest was assuming a failed build meant broken code. It usually meant two heavy jobs had landed in the same minute on the same box, and the fix was scheduling, not debugging.&lt;/p&gt;

&lt;p&gt;His corollary is worth sitting with: the scientific method is a classifier deliberately biased against accepting things as true, so &lt;strong&gt;a large class of true things will never be scientifically demonstrated&lt;/strong&gt;. If you are chasing an intermittent production issue with no reliable reproduction, you are not failing at engineering. You are working where proof is unavailable and probability is all you get.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 Step 3: emotion is information, not interference
&lt;/h2&gt;

&lt;p&gt;The third update gets dismissed fastest by technical readers, which is roughly the evidence for it. Dullien argues the reason-versus-emotion split is a &lt;strong&gt;western cultural construct&lt;/strong&gt;, not a neurological fact, and that in most non-western cultures integrating deliberation with emotion is the normal model.&lt;/p&gt;

&lt;p&gt;His trivia hook: your gut's &lt;strong&gt;enteric nervous system contains as many neurons as the entire cerebral cortex of a dog&lt;/strong&gt;, and your body forward-deploys neurons into muscles and extremities as latency optimisation. That information reaches you as feeling, not sentences. Then the logical argument, which is the part that convinces:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Attempting to eliminate a particular source of information almost certainly makes the quality of your decisions worse."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is an information-theoretic claim, not a self-help one. If a codebase makes your stomach drop every time you open it, that reaction encodes pattern-matching your verbal brain hasn't finished compiling. Same for the client whose emails you keep postponing, or the review where you couldn't articulate the objection but knew there was one.&lt;/p&gt;

&lt;p&gt;Note his caution: this is not an argument for acting on impulse, but for treating the feeling as a signal to investigate rather than noise to suppress.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; Maturity here isn't years on the job. It's noticing your incentives before they finish writing your opinions, dropping the assumption that one cause explains one effect, and treating your gut reaction as unlabelled data rather than a defect.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🛠️ Four things I'd actually change on Monday
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Write down the counter-narrative.&lt;/strong&gt; For your strongest technical opinion this quarter, write the best version of the opposite case. If you can't write a good one, you don't understand your own position.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stop saying "the root cause".&lt;/strong&gt; Say "the largest contributing cause". Running things on cheap infrastructure, I have almost never found a single cause.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log the environment, not just the error.&lt;/strong&gt; Free RAM, load average, and concurrent jobs at failure time. Half of your "bugs" will resolve into contention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat dread as a ticket.&lt;/strong&gt; When part of the codebase or a client relationship consistently makes you avoid it, open an issue for the avoidance itself.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  🚀 What this means for you
&lt;/h2&gt;

&lt;p&gt;If you are a student or a junior engineer here, take the second update seriously first. Coursework and tutorials run inside a deterministic sandbox that production, especially production you can afford, does not provide. Probabilistic thinking is the difference between an engineer who can operate a system and one who can only write code for it.&lt;/p&gt;

&lt;p&gt;If you are further along, the first update is the expensive one. Nobody will tell you that your architectural preference is really a job-security preference, or your rate opinion really a confidence problem. You catch that yourself, and the only tool that works is arguing against yourself in writing.&lt;/p&gt;

&lt;p&gt;The third is free. It costs nothing to stop treating your own discomfort as a bug in your reasoning. It is data. Check it, but don't throw it away.&lt;/p&gt;

&lt;p&gt;The original post is short and has no signup wall. &lt;a href="https://thomasdullien.github.io/posts/2026-08-21-three-important-steps-in-my-maturation-process/" rel="noopener noreferrer"&gt;Read it in full&lt;/a&gt; rather than trusting my summary.&lt;/p&gt;

</description>
      <category>career</category>
      <category>engineeringculture</category>
      <category>opinion</category>
    </item>
    <item>
      <title>Nvidia is buying power, not just selling GPUs</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Sat, 22 Aug 2026 03:05:16 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/nvidia-is-buying-power-not-just-selling-gpus-55l8</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/nvidia-is-buying-power-not-just-selling-gpus-55l8</guid>
      <description>&lt;p&gt;The &lt;strong&gt;Nvidia Cloverleaf data center partnership&lt;/strong&gt; announced on Friday tells you where the real constraint in AI has moved, and it is not the chip. &lt;a href="https://techcrunch.com/2026/08/21/nvidia-partners-with-data-center-developer-cloverleaf/" rel="noopener noreferrer"&gt;TechCrunch reported&lt;/a&gt; that Nvidia has taken a minority stake in &lt;strong&gt;Cloverleaf&lt;/strong&gt;, a company founded in &lt;strong&gt;2024&lt;/strong&gt; that raised &lt;strong&gt;$300 million&lt;/strong&gt; that year and sits between utility companies and data centers, arranging power and site infrastructure.&lt;/p&gt;

&lt;p&gt;Nvidia did not buy a chip designer. It bought a piece of the electricity supply chain. That reframing is worth thinking about if you build software from anywhere outside a well-supplied grid.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔌 The bottleneck moved from silicon to substations
&lt;/h2&gt;

&lt;p&gt;Read the deal literally. Cloverleaf's product is not compute. It is &lt;strong&gt;power sourcing and site infrastructure&lt;/strong&gt; — the interconnect agreements, the substations, the land next to a utility that can actually deliver load. Nvidia buying into that layer is an admission that shipping more GPUs does not help if nobody can plug them in.&lt;/p&gt;

&lt;p&gt;This was not a one-off either. The same week, per the reporting:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deal&lt;/th&gt;
&lt;th&gt;Announced&lt;/th&gt;
&lt;th&gt;Reported size&lt;/th&gt;
&lt;th&gt;What Nvidia bought into&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;SB Energy&lt;/strong&gt; (OpenAI-linked, Ohio)&lt;/td&gt;
&lt;td&gt;17 Aug 2026&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$1.5 billion&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Data center project with an energy parent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cloverleaf&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;21 Aug 2026&lt;/td&gt;
&lt;td&gt;Several hundred million (WSJ), minority stake (Reuters)&lt;/td&gt;
&lt;td&gt;Utility-to-data-center power intermediary&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Terms were not disclosed by either company, so treat the dollar figures as press reporting rather than filings.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; When the company that sells the shovels starts buying the mines, the scarce input is no longer shovels. For AI in 2026, the scarce input is grid capacity.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  💰 The circular money problem you should price in
&lt;/h2&gt;

&lt;p&gt;Here is the part I would not skip over. Nvidia sells GPUs to data centers. Nvidia is now also &lt;strong&gt;investing in the companies that build and power those data centers&lt;/strong&gt;. Some of that capital flows back as GPU orders.&lt;/p&gt;

&lt;p&gt;That is not illegal or even unusual in capital-intensive industries. Telecom vendors financed carriers for decades. But it has a specific consequence for you as a buyer of compute:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Capacity gets built ahead of demand&lt;/strong&gt;, because the supplier has a reason to underwrite it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rental prices can stay soft&lt;/strong&gt; while that overbuild is absorbed, which is good for small teams.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prices can also snap back&lt;/strong&gt; if the financing loop tightens, because the underwriting was never a market signal in the first place.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the cheap H100-hour you rent today is partly a subsidy artefact. Plan for it to be temporary, and do not architect anything that only works at today's price.&lt;/p&gt;




&lt;h2&gt;
  
  
  📊 Why this argues against buying hardware from Sri Lanka
&lt;/h2&gt;

&lt;p&gt;Every few months someone asks me whether they should import a used GPU and self-host a model. The Cloverleaf story is a decent argument for "no", and the reason is the same one Nvidia just paid for: &lt;strong&gt;power is the real line item&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A rented GPU-hour bundles things you cannot buy separately at small scale:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you actually need&lt;/th&gt;
&lt;th&gt;Rented cloud GPU&lt;/th&gt;
&lt;th&gt;Self-hosted box in Sri Lanka&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Electricity at industrial rates&lt;/td&gt;
&lt;td&gt;Included in the hourly price&lt;/td&gt;
&lt;td&gt;Domestic tariff, on your bill&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cooling&lt;/td&gt;
&lt;td&gt;Included&lt;/td&gt;
&lt;td&gt;Your problem, year-round&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Uptime through a grid fault&lt;/td&gt;
&lt;td&gt;Included&lt;/td&gt;
&lt;td&gt;UPS, and hope&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Import duty and forex exposure&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Paid up front, in USD&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ability to stop paying&lt;/td&gt;
&lt;td&gt;Instantly&lt;/td&gt;
&lt;td&gt;Never, you own it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I am not going to quote a break-even number here because it depends entirely on your utilisation and your tariff block. Run it yourself: our &lt;a href="https://induwara.lk/tools/ai-gpu-buy-vs-rent-calculator" rel="noopener noreferrer"&gt;GPU buy vs rent calculator&lt;/a&gt; handles the hardware side, and the &lt;a href="https://induwara.lk/tools/sri-lanka-electricity-bill-calculator" rel="noopener noreferrer"&gt;Sri Lanka electricity bill calculator&lt;/a&gt; will tell you what the extra load does to your monthly block.&lt;/p&gt;

&lt;p&gt;The short version: self-hosting wins only at high, sustained utilisation. Most small teams run their GPU at single-digit percent duty and pay for the other 90-plus percent of the month in idle electricity.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ What I'd actually do with this information
&lt;/h2&gt;

&lt;p&gt;Nothing about this deal changes your code today. It changes what you should assume about prices over the next two years. Concretely:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Treat compute as a variable cost, not an asset.&lt;/strong&gt; Rent. Keep the ability to walk away when a cheaper region or provider appears.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure cost per unit of work, not per hour.&lt;/strong&gt; An hour is meaningless. Tokens served, documents embedded, minutes transcribed — those are the units you can actually optimise. The &lt;a href="https://induwara.lk/tools/ai-rag-cost-calculator" rel="noopener noreferrer"&gt;RAG cost calculator&lt;/a&gt; is built around exactly that framing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Right-size the model before renting a bigger card.&lt;/strong&gt; Most projects I see are renting for a model that would fit in less VRAM with sane quantisation. Check with the &lt;a href="https://induwara.lk/tools/ai-llm-vram-calculator" rel="noopener noreferrer"&gt;LLM VRAM calculator&lt;/a&gt; before upgrading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep your inference layer provider-agnostic.&lt;/strong&gt; One interface, swappable backends. If the subsidised era ends, you migrate in a day instead of a quarter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do the boring optimisations first.&lt;/strong&gt; Caching, batching and a smaller model beat any hardware decision, and they cost nothing but attention.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; The infrastructure arms race is happening at a scale you will never participate in, and that is fine. Your advantage was never capital. It is that you can change your architecture in an afternoon and they cannot.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;If you are a student or a small-team builder in Sri Lanka, the Cloverleaf deal is not a story about Nvidia's balance sheet. It is a signal about three things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cheap compute is currently being underwritten&lt;/strong&gt; by the companies that benefit from it existing. Use it, enjoy it, do not depend on it forever.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The bottleneck is physical.&lt;/strong&gt; Grid interconnect and cooling are slow, permitted, regulated things. That means the compute glut arrives on a construction timetable, not a software one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your leverage is efficiency.&lt;/strong&gt; Nobody in Sri Lanka is going to out-capital a hyperscaler. But a well-cached, well-batched, right-sized pipeline running on rented hardware competes fine on output.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nvidia is buying substations because it has to. You do not have that constraint, and you should be glad about it. Build things that get cheaper when compute does, and that still work when it does not.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Facts in this post come from the TechCrunch report linked above. Deal terms were not disclosed by Nvidia or Cloverleaf; reported figures are attributed to the Wall Street Journal and Reuters via that article.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>nvidia</category>
      <category>cloudcosts</category>
    </item>
    <item>
      <title>OpenAI vs Anthropic: enterprise AI loyalty is near zero</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Fri, 21 Aug 2026 06:38:35 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/openai-vs-anthropic-enterprise-ai-loyalty-is-near-zero-47p3</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/openai-vs-anthropic-enterprise-ai-loyalty-is-near-zero-47p3</guid>
      <description>&lt;p&gt;New spending data on &lt;strong&gt;OpenAI vs Anthropic enterprise market share&lt;/strong&gt; says the two labs are swapping the lead every few months, and that churn is the part worth your attention. TechCrunch reported on &lt;a href="https://techcrunch.com/2026/08/20/openai-is-gaining-on-anthropic-with-business-users-new-data-indicates/" rel="noopener noreferrer"&gt;Ramp's corporate card data&lt;/a&gt; showing Anthropic still ahead but OpenAI closing through Q3.&lt;/p&gt;

&lt;p&gt;Most people will read that as a scoreboard. I read it as evidence that lock-in barely exists right now. If you build software from Sri Lanka on a small budget, that is leverage you should be using.&lt;/p&gt;




&lt;h2&gt;
  
  
  📊 What the data says, and what it doesn't cover
&lt;/h2&gt;

&lt;p&gt;Ramp is a corporate card and expense-management company. It tracks spending across &lt;strong&gt;more than 70,000 American businesses&lt;/strong&gt;, so it sees which AI vendors actually get charged, not which ones get announced.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Data point&lt;/th&gt;
&lt;th&gt;May 2026&lt;/th&gt;
&lt;th&gt;July 2026&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic share of AI spend&lt;/td&gt;
&lt;td&gt;41%&lt;/td&gt;
&lt;td&gt;~44%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI share of AI spend&lt;/td&gt;
&lt;td&gt;39%&lt;/td&gt;
&lt;td&gt;~40%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ramp customers paying for any AI&lt;/td&gt;
&lt;td&gt;~50% (March)&lt;/td&gt;
&lt;td&gt;~56%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Anthropic took the lead in May and has held it. But Q3-to-date growth is running OpenAI's way, which is the whole story behind the headline.&lt;/p&gt;

&lt;p&gt;Before you treat those percentages as gospel, note the limits TechCrunch itself lists:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;American businesses only.&lt;/li&gt;
&lt;li&gt;Excludes large enterprises that run spend management through providers like American Express.&lt;/li&gt;
&lt;li&gt;Covers corporate card and bill-pay spending, not every purchasing route.&lt;/li&gt;
&lt;li&gt;Skews toward tech, because Ramp is popular in Silicon Valley.&lt;/li&gt;
&lt;li&gt;Percentages only. No dollar amounts were shared.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;TechCrunch is blunt about it: "This isn't a measure of the total market." It's a directional read, not a census. Treat it that way.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🔁 Eight weeks is the new switching cost
&lt;/h2&gt;

&lt;p&gt;Look at the shape of the movement rather than the values. A lab ships a model, spend moves. Another lab ships, spend moves back. Ramp economist &lt;strong&gt;Ara Kharazian&lt;/strong&gt; told TechCrunch that &lt;strong&gt;GPT-5.6 Sol&lt;/strong&gt; is "really good, increasingly the choice for developers," while &lt;strong&gt;Fable 5&lt;/strong&gt; "disappointed both in adoption and real-world application."&lt;/p&gt;

&lt;p&gt;That is a market where a two-month-old quality gap rearranges buying decisions. Which means the switch itself has become cheap: same chat-completions shape, same tool-calling pattern, same streaming semantics. Changing providers is a config change and a round of eval runs, not a rewrite.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; Nobody has moat-grade lock-in on the model layer today. Design for that, and every future price cut and capability jump is yours to take. Design against it, and you're volunteering for a rewrite you didn't need.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🛠️ How to build for a market that flips every quarter
&lt;/h2&gt;

&lt;p&gt;Provider-agnostic does not mean a heavyweight abstraction framework. It means being disciplined about four things.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Hardcode to one vendor?&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model IDs&lt;/td&gt;
&lt;td&gt;Never&lt;/td&gt;
&lt;td&gt;Put them in config or env. They change more often than your code.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompts&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Store as data (files or DB rows), not string literals scattered in handlers.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool / function schemas&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Keep one JSON schema definition, map it per provider.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vendor-only features&lt;/td&gt;
&lt;td&gt;Isolate&lt;/td&gt;
&lt;td&gt;Caching, batch, computer use — behind a flag, with a plain fallback path.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A minimal seam is often enough:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;LLM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;opts&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;}):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;LLM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;LLM_PROVIDER&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;openaiAdapter&lt;/span&gt;
  &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;anthropicAdapter&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then three habits that make the seam real:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Keep an eval set.&lt;/strong&gt; Twenty to fifty of your own prompts with expected outputs. Without it, "the new model is better" is a vibe, and vendor benchmarks are marketing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log every call with the model ID.&lt;/strong&gt; When quality shifts, you want to know whether it was your prompt or their silent update.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-test on release, not on renewal.&lt;/strong&gt; Run your evals when a new model drops, then decide with numbers.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  💰 What lab churn does to your bill
&lt;/h2&gt;

&lt;p&gt;This is the part that helps anyone earning in rupees and paying in dollars. Competition this close is what keeps prices falling and free tiers alive. When 56% of a card provider's customer base is already paying for AI and the leaders are four points apart, neither side can afford to price like a monopoly.&lt;/p&gt;

&lt;p&gt;Practical moves for a small team here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Price the workload before you commit. Our &lt;a href="https://induwara.lk/tools/ai-token-counter" rel="noopener noreferrer"&gt;AI token counter&lt;/a&gt; and &lt;a href="https://induwara.lk/tools/ai-model-comparison" rel="noopener noreferrer"&gt;AI model comparison&lt;/a&gt; let you cost a feature before you write it.&lt;/li&gt;
&lt;li&gt;Split by task. Cheap models for classification, extraction, and formatting. Expensive ones only where reasoning quality shows up in the output.&lt;/li&gt;
&lt;li&gt;Re-price quarterly. If the leaderboard moves every two months, your cost assumptions from January are already stale.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  ⚠️ The retention clause matters more than the leaderboard
&lt;/h2&gt;

&lt;p&gt;One detail in the piece is worth more to Sri Lankan dev shops than the market share numbers: Anthropic warned users of its higher-end &lt;strong&gt;Fable&lt;/strong&gt; tier that it must &lt;strong&gt;retain their data for 30 days&lt;/strong&gt; under regulatory mandates. That caused a round of public anger, though the article notes the criticism was somewhat oversimplified, since Fable targets specific use cases rather than general chat.&lt;/p&gt;

&lt;p&gt;If you're doing client work for a bank, a hospital, or a European customer, retention terms are a procurement fact. Before you ship, get answers to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How long is prompt and output data retained, and can that be turned off?&lt;/li&gt;
&lt;li&gt;Does the tier you're on differ from the tier the marketing page describes?&lt;/li&gt;
&lt;li&gt;What does your own client contract promise about sub-processors?&lt;/li&gt;
&lt;li&gt;Can you point to the vendor's written policy, not a blog post?&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;A model that scores two points higher and violates your client's data clause is worth zero. Check the terms before the benchmarks.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Don't pick a "winner."&lt;/strong&gt; The lead changed in May and is being chased in August. Anyone telling you one lab has permanently won is selling something.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build a seam, not a framework.&lt;/strong&gt; One adapter interface, model IDs in config, prompts as data. That's a day of work and it buys you every future price drop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Own your evals.&lt;/strong&gt; Your twenty prompts beat any public benchmark for deciding what actually works on your product.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the retention terms first&lt;/strong&gt;, especially for client work under an NDA.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat volatility as a discount.&lt;/strong&gt; Two well-funded labs fighting over enterprise spend is the best pricing environment a small team in Colombo is going to get. Stay in a position to switch, and you keep collecting the benefit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The investors being warned about weak stickiness are right to worry. For the rest of us building things, weak stickiness is the good news.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>developertools</category>
      <category>srilankatech</category>
    </item>
    <item>
      <title>Waymo's cheaper robotaxi is a cost story, not an AI story</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Thu, 20 Aug 2026 10:34:52 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/waymos-cheaper-robotaxi-is-a-cost-story-not-an-ai-story-d9p</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/waymos-cheaper-robotaxi-is-a-cost-story-not-an-ai-story-d9p</guid>
      <description>&lt;p&gt;Waymo's next-generation robotaxi, the &lt;strong&gt;Ojai&lt;/strong&gt;, is now open to every rider in three US cities, and the part worth studying isn't the driving. It's the bill of materials. &lt;a href="https://techcrunch.com/2026/08/19/waymos-cheaper-next-gen-robotaxi-is-now-open-to-all-riders-in-these-three-cities/" rel="noopener noreferrer"&gt;TechCrunch reported on 19 August 2026&lt;/a&gt; that the Ojai is cheaper to build, operate and maintain than the Jaguar I-Pace it replaces.&lt;/p&gt;

&lt;p&gt;That one sentence is the entire strategy. Waymo isn't scaling because the car finally got smart enough. It's scaling because the car got cheap enough. If you build anything with a per-unit cost attached, that distinction is your problem too.&lt;/p&gt;




&lt;h2&gt;
  
  
  🚕 What was actually announced
&lt;/h2&gt;

&lt;p&gt;Stripping out the press-release energy, here is what the source states:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Detail&lt;/th&gt;
&lt;th&gt;What the source says&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Vehicle&lt;/td&gt;
&lt;td&gt;Waymo Ojai, running Waymo's &lt;strong&gt;sixth-generation&lt;/strong&gt; driving system&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open to all riders in&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Los Angeles, Phoenix, San Francisco&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Next cities&lt;/td&gt;
&lt;td&gt;Denver, Las Vegas, San Diego, later in 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;In-car assistant&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Google Gemini&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Built by&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Zeekr&lt;/strong&gt; (owned by China's Geely), on Zeekr's SEA-M platform&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Autonomy hardware fitted&lt;/td&gt;
&lt;td&gt;At Waymo's &lt;strong&gt;Arizona&lt;/strong&gt; factory, after the vehicles are imported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fleet in commercial service&lt;/td&gt;
&lt;td&gt;About &lt;strong&gt;300&lt;/strong&gt; Ojais&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Imported in July 2026 alone&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;725&lt;/strong&gt; vehicles&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Projection&lt;/td&gt;
&lt;td&gt;On pace for &lt;strong&gt;5,000&lt;/strong&gt; in the US by end of 2026, per research firm MoffettNathanson&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Notice what is missing from that list: any claim that the new car drives better. The upgrade being sold is economic.&lt;/p&gt;




&lt;h2&gt;
  
  
  💰 The gap between 300 and 725 is the real headline
&lt;/h2&gt;

&lt;p&gt;Around 300 Ojais are carrying passengers. But 725 landed in the country in a single month. That is not a demand signal, it's a throughput signal.&lt;/p&gt;

&lt;p&gt;Every one of those vehicles has to pass through the Arizona factory to get its autonomy hardware fitted before it can earn a cent. So the constraint on Waymo's growth right now is not the model, not the roads, and not rider appetite in Los Angeles. It's a retrofit line.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; When a company that has spent a decade on the hard AI problem starts optimising its factory instead of its model, the technology is no longer the bottleneck. The supply chain is.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I find that genuinely useful as a diagnostic. Ask it about your own project:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the thing slowing you down the quality of your output, or the cost of producing each unit of it?&lt;/li&gt;
&lt;li&gt;If you doubled demand tomorrow, what would break first — accuracy, or your bill?&lt;/li&gt;
&lt;li&gt;Are you still tuning the demo when the demo already works?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most small teams I talk to are still polishing the model when their actual ceiling is that every customer costs them money.&lt;/p&gt;




&lt;h2&gt;
  
  
  🌐 Why the border decides your hardware cost
&lt;/h2&gt;

&lt;p&gt;The source notes that tariffs on imported vehicles have added cost to this programme. Waymo is a company with effectively unlimited engineering budget, buying from a manufacturing partner it has worked with since 2021, and its unit cost is still partly set by customs policy rather than by the factory.&lt;/p&gt;

&lt;p&gt;Anyone who has imported hardware into Sri Lanka knows that feeling exactly. The invoice from the supplier is the small number. The landed cost is the real one, and it is decided by duty structures you do not control and that change without warning.&lt;/p&gt;

&lt;p&gt;Concrete version of the same lesson:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A GPU that costs $600 abroad is not a $600 GPU once it clears Colombo.&lt;/li&gt;
&lt;li&gt;An EV's sticker price abroad tells you almost nothing about what it costs to put on the road here — our &lt;a href="https://induwara.lk/tools/sri-lanka-ev-import-tax-calculator" rel="noopener noreferrer"&gt;Sri Lanka EV import tax calculator&lt;/a&gt; exists precisely because that gap surprises people.&lt;/li&gt;
&lt;li&gt;Any hardware-flavoured startup plan built on foreign retail prices is a fiction until you run the duty math.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;If your business model depends on hardware crossing a border, model the border before you model the product. Waymo can absorb a tariff surprise. You cannot.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🛠️ The levers Waymo pulled, translated for a small team
&lt;/h2&gt;

&lt;p&gt;What Waymo is doing here is not exotic. It's four ordinary cost moves executed at scale, and each one has a version you can run this month.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Waymo's lever&lt;/th&gt;
&lt;th&gt;Your equivalent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Replacing the Jaguar I-Pace with a purpose-built Zeekr vehicle&lt;/td&gt;
&lt;td&gt;Dropping the expensive managed service you used to ship v1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fitting the autonomy hardware in-house in Arizona&lt;/td&gt;
&lt;td&gt;Owning the integration work instead of paying per-seat forever&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Using &lt;strong&gt;Gemini&lt;/strong&gt; as the in-car assistant rather than a bespoke voice stack&lt;/td&gt;
&lt;td&gt;Calling a general model API instead of training your own&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Importing 725 units in one month&lt;/td&gt;
&lt;td&gt;Committing to volume pricing only once demand is proven&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern in all four: &lt;strong&gt;buy the commodity, build the differentiator.&lt;/strong&gt; Waymo's differentiator is the driving system. The vehicle, the assistant and the platform underneath are all bought in. They spent their scarce money on exactly one thing.&lt;/p&gt;

&lt;p&gt;That's the discipline worth copying. If you are a two-person team in Colombo building something on AI APIs, your differentiator is almost certainly not the model. It's your data, your distribution, or your understanding of a local problem nobody in San Francisco will bother to solve. Everything else should be rented as cheaply as you can rent it.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;Three things I'd actually act on:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Separate your "does it work" budget from your "does it pay" budget.&lt;/strong&gt; These are different projects. A working demo says nothing about whether unit 1,000 is profitable. Waymo proved the driving years ago and is only now solving the second problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Find your retrofit line.&lt;/strong&gt; Every product has one step that quietly caps how fast you can grow. It is rarely the glamorous part. Find it, measure it, then decide whether to widen it or design around it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Price the whole chain, not the component.&lt;/strong&gt; Tariffs got Waymo. Egress fees, per-seat licences, payment gateway cuts and import duty will get you. Cost the delivered thing, not the parts list.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The robotaxi headline will be read as an AI story. It isn't. It's a company that finished the research phase and discovered the second half of the work is manufacturing, logistics and arithmetic. That second half is where most products actually live or die, and it's the half that a small team in Sri Lanka can be just as good at as anyone in California.&lt;/p&gt;

</description>
      <category>waymo</category>
      <category>autonomousvehicles</category>
      <category>uniteconomics</category>
    </item>
    <item>
      <title>Linear's AI usage data: why 65 PRs a week proves little</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Wed, 19 Aug 2026 14:20:37 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/linears-ai-usage-data-why-65-prs-a-week-proves-little-nga</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/linears-ai-usage-data-why-65-prs-a-week-proves-little-nga</guid>
      <description>&lt;p&gt;&lt;strong&gt;AI usage patterns in software teams&lt;/strong&gt; got a rare piece of hard data this week: Linear published &lt;a href="https://linear.app/data" rel="noopener noreferrer"&gt;Edition 01 of its data report&lt;/a&gt;, written by Tim Qi, drawn from aggregated activity inside its own product. The number everyone will quote is that teams which connected a coding agent went from &lt;strong&gt;21 pull requests a week to 65&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That figure is real. It is also close to meaningless on its own, and Linear says as much inside the report. The parts worth your attention are quieter.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 Read the definitions before you quote the numbers
&lt;/h2&gt;

&lt;p&gt;Every adoption chart in the report rests on a specific definition, and the definition is generous.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;What it actually measures&lt;/th&gt;
&lt;th&gt;Sample&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"AI-active" user&lt;/td&gt;
&lt;td&gt;At least &lt;strong&gt;one&lt;/strong&gt; AI interaction (in-app or Slack conversation, or an agent session) in a &lt;strong&gt;28-day window&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;127,000 paid users&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PR volume&lt;/td&gt;
&lt;td&gt;Pull requests opened per workspace per week&lt;/td&gt;
&lt;td&gt;47,900 paid workspaces&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding-agent cohort&lt;/td&gt;
&lt;td&gt;4,280 teams with an agent vs 2,607 without&lt;/td&gt;
&lt;td&gt;6,887 paid teams&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Company-size adoption&lt;/td&gt;
&lt;td&gt;Same 28-day AI-active bar, split by headcount&lt;/td&gt;
&lt;td&gt;199,000 paid users&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So when the report says product roles moved from &lt;strong&gt;12% to 34%&lt;/strong&gt; AI-active between January and June 2026, that is not "a third of product managers use AI daily." It is "a third touched it at least once in a month." Those are very different claims, and the gap between them is where most AI-adoption headlines go wrong.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; A 28-day "at least once" bar measures curiosity, not dependence. Adoption charts built on it show how many people tried the thing, not how many rely on it.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  📊 The real shift is in who files work, not how much gets filed
&lt;/h2&gt;

&lt;p&gt;Strip out the throughput noise and one pattern is genuinely new: non-engineers are now shipping code artifacts.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Function&lt;/th&gt;
&lt;th&gt;Attaching PRs, Jun 2024 → Jun 2026&lt;/th&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Engineering&lt;/td&gt;
&lt;td&gt;20% → 34%&lt;/td&gt;
&lt;td&gt;+14 pts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Product managers&lt;/td&gt;
&lt;td&gt;3% → 10%&lt;/td&gt;
&lt;td&gt;+7 pts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Designers&lt;/td&gt;
&lt;td&gt;1% → 8%&lt;/td&gt;
&lt;td&gt;+7 pts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A designer group going from 1% to 8% is an eightfold relative move. Alongside it, AI authorship of issues went from &lt;strong&gt;fewer than one in a thousand&lt;/strong&gt; to &lt;strong&gt;just under half&lt;/strong&gt; of all issues created.&lt;/p&gt;

&lt;p&gt;Adoption growth by function over January–June 2026 tells the same story:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Product:&lt;/strong&gt; 12% → 34%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Engineering:&lt;/strong&gt; 12% → 30%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Founder:&lt;/strong&gt; 14% → 30%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Design:&lt;/strong&gt; 6% → 22%&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Go-to-market:&lt;/strong&gt; 5% → 18%&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Product overtook engineering. For a five-person team in Colombo where the founder already writes copy, files bugs and reviews Figma, this is not a disruption; it is the tooling finally matching how you already work. The interesting question is not "will roles blur" but "who reviews the output when everyone can produce it."&lt;/p&gt;




&lt;h2&gt;
  
  
  ⚡ "Motion rather than value" is their phrase, not mine
&lt;/h2&gt;

&lt;p&gt;The report is unusually honest about its own limits. Two caveats are stated outright:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Looking at pull requests indicates motion rather than value. An opened PR says nothing about the value of the change.&lt;/p&gt;

&lt;p&gt;The data is a picture of adoption inside our own customer base, not the market at large.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is &lt;strong&gt;no data in the report on defect rates, review burden, revert frequency, or whether any of those PRs shipped&lt;/strong&gt;. The authors say they intend to cover the full lifecycle, from token spend to outcomes, in future work. Until then, the +111% two-year rise in overall PR volume describes artifacts created, not problems solved.&lt;/p&gt;

&lt;p&gt;Run the arithmetic on the headline. Sixty-five PRs per workspace per week, with three reviewers, is roughly &lt;strong&gt;22 reviews per person per week&lt;/strong&gt; on top of their own work. Nobody's reading capacity tripled between 2024 and 2026. If generation went up 3x and review capacity stayed flat, the constraint simply moved, and the report has no visibility into what happened at the new constraint.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ What a small Sri Lankan team should actually do with this
&lt;/h2&gt;

&lt;p&gt;The report is drawn from paid Linear workspaces, which skews toward funded startups. Three things translate anyway:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Treat review capacity as the scarce resource.&lt;/strong&gt; Cap open PRs per author before you raise agent throughput. A queue of 40 unreviewed agent PRs is worse than 10 human ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enforce small diffs on agent output.&lt;/strong&gt; Agents happily produce 900-line changes. A 900-line diff does not get reviewed; it gets approved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Watch the currency mismatch.&lt;/strong&gt; Agent tokens are billed in USD while salaries are paid in LKR, so per-seat AI tooling costs a Sri Lankan team far more in relative terms than it costs the median workspace in this dataset. Estimate spend before you commit: our &lt;a href="https://induwara.lk/tools/ai-token-counter" rel="noopener noreferrer"&gt;AI token counter&lt;/a&gt; gives you a per-prompt figure, and the &lt;a href="https://induwara.lk/tools/lkr-exchange-rate" rel="noopener noreferrer"&gt;LKR exchange rate tool&lt;/a&gt; turns a USD subscription into a number your budget recognises.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For students and solo builders, the useful signal is different. Non-engineers filing PRs at 8-10% means the floor for "can contribute code" dropped. The differentiator is no longer typing the code; it is judging whether the change is correct.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 The metrics they didn't publish are the ones to track
&lt;/h2&gt;

&lt;p&gt;The interesting measurements are the ones absent from Edition 01, and every one of them is free to collect from your own git history.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Why it matters&lt;/th&gt;
&lt;th&gt;How to get it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Merged ÷ opened PRs&lt;/td&gt;
&lt;td&gt;Separates shipped work from motion&lt;/td&gt;
&lt;td&gt;GitHub API, or &lt;code&gt;gh pr list --state all&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time to first review&lt;/td&gt;
&lt;td&gt;Shows where the queue is backing up&lt;/td&gt;
&lt;td&gt;PR timestamps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Revert + hotfix rate&lt;/td&gt;
&lt;td&gt;The closest cheap proxy for quality&lt;/td&gt;
&lt;td&gt;Commit message grep&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;% of PRs over 400 lines&lt;/td&gt;
&lt;td&gt;Predicts rubber-stamped reviews&lt;/td&gt;
&lt;td&gt;&lt;code&gt;git log --shortstat&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A starting point that costs nothing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# hotfix/revert share of the last 200 commits&lt;/span&gt;
git log &lt;span class="nt"&gt;-200&lt;/span&gt; &lt;span class="nt"&gt;--oneline&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-icE&lt;/span&gt; &lt;span class="s1"&gt;'^\w+ (revert|hotfix|fix: regression)'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Track those four for a month before and after you add an agent. That is a better dataset for your team than any industry report, because it is measured on your codebase, with your reviewers.&lt;/p&gt;




&lt;h2&gt;
  
  
  🚀 What this means for you
&lt;/h2&gt;

&lt;p&gt;Linear's report is worth reading, and its honesty about its own limits is worth more than its charts. Take three things from it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Adoption numbers rest on soft definitions.&lt;/strong&gt; Check the window and the threshold before you repeat a percentage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The role boundary moved.&lt;/strong&gt; Designers and PMs attaching PRs at 8-10% is the durable finding, not the throughput jump.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nobody has published the quality half yet.&lt;/strong&gt; Until someone does, throughput claims about coding agents are unfinished sentences.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you run a small team here, the cheapest advantage available right now is measuring the thing the vendors have not measured. Review latency and revert rate take an afternoon to instrument, and they tell you whether your agents are helping or just producing.&lt;/p&gt;

</description>
      <category>aiengineering</category>
      <category>developerproductivity</category>
      <category>teampractices</category>
    </item>
    <item>
      <title>The agentic SDLC: steal the eval, skip the five agents</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Tue, 18 Aug 2026 18:01:58 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/the-agentic-sdlc-steal-the-eval-skip-the-five-agents-1n2e</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/the-agentic-sdlc-steal-the-eval-skip-the-five-agents-1n2e</guid>
      <description>&lt;p&gt;The &lt;strong&gt;agentic SDLC&lt;/strong&gt; is having its moment, and most of the coverage focuses on the wrong part of it. On the &lt;a href="https://stackoverflow.blog/2026/08/18/what-does-an-agentic-sdlc-actually-look-like/" rel="noopener noreferrer"&gt;Stack Overflow Podcast&lt;/a&gt;, Ryan talks to &lt;strong&gt;Suneet Malhotra&lt;/strong&gt;, Senior Manager of Test Engineering at Motorola Solutions, about a &lt;strong&gt;five-agent&lt;/strong&gt; end-to-end pipeline built on &lt;strong&gt;MCP&lt;/strong&gt; servers.&lt;/p&gt;

&lt;p&gt;My reaction was not "I want five agents." It was: the two smallest ideas in that episode are the only ones a two-person team in Colombo can actually run, and they're the ones nobody is copying.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 The agent count is the least interesting number
&lt;/h2&gt;

&lt;p&gt;Five agents wired across the lifecycle is an org-chart diagram more than an engineering insight. It tells you Motorola Solutions has enough scale to justify a dedicated stage per phase. It does not tell you that five is correct, or that four fails.&lt;/p&gt;

&lt;p&gt;What the shape actually admits is more useful:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agents are placed &lt;strong&gt;at handoffs&lt;/strong&gt;, not at tasks. Design → spec, spec → code, code → test.&lt;/li&gt;
&lt;li&gt;Handoffs are where requirements quietly get lost, so that's where the cost sits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP&lt;/strong&gt; is the plumbing choice, not the idea. It's what makes each stage able to read the ticket, the repo, and the test results without a bespoke integration each time.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;If you take one structural lesson: agents earn their keep at the seams between phases, not inside them. Automating "write the code" is crowded. Automating "did the spec survive the handoff" is not.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🛠️ Specification enrichment is the cheapest stage in the whole pipeline
&lt;/h2&gt;

&lt;p&gt;The technique I'd implement tomorrow is &lt;strong&gt;specification enrichment&lt;/strong&gt;: an extra stage that runs immediately after design, before anyone writes implementation code. QA moves left, into the requirements themselves.&lt;/p&gt;

&lt;p&gt;Concretely, it's one prompt against your design doc that asks the questions a good tester would ask in review:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What are the boundary values, and what happens at each one?&lt;/li&gt;
&lt;li&gt;Which inputs are unspecified, and what does the system do when it gets them?&lt;/li&gt;
&lt;li&gt;What's the failure mode when a dependency is down?&lt;/li&gt;
&lt;li&gt;Which acceptance criteria are untestable as written?&lt;/li&gt;
&lt;li&gt;What did this spec assume without saying?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For a Sri Lankan tool like an EPF or tax calculator, that stage is the difference between shipping and shipping wrong. "Calculate the tax" passes review. "What happens at exactly the bracket boundary, at zero income, at a negative deduction, when the effective date falls mid-year" does not, until you answer it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; the bug you catch in the spec costs one prompt. The same bug caught after a user files it costs a rebuild, a redeploy, and your credibility on a page Google already ranked.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You don't need an agent framework for this. You need a checklist prompt and the discipline to run it before you open the editor.&lt;/p&gt;




&lt;h2&gt;
  
  
  📊 Cohen's kappa is the part I'd steal first
&lt;/h2&gt;

&lt;p&gt;The genuinely technical idea in the episode is using &lt;strong&gt;Cohen's kappa&lt;/strong&gt; to compare multiple LLMs acting as judges. This matters because "LLM-as-a-judge" has a silent failure mode: your judge agrees with you often enough to feel right, and you never check how much of that agreement is luck.&lt;/p&gt;

&lt;p&gt;Kappa corrects for chance agreement:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;κ = (p_o − p_e) / (1 − p_e)

p_o = observed agreement
p_e = agreement expected by chance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Worked example, my numbers, not the episode's. You hand-label 100 outputs pass/fail and 80 are passes. Your judge also calls 80 passes, and agrees with you on 84 of the 100 items.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Quantity&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Observed agreement (p_o)&lt;/td&gt;
&lt;td&gt;0.84&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Expected by chance (p_e)&lt;/td&gt;
&lt;td&gt;0.80×0.80 + 0.20×0.20 = &lt;strong&gt;0.68&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cohen's κ&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;(0.84 − 0.68) / (1 − 0.68) = &lt;strong&gt;0.50&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;84% accuracy reads like a working judge. A kappa of 0.50 says half the headroom above chance is unexplained. On an unbalanced dataset, raw accuracy flatters you badly.&lt;/p&gt;

&lt;p&gt;The conventional reading of kappa (the Landis and Koch bands, widely used in inter-rater work):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;κ range&lt;/th&gt;
&lt;th&gt;Interpretation&lt;/th&gt;
&lt;th&gt;Would I ship on it?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&amp;lt; 0.20&lt;/td&gt;
&lt;td&gt;Slight&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.21 – 0.40&lt;/td&gt;
&lt;td&gt;Fair&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.41 – 0.60&lt;/td&gt;
&lt;td&gt;Moderate&lt;/td&gt;
&lt;td&gt;Only with human spot-checks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.61 – 0.80&lt;/td&gt;
&lt;td&gt;Substantial&lt;/td&gt;
&lt;td&gt;Yes, with sampling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.81 – 1.00&lt;/td&gt;
&lt;td&gt;Almost perfect&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Run this across two or three candidate judge models and you get something an accuracy score can't give you: evidence about which model to trust, and whether your rubric is the actual problem. If every model scores fair, the rubric is ambiguous, not the models.&lt;/p&gt;




&lt;h2&gt;
  
  
  💰 Running this on a learning budget
&lt;/h2&gt;

&lt;p&gt;The gap between a Motorola Solutions pipeline and a solo build in Sri Lanka is not intelligence. It's tokens per change. Five agents on every commit is a monthly bill; two agents at the two highest-leverage points is a rounding error.&lt;/p&gt;

&lt;p&gt;Where I'd spend, in order:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Worth it solo?&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Spec enrichment&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Yes, first&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One call, catches requirement bugs before code exists&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Judge / eval agent&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Yes, second&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tells you whether anything else you automate is working&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code generation&lt;/td&gt;
&lt;td&gt;Already have it&lt;/td&gt;
&lt;td&gt;Your editor does this&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Test generation&lt;/td&gt;
&lt;td&gt;Maybe&lt;/td&gt;
&lt;td&gt;Useful once specs are enriched, weak before&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Release / deploy agent&lt;/td&gt;
&lt;td&gt;Not yet&lt;/td&gt;
&lt;td&gt;Needs infra maturity you probably don't have&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two practical notes. First, a hand-labelled set of &lt;strong&gt;50 to 100 examples&lt;/strong&gt; is enough to compute a meaningful kappa, and you can build one in an afternoon. Second, before you commit to a judge model, price the workload: our &lt;a href="https://induwara.lk/tools/ai-token-counter" rel="noopener noreferrer"&gt;AI token counter&lt;/a&gt; shows how much context each eval call actually burns, and the &lt;a href="https://induwara.lk/tools/ai-model-comparison" rel="noopener noreferrer"&gt;model comparison tool&lt;/a&gt; lists input and output pricing side by side. Judges run on every candidate output, so per-token cost compounds faster than you'd expect.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;Read the source episode as a report from a large test-engineering org, not a blueprint. The transferable parts are small and cheap:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Put an agent at the handoff, not the task.&lt;/strong&gt; Design-to-spec is the highest-yield seam.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enrich the spec before writing code.&lt;/strong&gt; Five questions, one prompt, before the editor opens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never trust a judge you haven't measured.&lt;/strong&gt; Compute kappa on a hand-labelled set. Accuracy alone will lie to you on unbalanced data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Buy the two stages that pay, skip the three that don't.&lt;/strong&gt; Agent count is not a maturity score.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;Bottom line: an agentic SDLC isn't a fleet of agents. It's the two places where you stopped guessing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;I'm commenting on the Stack Overflow Podcast episode linked above, not reproducing it. The five-agent pipeline, MCP plumbing, Cohen's kappa evaluation and specification enrichment are theirs; the worked numbers, budget ordering and opinions here are mine.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiengineering</category>
      <category>softwaretesting</category>
      <category>llmevaluation</category>
    </item>
    <item>
      <title>A photorealistic Magic card browser and the cost of pretty UI</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Mon, 17 Aug 2026 21:39:37 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/a-photorealistic-magic-card-browser-and-the-cost-of-pretty-ui-3cjc</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/a-photorealistic-magic-card-browser-and-the-cost-of-pretty-ui-3cjc</guid>
      <description>&lt;p&gt;&lt;strong&gt;Why beautiful websites feel slow&lt;/strong&gt; is a question I keep running into, and this week a Magic: The Gathering card browser gave me a very clean example of it. &lt;a href="https://magic-oracle.com/" rel="noopener noreferrer"&gt;Oracle&lt;/a&gt; went up on Hacker News with the tagline &lt;em&gt;"Search every Magic card ever printed"&lt;/em&gt;, built, per the site's own footer, by &lt;strong&gt;Egstad from DBCo&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It got &lt;strong&gt;3 points and 2 comments&lt;/strong&gt;. The first comment was not about the design. It was a complaint that the logo animation was eating CPU. That gap between what was built and what was noticed first is the part worth writing about.&lt;/p&gt;




&lt;h2&gt;
  
  
  🎴 What Oracle actually ships
&lt;/h2&gt;

&lt;p&gt;Ignore the visuals for a second and the feature list is a serious search tool. From the site itself:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;What it covers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Colour filters&lt;/td&gt;
&lt;td&gt;Mono, two, three and five-colour combinations, named guilds like &lt;strong&gt;Azorius&lt;/strong&gt; and &lt;strong&gt;Rakdos&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Type filters&lt;/td&gt;
&lt;td&gt;Planeswalkers, Legends, Battles, Sagas, Equipment, Vehicles, Lands, Artifacts, Enchantments, Tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Format filters&lt;/td&gt;
&lt;td&gt;Standard, Pioneer, Modern, Legacy, Vintage, Pauper, Commander with &lt;strong&gt;EDHREC&lt;/strong&gt; ranking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Effect search&lt;/td&gt;
&lt;td&gt;Counterspells, board wipes, removal, card draw, ramp, mill, lifegain, landfall, reanimation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sorting&lt;/td&gt;
&lt;td&gt;Name, release date, set number, rarity, colour, price in &lt;strong&gt;USD / EUR / TIX&lt;/strong&gt;, mana value, power/toughness, artist, EDHREC rank&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Display modes&lt;/td&gt;
&lt;td&gt;Small grid, large grid, list, text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Result granularity&lt;/td&gt;
&lt;td&gt;One per card, every artwork, every printing, or paper only&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is a detail most people would skip. "Every printing" versus "one per card" is a real data-modelling decision, and exposing it as a user-facing toggle means someone thought hard about who is actually searching.&lt;/p&gt;

&lt;p&gt;The site does not state where its card data comes from, so I won't guess. What it does state is the intent, and the creator said it plainly in the thread.&lt;/p&gt;




&lt;h2&gt;
  
  
  🎭 The real bet: the data is commodity, the feel is the product
&lt;/h2&gt;

&lt;p&gt;The top comment made the obvious objection: &lt;strong&gt;Gatherer&lt;/strong&gt; already "has all the same info in a much snappier package." The creator's reply was that they knew about both Gatherer and Scryfall and wanted something that felt "a bit more cinematic and photorealistic."&lt;/p&gt;

&lt;p&gt;That is not a dodge. It is the entire product thesis, stated honestly.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; when the underlying data is public and everyone can have it, the only remaining place to compete is the experience layer. That's a legitimate strategy, but it means presentation is no longer decoration. It &lt;em&gt;is&lt;/em&gt; the feature, and it gets judged as harshly as a wrong tax bracket would be.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This matters for anyone here building tools. Almost every dataset a small team can build on is already available to competitors. Bus timetables, exchange rates, exam results, public gazette data. You will not win on having the numbers. You might win on the ten seconds it takes someone to find the one number they came for.&lt;/p&gt;

&lt;p&gt;But if the experience is your product, then a stutter is a bug in your core feature, not a cosmetic nit. Which is exactly what happened here.&lt;/p&gt;




&lt;h2&gt;
  
  
  ⚡ What "cinematic" costs on a mid-range phone
&lt;/h2&gt;

&lt;p&gt;The complaint was specific: a particle animation on the logo pulling roughly &lt;strong&gt;5% CPU&lt;/strong&gt;, on a page whose job is to sit there while you read card text. The creator asked, reasonably, what hardware that was measured on.&lt;/p&gt;

&lt;p&gt;Both people are right, and that's the interesting part. On a recent laptop, 5% is noise. Here is what the same background cost looks like as you move down the device ladder:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Device class&lt;/th&gt;
&lt;th&gt;5% sustained CPU means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Recent laptop, plugged in&lt;/td&gt;
&lt;td&gt;Genuinely nothing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Laptop on battery&lt;/td&gt;
&lt;td&gt;Fan spins up, measurable battery drain on a long session&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mid-range Android, 4GB RAM&lt;/td&gt;
&lt;td&gt;Competes with your own scroll and image decode; jank shows up first here&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Older phone, thermally throttled&lt;/td&gt;
&lt;td&gt;The animation and the content fight, and the content loses&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The device most Sri Lankan students and freelancers are browsing on is not the device most of us develop on. An effect that costs nothing on a MacBook is a frame budget you've already spent before the user has scrolled once.&lt;/p&gt;

&lt;p&gt;There's a second bill too. A card browser is, by definition, a very large number of high-quality images. Photorealistic rendering of card faces and foils means bigger assets, and mobile data in Sri Lanka is bought in finite bundles. Anything image-heavy should be compressing aggressively before it ships. We keep a free &lt;a href="https://induwara.lk/tools/image-compressor" rel="noopener noreferrer"&gt;image compressor&lt;/a&gt; and an &lt;a href="https://induwara.lk/tools/html-css-js-minifier" rel="noopener noreferrer"&gt;HTML/CSS/JS minifier&lt;/a&gt; online for exactly this reason, and both run in your browser rather than uploading anything.&lt;/p&gt;




&lt;h2&gt;
  
  
  🧭 The Escape-key problem
&lt;/h2&gt;

&lt;p&gt;The second complaint was navigation: the commenter couldn't get out of a full-screen card view without pressing Escape. The creator's answer was that cards open in a lightbox, so Escape returns you to the grid, and called it a common pattern.&lt;/p&gt;

&lt;p&gt;He's right that it's a common pattern. He's also, I'd argue, losing the argument, because the user reported being stuck anyway.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A convention only works when the user recognises it. If someone has to be told the pattern after the fact, the pattern didn't do its job on their screen.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Two things break the lightbox convention quietly:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;On touch, there is no Escape key.&lt;/strong&gt; The entire keyboard affordance is unavailable to a phone user, and the visual close target has to carry the full load.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A photorealistic full-bleed view can swallow the exit.&lt;/strong&gt; The more immersive the view, the more the close control looks like part of the artwork instead of a control.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The fix is cheap and unglamorous. A visible close button with a real hit area, a click on the backdrop, and browser-back mapped to close. Three of them, not one, because you don't know which one any given person will reach for.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;Oracle is a well-made thing built on a defensible bet, and I'd rather see this than the tenth identical minimal-grey search page. But the feedback loop it hit in public is the one to learn from. Here's what I take from it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pick your competitive layer deliberately.&lt;/strong&gt; If the data is commodity, say out loud that experience is the product, then hold yourself to that standard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set a frame budget before you add the effect.&lt;/strong&gt; Idle animation should be near-zero cost, because "idle" is where users spend most of their time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test on the cheapest device your audience owns&lt;/strong&gt;, not the one on your desk. Chrome DevTools CPU throttling at 4x costs nothing to turn on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never let a keyboard shortcut be the only exit.&lt;/strong&gt; Half your traffic has no keyboard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compress before you ship.&lt;/strong&gt; Especially with image-heavy pages on metered mobile data.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Delight is worth paying for. Just know the price, and know whose device is paying it.&lt;/p&gt;

</description>
      <category>webperf</category>
      <category>productdesign</category>
      <category>frontend</category>
    </item>
    <item>
      <title>LLM gender bias hides in how you write, not who you are</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Mon, 17 Aug 2026 01:38:15 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/llm-gender-bias-hides-in-how-you-write-not-who-you-are-3e7f</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/llm-gender-bias-hides-in-how-you-write-not-who-you-are-3e7f</guid>
      <description>&lt;p&gt;&lt;strong&gt;LLM gender bias&lt;/strong&gt; turns out not to live where most of us have been looking for it. A paper posted to arXiv on &lt;strong&gt;13 August 2026&lt;/strong&gt;, &lt;a href="https://arxiv.org/abs/2608.13328" rel="noopener noreferrer"&gt;&lt;em&gt;It's How You Ask: Gender-Associated Linguistic Bias in LLMs&lt;/em&gt;&lt;/a&gt; by &lt;strong&gt;Katherine Van Koevering&lt;/strong&gt; and &lt;strong&gt;Anjalie Field&lt;/strong&gt;, reports something more awkward than a name-swap test: putting a woman's name in the prompt changed nothing, but writing in a hedged, polite register produced large and consistent differences in what the model gave back.&lt;/p&gt;

&lt;p&gt;I want to argue that this finding is much bigger than gender, and that if you ship anything with an LLM behind it, it is now your bug to fix.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 What the paper actually measured
&lt;/h2&gt;

&lt;p&gt;The authors tested prompts across &lt;strong&gt;four models&lt;/strong&gt; and &lt;strong&gt;three document types&lt;/strong&gt;, varying the linguistic register rather than the stated identity of the user. The features they looked at are the ones sociolinguists have long associated with women's speech in English:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hedges&lt;/strong&gt; — "I think maybe", "sort of", "just wondering if"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tag questions&lt;/strong&gt; — "…that would work, wouldn't it?"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Collective reference&lt;/strong&gt; — "we should", "our team needs", instead of "write me"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Prompts carrying those features got responses that were &lt;strong&gt;shorter, less sophisticated, and less formal&lt;/strong&gt;, and the effect held after controlling for prompt complexity. So this is not the model correctly reading a vaguer request and giving a vaguer answer.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal in the prompt&lt;/th&gt;
&lt;th&gt;Effect on the response&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Explicit gender cue (a name)&lt;/td&gt;
&lt;td&gt;No effect reported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hedges, tag questions, collective reference&lt;/td&gt;
&lt;td&gt;Large, consistent effect&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt complexity (controlled for)&lt;/td&gt;
&lt;td&gt;Not the explanation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; The models are not discriminating on identity. They are discriminating on &lt;em&gt;style&lt;/em&gt; — and style correlates with identity, which produces the same outcome while passing every name-swap fairness test you were running.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🌐 Why Sri Lankan English sits directly in the blast radius
&lt;/h2&gt;

&lt;p&gt;Here is the part that made me stop and rewrite my own prompt templates.&lt;/p&gt;

&lt;p&gt;The register the paper flags as penalised is, more or less, the default polite register of Sri Lankan professional English. Think about how a typical email from a Colombo office actually opens:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I was just wondering if you could kindly help me with a small thing.
We were hoping to maybe put together a short proposal, if that's okay?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is three hedges, one collective reference, and a softener, in two sentences. Nobody wrote it that way because they were unsure. They wrote it that way because in Sri Lankan English, going straight to "Write a proposal." reads as rude.&lt;/p&gt;

&lt;p&gt;The paper's own framing is that these features are &lt;strong&gt;culturally embedded and outside conscious control&lt;/strong&gt;. That is the sharp end of it. If the workaround is "just write more assertively", then the workaround is asking people to code-switch out of their own dialect to get the same service everyone else gets by default. For a student in Kandy writing a scholarship statement, or a freelancer pitching an overseas client, that is a real and invisible tax.&lt;/p&gt;

&lt;p&gt;I have not seen data on non-native or South Asian English registers specifically, and the paper does not claim to have tested that. So treat this as my inference, not their finding. But the mechanism generalises cleanly: if politeness markers depress output quality, every high-politeness English variety is exposed.&lt;/p&gt;




&lt;h2&gt;
  
  
  ⚡ Why "just add a system prompt" won't fix it
&lt;/h2&gt;

&lt;p&gt;The instinct for most of us building on top of an API is to patch this at the instruction layer: append &lt;em&gt;"treat all requests with equal rigour regardless of phrasing."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The paper closes that door. It reports that these linguistic features are &lt;strong&gt;encoded in early transformer layers and entangled with other features&lt;/strong&gt;. In plain terms:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The register signal is picked up &lt;strong&gt;early&lt;/strong&gt;, well before whatever your system prompt is trying to steer.&lt;/li&gt;
&lt;li&gt;It is &lt;strong&gt;entangled&lt;/strong&gt; — you cannot cleanly isolate and suppress "hedginess" without dragging along whatever else those representations carry.&lt;/li&gt;
&lt;li&gt;So mitigation has to happen either in training, or &lt;strong&gt;outside the model entirely&lt;/strong&gt;, in your application layer.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Option three is the only one available to a small team in Sri Lanka. Which is fine, because option three is also the cheapest.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ What to actually build if you ship an LLM feature
&lt;/h2&gt;

&lt;p&gt;Concretely, a &lt;strong&gt;prompt normalisation step&lt;/strong&gt; between your user's text and the model:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Normaliser&lt;/td&gt;
&lt;td&gt;Rewrites the user's raw input into a flat, direct instruction before it hits the model&lt;/td&gt;
&lt;td&gt;One cheap small-model call, or pure regex for common hedges&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eval pair&lt;/td&gt;
&lt;td&gt;Runs the same request in hedged and direct form, compares outputs&lt;/td&gt;
&lt;td&gt;Runs in CI, not per-request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Length floor&lt;/td&gt;
&lt;td&gt;Flags responses well below the median for that request type&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A regex-only normaliser gets you a surprising distance. Stripping a fixed list of openers ("I was just wondering if", "would it be possible to", "sorry to bother you but") costs nothing per request and is fully auditable, which matters more than elegance when you are debugging why one user's output looks thin.&lt;/p&gt;

&lt;p&gt;If you want to sanity-check your own phrasing before wiring anything up, our &lt;a href="https://induwara.lk/tools/ai-prompt-formatter" rel="noopener noreferrer"&gt;AI Prompt Formatter&lt;/a&gt; turns rough or over-softened notes into a direct instruction, and the &lt;a href="https://induwara.lk/tools/word-counter" rel="noopener noreferrer"&gt;Word Counter&lt;/a&gt; is enough to measure the response-length gap between two phrasings of the same request. That is the whole test, and it costs you nothing.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Run this once today:&lt;/strong&gt; take one real prompt from your app. Write it twice — once hedged, once blunt. Send both. Count the words in each reply. If the gap is large, you have shipped the bug.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If you use LLMs for work:&lt;/strong&gt; you are probably leaving quality on the table through politeness alone. Be blunt with the model and polite with the human. They are different audiences.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you build on an API:&lt;/strong&gt; name-swap fairness tests are no longer sufficient evidence that your product treats users equally. Add register to your eval set.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you teach or mentor:&lt;/strong&gt; telling students to "prompt better" is fair advice, but be honest that you are teaching a dialect, not a skill.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The uncomfortable version of this result is that the fairness checks the industry standardised on were testing the variable that turned out not to matter. Names are easy to test, so we tested names. Register is hard to test, so we mostly didn't.&lt;/p&gt;

&lt;p&gt;The useful version is that the fix is within reach of a two-person team with a free-tier API key and an afternoon. Normalise the input, measure the output, ship it. You do not need to wait for a model provider to solve this for you.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>bias</category>
    </item>
  </channel>
</rss>
