<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Induwara Ashinsana</title>
    <description>The latest articles on DEV Community by Induwara Ashinsana (@induwara_ashinsana_9e4d5b).</description>
    <link>https://dev.to/induwara_ashinsana_9e4d5b</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3958655%2F6fa1c062-e3e8-4949-affc-f60cccfc2dfb.jpg</url>
      <title>DEV Community: Induwara Ashinsana</title>
      <link>https://dev.to/induwara_ashinsana_9e4d5b</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/induwara_ashinsana_9e4d5b"/>
    <language>en</language>
    <item>
      <title>Best Free AI Video Generators 2026 — Honest Comparison (Watermarks, Limits, Real Quality)</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Thu, 10 Sep 2026 02:45:23 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/best-free-ai-video-generators-2026-honest-comparison-watermarks-limits-real-quality-16g4</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/best-free-ai-video-generators-2026-honest-comparison-watermarks-limits-real-quality-16g4</guid>
      <description>&lt;p&gt;Every "best &lt;a href="https://induwara.lk/tools/ai-voice-generator" rel="noopener noreferrer"&gt;free AI&lt;/a&gt; video generator" list on the internet right now misleads you the same way. They list ten tools, call them "free," and never mention that the free tier gives you four watermarked seconds of 480p, signs you up for a 200-person queue, or quietly burns through your monthly credits in three clicks.&lt;/p&gt;

&lt;p&gt;This is the honest version. I went through every tool below in the last 72 hours. Here is what "free" actually buys you in mid-2026, where the trade-offs hide, and which tool fits which job.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR — what "free" actually means
&lt;/h2&gt;

&lt;p&gt;Across every consumer AI video tool, "free" comes in four flavours, and most lists conflate them:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Free tier with credits&lt;/strong&gt; — Runway, Luma, Pika. You get N credits at signup, a smaller monthly refill, watermarks, and clear length caps. Useful for testing, painful for anything regular.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free with a daily or weekly cap&lt;/strong&gt; — Hailuo, Pixverse, Kling. No credit count, but you queue alongside everyone else, and quality drops at peak times.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open source, self-hosted&lt;/strong&gt; — Stable Video Diffusion, AnimateDiff. Free if you have a GPU or a Hugging Face Space allowance. Not free if you're on a phone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Free" via someone else's API key&lt;/strong&gt; — the entire premise of the viral Facebook "make a free Sora 2 playground" post. This works for about thirty minutes before the host platform removes the listing for ToS violation. Skip it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Nothing is free as in "[[[&lt;a href="https://induwara.lk/tools/freelancer-hourly-rate-calculator" rel="noopener noreferrer"&gt;no signup&lt;/a&gt;](&lt;a href="https://induwara.lk/tools/invoice-generator)%5D(https://induwara.lk/tools/speech-to-text)%5D(https://induwara.lk/tools/text-to-speech" rel="noopener noreferrer"&gt;https://induwara.lk/tools/invoice-generator)](https://induwara.lk/tools/speech-to-text)](https://induwara.lk/tools/text-to-speech&lt;/a&gt;), no watermark, no limit, unlimited generations." If a site claims otherwise in 2026, it is either burning VC money to pull you in for an upsell or proxying a stolen API key. Either way it disappears within a week.&lt;/p&gt;

&lt;h2&gt;
  
  
  The comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Free credits / limit&lt;/th&gt;
&lt;th&gt;Max clip&lt;/th&gt;
&lt;th&gt;Watermark&lt;/th&gt;
&lt;th&gt;Quality&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Luma Dream Machine&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~30 generations / month&lt;/td&gt;
&lt;td&gt;5 sec&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;YouTube Shorts, social&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Runway&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;525 credits at signup, then 125 / month&lt;/td&gt;
&lt;td&gt;4–10 sec&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Highest (Gen-3)&lt;/td&gt;
&lt;td&gt;Pro-quality short clips&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pika 1.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~80 generations / month&lt;/td&gt;
&lt;td&gt;3 sec&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Mid&lt;/td&gt;
&lt;td&gt;Quick stylistic clips&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hailuo (MiniMax)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Daily queue, ~10–20 / day&lt;/td&gt;
&lt;td&gt;6 sec&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Mid-high&lt;/td&gt;
&lt;td&gt;Anime, stylised motion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Kling 1.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~6 generations / day&lt;/td&gt;
&lt;td&gt;5–10 sec&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Photoreal motion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pixverse&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~5 / day&lt;/td&gt;
&lt;td&gt;4 sec&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Mid&lt;/td&gt;
&lt;td&gt;Fast iteration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Stable Video Diffusion&lt;/strong&gt; (open source)&lt;/td&gt;
&lt;td&gt;Unlimited if you bring a GPU&lt;/td&gt;
&lt;td&gt;4 sec&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Mid&lt;/td&gt;
&lt;td&gt;Local, full control&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Sora (OpenAI)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Requires ChatGPT Plus ($20 / month)&lt;/td&gt;
&lt;td&gt;5–20 sec&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Highest&lt;/td&gt;
&lt;td&gt;Not actually free&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Where I would actually start
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;For YouTube Shorts and Instagram Reels:&lt;/strong&gt; Luma Dream Machine. The free monthly credit gives you about thirty attempts, quality sits at the top of the free pack, and a 5-second clip is plenty for short-form. Plan your prompt; don't iterate in panic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For one-off "wow, look what AI can do" clips:&lt;/strong&gt; Runway. Save the 525 signup credits for the things you actually want to keep — Gen-3 is noticeably better than anything else in the free pool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For repeatable workflows (product clips, social ad tests):&lt;/strong&gt; Hailuo or Pixverse. Daily limits are forgiving, the queues move, and you can run a steady drip rather than burn credits in a sprint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For control and zero ongoing cost:&lt;/strong&gt; Stable Video Diffusion via a Hugging Face Space (or local if you have a 12 GB+ GPU). Less convenient than the hosted tools, but you own the workflow and there is no rate limit.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is NOT actually free — calling out the hype
&lt;/h2&gt;

&lt;p&gt;Three claims circulating right now are misleading, and they're costing people time:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;"Free Sora 2 API"&lt;/strong&gt; — there is no public Sora 2 REST API as of mid-2026. Sora 1 access goes through ChatGPT Plus or Pro. Sites and Facebook posts offering a "free Sora 2 key" are reselling cracked accounts or selling you a fantasy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Free Sora playground on Hugging Face Spaces"&lt;/strong&gt; — the viral self-host trick. The moment you password-gate a Space for commercial use, you violate Hugging Face's free-tier terms and the Space is removed. Lifespan in practice: a few hours, not "free for life."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Free unlimited AI video"&lt;/strong&gt; — every consumer tool runs on someone's GPU. Someone is paying for that GPU. If you are not paying with money, you are paying with watermarks, length limits, ads, queues, &lt;a href="https://induwara.lk/tools/image-upscaler" rel="noopener noreferrer"&gt;or your&lt;/a&gt; prompts becoming training data. Pick which one is acceptable to you and move on.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The free-tier landscape is workable if you set expectations correctly. The "completely free, unlimited" landscape does not exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  A workflow that combines the free tools
&lt;/h2&gt;

&lt;p&gt;The way to get the most out of free AI video is to pair it with the rest of the free creator stack:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Script your idea.&lt;/strong&gt; Keep it tight — every clip is 3–10 seconds. Use a free &lt;a href="https://induwara.lk/tools/ai-text-summarizer" rel="noopener noreferrer"&gt;AI text summarizer&lt;/a&gt; to compress a longer concept into a punchy prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generate the video.&lt;/strong&gt; Pick a tool from above based on what you are making.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add narration.&lt;/strong&gt; Generate a voiceover with our &lt;a href="https://induwara.lk/tools/text-to-speech" rel="noopener noreferrer"&gt;free text-to-speech tool&lt;/a&gt; and match the voice tone to your visual style.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sequence and trim.&lt;/strong&gt; CapCut Web (free) or DaVinci Resolve (free desktop) handle stitching, captions, and a music bed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Share privately first.&lt;/strong&gt; Before publishing to a million strangers, send a draft to collaborators using our &lt;a href="https://induwara.lk/tools/secret-file" rel="noopener noreferrer"&gt;self-destructing file share&lt;/a&gt; — the only copy floating around is a one-time encrypted link.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most "AI video for free" guides stop at step 2. Steps 3–5 are where a free workflow quietly becomes a professional one.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is Sora 2 free?&lt;/strong&gt;&lt;br&gt;
No. Sora 1 is bundled with ChatGPT Plus and Pro ($20 / $200 per month). Sora 2 has no public consumer pricing yet. Anyone offering a "free Sora 2 API key" is misleading you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the best free AI video generator for YouTube Shorts in 2026?&lt;/strong&gt;&lt;br&gt;
Luma Dream Machine. Roughly thirty free generations per month, 5-second clips at the highest free-tier quality, and the watermark is small enough to crop or cover.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I remove the watermark from free AI video?&lt;/strong&gt;&lt;br&gt;
You cannot, legally. The watermark is the price of the free tier. Paid plans on Runway, Pika, and Luma all remove it; subscription cost runs $10–$95 per month depending on tier.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I use AI-generated video commercially?&lt;/strong&gt;&lt;br&gt;
It depends on the tool's terms, the tier you are on, and the model used. Runway and Luma allow commercial use on paid tiers; many free tiers explicitly prohibit it. Read the specific tool's terms before shipping to a paying client.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the longest clip I can generate for free?&lt;/strong&gt;&lt;br&gt;
Usually 5–10 seconds. Kling tops the free list at 10 seconds; most others cap at 4–6. Longer clips require paid tiers or extensions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the best free open-source AI video model?&lt;/strong&gt;&lt;br&gt;
Stable Video Diffusion from Stability AI. You can run it locally on a 12 GB+ GPU, or on a Hugging Face Space within free quota. Less polished than Runway or Luma, but no ongoing cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is "free AI video" actually free?&lt;/strong&gt;&lt;br&gt;
The platforms are free at the point of use within their limits. They are not free for the company hosting them — GPU rental costs $1–$3 per hour even at scale. Free tiers exist as marketing funnels for paid plans. Use them; don't pretend they are a business model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;If your goal is to make and ship video, pick &lt;strong&gt;Luma for quality&lt;/strong&gt;, &lt;strong&gt;Hailuo for volume&lt;/strong&gt;, and pair them with a free TTS and a free editor. If your goal is to build a "free AI video site" on top of someone else's API, the Facebook posts are selling you a fantasy — the unit economics don't work for anyone who isn't selling the course about it.&lt;/p&gt;

&lt;p&gt;For the rest of our free tools — calculators, encrypted chat, code playgrounds, and more — see &lt;a href="https://induwara.lk/tools" rel="noopener noreferrer"&gt;induwara.lk/tools&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>video</category>
      <category>freetools</category>
      <category>comparison</category>
    </item>
    <item>
      <title>Xiaomi 18 Fold: the foldable spec war just moved to silicon</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:36:31 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/xiaomi-18-fold-the-foldable-spec-war-just-moved-to-silicon-19nh</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/xiaomi-18-fold-the-foldable-spec-war-just-moved-to-silicon-19nh</guid>
      <description>&lt;p&gt;The &lt;strong&gt;Xiaomi 18 Fold&lt;/strong&gt; is the clearest sign yet that the foldable phone race has stopped being about hinges and started being about who owns the chip. Xiaomi has shipped a book-style foldable that beats the &lt;strong&gt;Samsung Galaxy Z Fold 8&lt;/strong&gt; on battery, camera and display brightness, running on a processor Xiaomi designed itself.&lt;/p&gt;

&lt;p&gt;Dominic Preston at The Verge got hands-on time with the phone at IFA and wrote it up in &lt;a href="https://www.theverge.com/tech/991008/xiaomi-18-fold-hands-on-impressions-specs-wide" rel="noopener noreferrer"&gt;Xiaomi's wide foldable promises more power than Samsung's&lt;/a&gt;. I'm not going to repeat his impressions. I want to look at why this phone matters even if you never buy one, and what it means if you build Android apps in Sri Lanka.&lt;/p&gt;




&lt;h2&gt;
  
  
  📊 The spec gap is real, and it favours Xiaomi
&lt;/h2&gt;

&lt;p&gt;Here is what The Verge reported, side by side with the Samsung numbers it quoted for comparison.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Spec&lt;/th&gt;
&lt;th&gt;Xiaomi 18 Fold&lt;/th&gt;
&lt;th&gt;Samsung Galaxy Z Fold 8 (per The Verge)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Chipset&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Xring O3&lt;/strong&gt; (Xiaomi in-house)&lt;/td&gt;
&lt;td&gt;Not stated in source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Arm Mali G2-Ultra NX&lt;/strong&gt; (first phone to use it)&lt;/td&gt;
&lt;td&gt;Not stated in source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Battery&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;6,000mAh&lt;/strong&gt; silicon-carbon&lt;/td&gt;
&lt;td&gt;4,800mAh&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Charging&lt;/td&gt;
&lt;td&gt;67W wired, 50W wireless&lt;/td&gt;
&lt;td&gt;Not stated in source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inner screen&lt;/td&gt;
&lt;td&gt;7.58-inch OLED, ~1.4:1&lt;/td&gt;
&lt;td&gt;Not stated in source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Outer screen&lt;/td&gt;
&lt;td&gt;5.38-inch OLED, ~1.4:1&lt;/td&gt;
&lt;td&gt;Not stated in source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Peak brightness&lt;/td&gt;
&lt;td&gt;Up to 4,000 nits&lt;/td&gt;
&lt;td&gt;Not stated in source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rear cameras&lt;/td&gt;
&lt;td&gt;200MP main (1/1.56-inch type) + 50MP 3.5x periscope + 50MP ultrawide&lt;/td&gt;
&lt;td&gt;Lacks a triple camera&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IP rating&lt;/td&gt;
&lt;td&gt;None announced&lt;/td&gt;
&lt;td&gt;Not stated in source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weight and thickness&lt;/td&gt;
&lt;td&gt;Slightly heavier and thicker than the Z Fold 8&lt;/td&gt;
&lt;td&gt;Lighter and thinner&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;China price&lt;/td&gt;
&lt;td&gt;¥10,999 (about $1,600)&lt;/td&gt;
&lt;td&gt;Approaching $2,000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I have left the Samsung column blank where The Verge did not give a number. I would rather show a gap than fill it with a guess.&lt;/p&gt;

&lt;p&gt;The pattern is obvious. Xiaomi wins on everything you can put on a spec sheet. Samsung keeps the wins that only show up in your hand: build quality, a less visible crease, lower weight, and water resistance, since Xiaomi has announced no rating at all. The Verge notes Oppo's &lt;strong&gt;Find N6&lt;/strong&gt; is still the only foldable to have all but removed the crease. Those boring advantages are what decide whether you still like a phone in month eighteen, and they are why I would not call Samsung finished.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; A &lt;strong&gt;25% larger battery&lt;/strong&gt;, a proper triple camera and a claimed flagship-class chip for roughly &lt;strong&gt;$400 less&lt;/strong&gt; than the Samsung. That is not an incremental refresh. That is a pricing challenge.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  ⚡ The Xring O3 is the story, not the fold
&lt;/h2&gt;

&lt;p&gt;The most important line in the source is easy to skim past: this is the &lt;strong&gt;first Xiaomi phone powered by its own Xring O3 chipset&lt;/strong&gt;, and Xiaomi claims Snapdragon 8 Elite-level performance from it.&lt;/p&gt;

&lt;p&gt;Think about who else ships their own phone silicon at flagship scale:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Apple&lt;/strong&gt; with the A-series.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google&lt;/strong&gt; with Tensor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Samsung&lt;/strong&gt; with Exynos, in some markets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Huawei&lt;/strong&gt; with Kirin.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Xiaomi joining that list changes its cost structure and its roadmap independence. A company that buys Qualcomm chips ships when Qualcomm ships. A company with its own silicon can decide what the GPU does, and this one debuts &lt;strong&gt;Arm's Mali G2-Ultra NX&lt;/strong&gt; with what The Verge describes as DLSS-style graphics acceleration on Android, meaning upscaling done by the GPU rather than brute-force rendering.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Bottom line for developers:&lt;/strong&gt; Upscaling on mobile is coming to Android the way it came to PC gaming. If you build anything GPU-heavy, whether a game, a 3D viewer or an on-device model, expect a new performance tier where the rendered resolution and the displayed resolution are not the same number.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;One caution. "Snapdragon 8 Elite-level" is Xiaomi's claim, quoted by The Verge, not a benchmark result. First-generation in-house silicon tends to be close on paper and warmer in practice. Wait for independent thermals before treating the O3 as proven.&lt;/p&gt;




&lt;h2&gt;
  
  
  📐 The 1.4:1 screen is a new layout problem for your app
&lt;/h2&gt;

&lt;p&gt;Xiaomi has followed Huawei and Samsung into the short, wide book-foldable shape, and then gone further. Both screens sit at roughly &lt;strong&gt;1.4:1&lt;/strong&gt;, which The Verge points out is almost exactly the ratio of international paper sizes and IMAX screens.&lt;/p&gt;

&lt;p&gt;For an Android developer, the practical consequence is this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Screen state&lt;/th&gt;
&lt;th&gt;Approx. shape&lt;/th&gt;
&lt;th&gt;What breaks&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Closed, 5.38-inch&lt;/td&gt;
&lt;td&gt;Short and wide, ~1.4:1&lt;/td&gt;
&lt;td&gt;Apps built for tall 20:9 phones get cramped vertically; fixed-height headers eat the screen&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open, 7.58-inch&lt;/td&gt;
&lt;td&gt;Near-square, ~1.4:1&lt;/td&gt;
&lt;td&gt;Single-column phone layouts waste half the width; tablet layouts assume wider&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mid-transition&lt;/td&gt;
&lt;td&gt;Resize event&lt;/td&gt;
&lt;td&gt;State loss if you don't handle configuration changes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Verge itself expects "some apps running awkwardly on the short screen". If you maintain an Android app, even a small one for a local business, three checks are worth doing now:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Test at ~1.4:1&lt;/strong&gt; in the emulator, both a small-phone size and a near-square tablet size. Don't rely on the default device profiles.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use window size classes&lt;/strong&gt; rather than the old "is this a tablet" boolean. The open 18 Fold is neither a phone nor a tablet by the old definitions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Handle configuration changes&lt;/strong&gt; properly so an open-to-close fold does not restart your activity and drop the user's form input.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of this is Xiaomi-specific. Samsung and Huawei are already here, and Apple may join them this week.&lt;/p&gt;




&lt;h2&gt;
  
  
  💰 What it would cost to bring one to Sri Lanka
&lt;/h2&gt;

&lt;p&gt;Here is the catch for anyone in Colombo already pricing this up: &lt;strong&gt;Xiaomi has not confirmed an international launch at all&lt;/strong&gt;. The Verge reports China-only preorders at ¥10,999, with sales starting September 10. That is the whole known availability picture.&lt;/p&gt;

&lt;p&gt;So the realistic Sri Lankan path, at least initially, is a grey import. Before you get excited by "about $1,600":&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The yuan-to-rupee figure moves daily. Run the ¥10,999 through the &lt;a href="https://induwara.lk/tools/lkr-exchange-rate" rel="noopener noreferrer"&gt;LKR exchange rate tool&lt;/a&gt; on the day you are actually deciding, not today.&lt;/li&gt;
&lt;li&gt;Sri Lanka taxes imported phones on arrival, and the duty structure is not a single flat percentage. The &lt;a href="https://induwara.lk/tools/sri-lanka-mobile-phone-import-tax-calculator" rel="noopener noreferrer"&gt;Sri Lanka mobile phone import tax calculator&lt;/a&gt; will give you the landed cost from a declared value.&lt;/li&gt;
&lt;li&gt;A China-only unit means a China ROM. Expect to deal with Google services, regional app stores and update timing yourself.&lt;/li&gt;
&lt;li&gt;No announced IP rating on a device you will use through a Sri Lankan monsoon is a real risk, not a spec-sheet footnote.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning:&lt;/strong&gt; Grey-market foldables have no local warranty and a hinge is the single most repair-prone part of any phone. Price in the cost of a replacement before you price in the savings.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;If you are a &lt;strong&gt;buyer in Sri Lanka&lt;/strong&gt;: wait. Nothing about this phone is purchasable here through a normal channel yet, and Apple's expected foldable announcement will reprice the whole category within days. Run the landed-cost numbers, but don't act on them until Xiaomi says the word "international".&lt;/p&gt;

&lt;p&gt;If you are an &lt;strong&gt;Android developer&lt;/strong&gt;: the short, wide, near-square foldable is now a three-vendor form factor, with possibly a fourth this week. Test at 1.4:1. Adopt window size classes. Stop treating the fold as a niche.&lt;/p&gt;

&lt;p&gt;If you are a &lt;strong&gt;student or builder watching the industry&lt;/strong&gt;: the real headline is not the 6,000mAh battery. It is that another major phone maker now designs its own flagship chip and is using it to undercut the market leader by a few hundred dollars. Hardware margins are moving to whoever owns the silicon. That trend will outlast this phone.&lt;/p&gt;

</description>
      <category>xiaomi</category>
      <category>foldablephones</category>
      <category>android</category>
    </item>
    <item>
      <title>Jensen Huang says AGI has arrived. Watch the GPU count</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Mon, 07 Sep 2026 22:20:47 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/jensen-huang-says-agi-has-arrived-watch-the-gpu-count-d65</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/jensen-huang-says-agi-has-arrived-watch-the-gpu-count-d65</guid>
      <description>&lt;p&gt;Nvidia CEO &lt;strong&gt;Jensen Huang&lt;/strong&gt; says &lt;strong&gt;"AGI has arrived"&lt;/strong&gt;, and the same short post tells you why to read it slowly. Huang also notes that OpenAI's new model was trained on Nvidia chips, and that &lt;strong&gt;400,000 GPUs&lt;/strong&gt; are coming online next. Business Insider &lt;a href="https://www.businessinsider.com/nvidia-jensen-huang-agi-openai-astra-ai-2026-9" rel="noopener noreferrer"&gt;reported the post on 6 September 2026&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I don't think Huang is being dishonest. I think this is a demand forecast wearing a lab coat, and the difference matters a lot if you're shipping software from Sri Lanka on a rupee budget.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 Four people, four definitions, one week
&lt;/h2&gt;

&lt;p&gt;The strongest evidence that "AGI" is not a technical milestone right now is that the people announcing it can't agree on what it means. All four of these statements come from the same news cycle:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Who&lt;/th&gt;
&lt;th&gt;What they said&lt;/th&gt;
&lt;th&gt;Where&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Jensen Huang&lt;/strong&gt;, Nvidia CEO&lt;/td&gt;
&lt;td&gt;"From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team."&lt;/td&gt;
&lt;td&gt;Post on X, Sunday&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Greg Brockman&lt;/strong&gt;, OpenAI president&lt;/td&gt;
&lt;td&gt;"Welcome to the AGI era." And: "For me personally, I do think we're there."&lt;/td&gt;
&lt;td&gt;Press call, Thursday&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Sam Altman&lt;/strong&gt;, OpenAI CEO&lt;/td&gt;
&lt;td&gt;AGI is "a very poorly defined term… it's like an irrelevant marketing term."&lt;/td&gt;
&lt;td&gt;"Sources" podcast&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Gary Marcus&lt;/strong&gt;, AI researcher and critic&lt;/td&gt;
&lt;td&gt;Huang "gave no evidence and no definitions, which feels to me like an effort at a takeover of a scientific question by corporate fiat."&lt;/td&gt;
&lt;td&gt;Substack, Sunday&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Marcus published his own 10-point definition of AGI and counted that &lt;strong&gt;Astra&lt;/strong&gt; meets one or two of them. His summary: "By conventional definitions, Astra still falls short."&lt;/p&gt;

&lt;p&gt;You do not have to pick a side in that argument. You only have to notice that the CEO of the company selling the term's most valuable meaning, and the CEO of the company that built the model, are describing the same word in opposite registers.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; When a word means "civilisational milestone" to the supplier and "irrelevant marketing term" to the buyer's own CEO, it is not a spec. Don't plan a roadmap around it.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  📊 The number in that post that isn't "AGI"
&lt;/h2&gt;

&lt;p&gt;Strip the adjective out of Huang's post and what's left is a capacity announcement. Here is the money side of the story, using only the figures Business Insider reported:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Figure&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Nvidia quarterly revenue reported in August&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$96.2 billion&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Change vs. same period a year earlier&lt;/td&gt;
&lt;td&gt;More than &lt;strong&gt;double&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data centre segment (includes AI chips)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$89 billion&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPUs Huang says are "coming online next"&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;400,000&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And the dependency runs both directions. In a March funding announcement, OpenAI called Nvidia "the foundation of our infrastructure" and said its "training fleet and the majority of our inference stack continue to run on Nvidia GPUs."&lt;/p&gt;

&lt;p&gt;So the person certifying that the milestone has been reached is also the supplier being paid for reaching it. That doesn't make him wrong. It does mean his post is not independent verification, and treating it as such is a category error.&lt;/p&gt;




&lt;h2&gt;
  
  
  🧪 "AGI" is not something you can build against
&lt;/h2&gt;

&lt;p&gt;OpenAI's own published definition is "highly autonomous systems that outperform humans at most economically valuable work." Try turning that into an acceptance test for your project. You can't. It has no threshold, no task list, and no measurement procedure.&lt;/p&gt;

&lt;p&gt;What OpenAI said about Astra specifically is more useful, because it's narrower: the company called it the world's "most intelligent and aligned model" and said it can perform "the most demanding professional work with unmatched speed, accuracy, and judgment." Vendor claims, but at least testable ones. Note also that Astra was announced Thursday and described as rolling out to customers this week, so on the day I'm writing this there is no broad independent record of how it behaves on real workloads.&lt;/p&gt;

&lt;p&gt;Which brings me to the only benchmark that matters for a small team: &lt;strong&gt;yours&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ Build the 30-task eval instead
&lt;/h2&gt;

&lt;p&gt;If you run a two-person shop in Colombo, or you're a final-year student picking a model for your project, this is the process I'd follow instead of reading launch coverage:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Write down 30 real tasks&lt;/strong&gt; from your actual product. Not riddles. The Sinhala-English support ticket you had to summarise, the invoice PDF you had to parse, the SQL you had to review.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write the pass condition first&lt;/strong&gt;, before you run anything. "Correct total, correct currency, no hallucinated line items." Vague criteria produce vague conclusions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run every candidate model on the same 30&lt;/strong&gt;, including the cheap ones and the open-weight ones you can self-host.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Record cost per solved task&lt;/strong&gt;, not cost per million tokens. A model that's 3× the price and solves twice as many tasks first-try may still be cheaper once you count your own debugging hours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-run it quarterly.&lt;/strong&gt; This is the part people skip, and it's the part that catches silent regressions and price changes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Our &lt;a href="https://induwara.lk/tools/ai-model-comparison" rel="noopener noreferrer"&gt;AI model comparison tool&lt;/a&gt; is a starting point for step 3, and the &lt;a href="https://induwara.lk/tools/ai-agent-cost-calculator" rel="noopener noreferrer"&gt;AI agent cost calculator&lt;/a&gt; helps with step 4 once you know your average tokens per task. Both are free and need no signup.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning:&lt;/strong&gt; The most expensive mistake available this week is rewriting a working pipeline around a model that shipped four days ago. Run your eval first. If the new model wins on your 30 tasks, migrate. If it doesn't, you saved a sprint.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  💰 What 400,000 GPUs might mean for a rupee budget
&lt;/h2&gt;

&lt;p&gt;Here is where the capacity number is genuinely more interesting to us than the AGI claim.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;More supply usually pushes inference prices down over time.&lt;/strong&gt; That has been the pattern in this market, and it's the pattern that has made frontier models usable on a Sri Lankan freelancer's budget at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is a forecast, not a price cut.&lt;/strong&gt; Huang said the GPUs are coming online. Nobody announced what they'll cost to rent or what per-token pricing will look like afterwards. Budget on today's published rates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacity gets allocated to the biggest buyers first.&lt;/strong&gt; If your workload is small and latency-tolerant, batch APIs and off-peak scheduling will do more for your bill this quarter than any datacentre build-out.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;The announcement changes the vocabulary, not your stack. Nothing about your rate limits, your budget, or your product's failure modes changed because a chip vendor used a three-letter word on a Sunday.&lt;/p&gt;

&lt;p&gt;Three things worth doing this week:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Keep your own eval set.&lt;/strong&gt; It is the only defence against both hype and FUD, and it costs one afternoon to build.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read who benefits before you read the claim.&lt;/strong&gt; Huang sells the compute. Brockman sells the model. Altman calls the term meaningless. Marcus sells the counter-argument. Everyone in that table has a position.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optimise the thing you control.&lt;/strong&gt; Prompt size, caching, batch scheduling, and picking the smallest model that passes your tests will move your monthly bill more than any frontier launch will.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If Astra really does clear the bar Brockman thinks it clears, you'll find out from your own 30 tasks within a month, and you'll know exactly which ones it changed. That's a better source than an X post from the person selling the GPUs.&lt;/p&gt;

</description>
      <category>aiindustry</category>
      <category>openai</category>
      <category>nvidia</category>
    </item>
    <item>
      <title>Pigeon: a signed pass for what your sub-agent may do</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Mon, 07 Sep 2026 01:56:51 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/pigeon-a-signed-pass-for-what-your-sub-agent-may-do-1b2p</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/pigeon-a-signed-pass-for-what-your-sub-agent-may-do-1b2p</guid>
      <description>&lt;p&gt;Almost every AI sub-agent permissions bug I have seen starts the same way: the parent agent spawns a child and hands it the same API key. The child now has everything the parent had. Deploy to production. Read the payments table. Merge to main.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pigeon&lt;/strong&gt; (&lt;a href="https://github.com/pigeonlabsHQ/pigeon" rel="noopener noreferrer"&gt;pigeonlabsHQ/pigeon on GitHub&lt;/a&gt;) is a small Python library that attacks exactly this. Instead of copying the key, you mint the child a &lt;strong&gt;Pigeon Pass&lt;/strong&gt;: a signed credential describing what it may do, and nothing wider.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔑 The real failure is key copying, not model misbehaviour
&lt;/h2&gt;

&lt;p&gt;Most of the agent-safety conversation is about the model: will it hallucinate, will it get prompt-injected, will it do something silly. That is the wrong layer to fix first. The thing that turns a silly action into an incident is that the silly actor was holding a key with full scope.&lt;/p&gt;

&lt;p&gt;Pigeon's framing is blunt and I think it is right:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Identity tells you who the agent is. Authority tells you what it may do.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;We have spent years getting identity right for humans and then handed our agents a shared secret with no scope at all. If you have ever stuffed a &lt;code&gt;scope&lt;/code&gt; claim into a token and hoped the downstream service checked it, this will feel familiar — you can &lt;a href="https://induwara.lk/tools/jwt-decoder" rel="noopener noreferrer"&gt;decode a JWT here&lt;/a&gt; and see how little most of them actually constrain.&lt;/p&gt;




&lt;h2&gt;
  
  
  📋 What is actually on a Pass
&lt;/h2&gt;

&lt;p&gt;A Pass is not a profile of JWT, macaroons, Biscuit, or UCAN. The spec says so explicitly. It is its own format, signed with &lt;strong&gt;Ed25519&lt;/strong&gt;, and every field is mandatory — unknown or missing fields are malformed, not ignored.&lt;/p&gt;

&lt;p&gt;The permission model is three-dimensional:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;th&gt;Child may&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;capabilities&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;["deploy", "open_pr"]&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;only subset, exact match, no wildcards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;resources&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;["environment:staging", "repo:acme/api"]&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;only narrower patterns; trailing &lt;code&gt;*&lt;/code&gt; allowed, &lt;code&gt;*&lt;/code&gt; alone is root-only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;constraints&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;{"max_deploys_per_hour": 3}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;keep every parent dimension; may add more&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The core invariant is one sentence: a child must never carry more effective authority than its parent. Try it anyway and you get a &lt;code&gt;DelegationError&lt;/code&gt; with &lt;code&gt;reason_code == "PRIVILEGE_ESCALATION"&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pigeon&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;delegate&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;grant&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;verify&lt;/span&gt;

&lt;span class="n"&gt;parent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;grant&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;subject&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agent:orchestrator&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;capabilities&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deploy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;open_pr&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;resources&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;environment:staging&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;repo:acme/api&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;worker&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;delegate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;subject&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agent:pr-bot&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                  &lt;span class="n"&gt;capabilities&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;open_pr&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;resources&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;repo:acme/api&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="n"&gt;denied&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;worker&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deploy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;resource&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;environment:staging&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;denied&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reason_code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CAPABILITY_NOT_GRANTED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The detail I like most: &lt;code&gt;verify&lt;/code&gt; never returns a bare boolean. A denial carries a reason code, a message, and the &lt;code&gt;requested&lt;/code&gt; vs &lt;code&gt;allowed&lt;/code&gt; comparison that failed. That is the difference between a library you can debug at 2am and one you rip out.&lt;/p&gt;




&lt;h2&gt;
  
  
  ⚠️ The security file is the reason I trust it
&lt;/h2&gt;

&lt;p&gt;Most agent-security projects oversell. Pigeon's &lt;code&gt;SECURITY.md&lt;/code&gt; does the opposite, and it is the strongest signal in the whole repo.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pigeon does&lt;/th&gt;
&lt;th&gt;Pigeon does not&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fails closed when narrowing cannot be proven&lt;/td&gt;
&lt;td&gt;Stop prompt injection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Verifies the whole chain, not just the leaf&lt;/td&gt;
&lt;td&gt;Enforce a dimension you did not write&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Invalidates the signature on any tampered field&lt;/td&gt;
&lt;td&gt;See revocations issued after an offline Pass was minted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Counts &lt;code&gt;rate&lt;/code&gt; and &lt;code&gt;count&lt;/code&gt; against every ancestor&lt;/td&gt;
&lt;td&gt;Manage, rotate, or recover your keys&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three admissions stand out.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Same-process crypto is nearly theatre.&lt;/strong&gt; If the issuer and verifier are the same process, the signature adds little over a plain data-structure check. It earns its keep when the Pass crosses a process, machine, or organisation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;v0.1 ships no durable store.&lt;/strong&gt; In-memory replay, revocation, and usage stores vanish on restart, so rate and count budgets reset. That is more permissive than most people would assume.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The recommended default TTL is one hour&lt;/strong&gt;, because short expiry is the only revocation an offline verifier really has.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; the enforcement point is the whole product. As the README puts it, &lt;em&gt;"If the runner never calls &lt;code&gt;verify&lt;/code&gt;, the Pass is decoration."&lt;/em&gt; Signing changes nothing if the tool runs regardless of the answer.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🛠️ Why this matters if you are building on a free tier
&lt;/h2&gt;

&lt;p&gt;Here is the part that makes it relevant for a solo developer or a three-person team in Colombo rather than a security team at a bank.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;There is no server.&lt;/strong&gt; No control plane to host, no per-seat pricing, no vendor. You change two places in code you already wrote: the spawn site and the tool site.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is MIT-licensed Python 3.12+&lt;/strong&gt;, installed with &lt;code&gt;git clone&lt;/code&gt; and &lt;code&gt;pip install .&lt;/code&gt;. Total cost: nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It forces you to write the policy down.&lt;/strong&gt; Most of us have never actually enumerated what our automation is allowed to touch. Filling in &lt;code&gt;capabilities&lt;/code&gt;, &lt;code&gt;resources&lt;/code&gt;, and &lt;code&gt;constraints&lt;/code&gt; is uncomfortable in a productive way.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That third point is the real value, and it survives even if you never ship Pigeon. The exercise of listing what a sub-agent may do is worth an afternoon on its own. I run an autonomous build pipeline on this site, and the honest answer to "what may the build stage touch?" was, for a long time, "whatever the process could reach."&lt;/p&gt;

&lt;p&gt;There is also an &lt;strong&gt;MCP middleware&lt;/strong&gt; helper: the client mints a narrower Pass per tool call, the server verifies before the handler runs. The repo is careful to say this is an enforcement point and not part of the MCP specification. Given how many people are now wiring MCP servers into agents without any per-tool boundary, that is a pattern worth copying by hand even if you skip the library.&lt;/p&gt;




&lt;h2&gt;
  
  
  🤔 Where I would push back
&lt;/h2&gt;

&lt;p&gt;Two honest reservations.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;v0.1 means v0.1.&lt;/strong&gt; Nine constraint ops, no persistent store, and a protocol the author invites people to find escalation bugs in. I would not put it in front of a payments flow this month.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An omitted dimension is not enforced.&lt;/strong&gt; If you did not put &lt;code&gt;environment:production&lt;/code&gt; out of reach, production is in reach. The protocol will not invent a policy you did not sign. Your Pass is exactly as good as your imagination about what could go wrong, which is a familiar and uncomfortable property of every allowlist ever written.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What this means for you
&lt;/h2&gt;

&lt;p&gt;If you are running agents that spawn other agents, do this today regardless of whether you adopt Pigeon:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Stop passing the parent key down.&lt;/strong&gt; Keep the real secret on the runner.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write the allowed action list somewhere machine-readable&lt;/strong&gt;, even if it starts as a dict.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put the check where the side effect happens&lt;/strong&gt;, not where the agent is spawned. That is the only place it counts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set an expiry in hours, not weeks&lt;/strong&gt;, on anything you do mint.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make denials explain themselves&lt;/strong&gt; — reason code plus requested-vs-allowed, or you will disable the check the first time it blocks you unfairly.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Pigeon calls itself a small primitive, not a platform. That modesty is the point. The idea it encodes, that authority should narrow every time it is delegated, is older than AI agents and does not need a library to be useful. But having it in twenty lines of Python, MIT-licensed and serverless, removes the last excuse for not doing it.&lt;/p&gt;

</description>
      <category>aiagents</category>
      <category>security</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Oura's IPO shows the smart ring moat isn't the sensors</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Sun, 06 Sep 2026 01:44:35 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/ouras-ipo-shows-the-smart-ring-moat-isnt-the-sensors-334h</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/ouras-ipo-shows-the-smart-ring-moat-isnt-the-sensors-334h</guid>
      <description>&lt;p&gt;The &lt;strong&gt;smart ring&lt;/strong&gt; market just got its first real financial disclosure, and it says something different from what the product marketing says. Oura filed to go public on &lt;strong&gt;September 3, 2026&lt;/strong&gt;, and the numbers in that filing tell you the company is not primarily a sensor business.&lt;/p&gt;

&lt;p&gt;TechCrunch's &lt;a href="https://techcrunch.com/2026/09/05/oura-is-going-public-but-these-smart-ring-companies-are-coming-for-its-crown/" rel="noopener noreferrer"&gt;rundown of the challengers coming for Oura's crown&lt;/a&gt; lists six players trying different angles. I read that list and saw three separate moats, only one of which is technical.&lt;/p&gt;




&lt;h2&gt;
  
  
  📊 The arithmetic in the filing
&lt;/h2&gt;

&lt;p&gt;The disclosed figures are worth sitting with before looking at any competitor spec sheet:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Figure&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Revenue, nine months to June 30&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$1.21 billion&lt;/strong&gt; (nearly doubled)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rings sold, past year&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.6 million&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Paid members&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~5 million&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latest product&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Oura Ring 5&lt;/strong&gt;, billed as its slimmest and lightest&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Look at the last two rows together. Paid members exceed rings sold in the past year by roughly 1.4 million. That gap is people who bought hardware in an earlier year and are still paying every month. Hardware got them in; the subscription is what compounds.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; A competitor can match Oura's sensors in one hardware cycle. Matching a base of five million people already in the habit of paying a recurring fee takes years, and no spec sheet shortens it.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  ⚖️ The second moat is a patent docket
&lt;/h2&gt;

&lt;p&gt;The most instructive fact in the whole story has nothing to do with heart rate accuracy. &lt;strong&gt;Ultrahuman&lt;/strong&gt;'s US business was disrupted in &lt;strong&gt;October 2025&lt;/strong&gt; after an &lt;strong&gt;ITC ruling&lt;/strong&gt; went Oura's way in a patent dispute. Not a bad review, not a failed sensor. A trade ruling.&lt;/p&gt;

&lt;p&gt;For anyone in Sri Lanka building hardware or a hardware-adjacent product aimed at Western markets, that is the lesson to file away:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your technical roadmap can be correct and your market access can still be switched off by a body you never presented to.&lt;/li&gt;
&lt;li&gt;Freedom-to-operate work is not a legal formality you do after product-market fit. In a crowded sensor category it &lt;em&gt;is&lt;/em&gt; part of the design constraint.&lt;/li&gt;
&lt;li&gt;The incumbent with revenue can afford to litigate for years. You probably cannot.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ultrahuman has since raised &lt;strong&gt;$70 million&lt;/strong&gt; with Qualcomm venture backing and is shipping the &lt;strong&gt;Ring Pro at $479&lt;/strong&gt; in the US from mid-September. Money solves the appeal; it does not give back the year.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔋 Squeezing a dual-core CPU onto a finger
&lt;/h2&gt;

&lt;p&gt;Here is the part that actually interests me as an engineering problem. The Ultrahuman Ring Pro carries a &lt;strong&gt;dual-core processor&lt;/strong&gt; and redesigned heart-rate sensing, with the stated aim of running on-device software for AI interactions and even games.&lt;/p&gt;

&lt;p&gt;A ring is the tightest power and thermal budget in consumer hardware. There is no room for a fan, the battery is a sliver, and the whole thing sits against skin, so waste heat is a comfort problem before it is a reliability problem. Putting general compute in there is not a gimmick decision. It is a margin decision:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Cloud inference costs money per user, per day, forever.&lt;/strong&gt; A subscription business with millions of members pays that bill every single month.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On-device inference costs money once&lt;/strong&gt;, in silicon, at manufacture.&lt;/li&gt;
&lt;li&gt;Latency and privacy improvements are real, but they are the pleasant side effects. The spreadsheet is the driver.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That trade should feel familiar if you have built anything on a free tier. It is the same reason we run several tools on induwara.lk fully client-side rather than paying for a server round trip per request. When per-user marginal cost is your enemy, you push work to the edge, whether that edge is a browser or a titanium band.&lt;/p&gt;




&lt;h2&gt;
  
  
  🌐 The field, and what each one is betting on
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Company&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;th&gt;Availability&lt;/th&gt;
&lt;th&gt;The bet&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Oura&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not disclosed in filing coverage&lt;/td&gt;
&lt;td&gt;Shipping (Ring 5)&lt;/td&gt;
&lt;td&gt;Installed base + subscription&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Ultrahuman&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$479&lt;/strong&gt; (Ring Pro)&lt;/td&gt;
&lt;td&gt;US from mid-Sept 2026&lt;/td&gt;
&lt;td&gt;On-device compute, apps, games&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;RingConn&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;From &lt;strong&gt;$349&lt;/strong&gt; (Gen 3)&lt;/td&gt;
&lt;td&gt;Since &lt;strong&gt;May 2026&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;Price, plus vascular health insights&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Samsung&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$399&lt;/strong&gt; (Galaxy Ring, 2024)&lt;/td&gt;
&lt;td&gt;Shipping&lt;/td&gt;
&lt;td&gt;Bundling into the Galaxy ecosystem&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Circular&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not announced (Ring 3)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Early 2027&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Medical-grade claims + NFC payments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Dreame&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not announced&lt;/td&gt;
&lt;td&gt;Not announced&lt;/td&gt;
&lt;td&gt;Haptics and an on-ring touchpad&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two of these are notable for what they lack or promise. Samsung's Galaxy Ring, at $399 since 2024, still does &lt;strong&gt;not&lt;/strong&gt; do sleep apnea or AFib detection, which is a striking gap for the company with the most distribution. Circular is going the other way entirely: the Ring 3 Pro claims &lt;strong&gt;FDA-cleared ECG for AFib detection&lt;/strong&gt;, plus blood pressure and glucose tracking.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If Circular ships glucose tracking from a ring on the stated timeline, that is a bigger story than the IPO. Treat unshipped medical claims with the scepticism you would apply to any pre-launch spec sheet. Early 2027 is not a shipping date, it is a window.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  💰 What one of these actually costs in Sri Lanka
&lt;/h2&gt;

&lt;p&gt;The listed price is the smallest part of the bill here. A realistic total cost of ownership has four lines:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The hardware.&lt;/strong&gt; $349 to $479 for the ones with announced prices.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Import duty and clearance&lt;/strong&gt;, if it comes in as a personal import or in your baggage. Rates vary by category and consignment value, so check the current schedule with our &lt;a href="https://induwara.lk/tools/sri-lanka-customs-baggage-allowance-calculator" rel="noopener noreferrer"&gt;Sri Lanka customs baggage allowance calculator&lt;/a&gt; rather than guessing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The subscription.&lt;/strong&gt; Oura's five million paid members are paying for something recurring. Budget for a monthly fee in LKR, converted at whatever the rate is on renewal day, not on purchase day. Our &lt;a href="https://induwara.lk/tools/currency-converter" rel="noopener noreferrer"&gt;currency converter&lt;/a&gt; is the boring but necessary step here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sizing and returns.&lt;/strong&gt; A ring that does not fit is a paperweight, and cross-border returns from Sri Lanka are painful and slow.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Before you spend that, be honest about what you would do with the data. If the goal is better sleep or smarter training, most of the value comes from acting on simple numbers you can get for free. A &lt;a href="https://induwara.lk/tools/sleep-cycle-calculator" rel="noopener noreferrer"&gt;sleep cycle calculator&lt;/a&gt; will tell you when to go to bed tonight, and a &lt;a href="https://induwara.lk/tools/heart-rate-zone-calculator" rel="noopener noreferrer"&gt;heart rate zone calculator&lt;/a&gt; will tell you what pace to run at, both for zero rupees. A ring measures adherence. It does not create it.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If you are buying:&lt;/strong&gt; wait. Ultrahuman ships this month, Circular's medical-claim ring lands early 2027, and RingConn is already at $349. Prices in this category are heading down while the feature floor is heading up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you are building hardware:&lt;/strong&gt; budget for freedom-to-operate research the way you budget for tooling. The ITC ruling took Ultrahuman out of its biggest market without a single technical failure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you are building software:&lt;/strong&gt; the on-device compute shift in a $479 ring is the same economics that should push your own inference to the client. Recurring per-user cost is what kills small-team margins, not the initial build.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you are studying this space:&lt;/strong&gt; the filing is the most useful public document in consumer wearables right now. Subscriptions carried it, not sensors.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rings are converging. The business models are not, and that is where this fight gets decided.&lt;/p&gt;

</description>
      <category>wearables</category>
      <category>hardware</category>
      <category>embedded</category>
    </item>
    <item>
      <title>Crusoe's $30B valuation is a bet on power, not AI</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Sat, 05 Sep 2026 05:05:47 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/crusoes-30b-valuation-is-a-bet-on-power-not-ai-2pa0</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/crusoes-30b-valuation-is-a-bet-on-power-not-ai-2pa0</guid>
      <description>&lt;p&gt;&lt;strong&gt;Crusoe's reported $3 billion raise at a $30 billion valuation&lt;/strong&gt; is not really an AI story. It is an electricity story wearing an AI jacket. &lt;a href="https://techcrunch.com/2026/09/03/crusoe-reportedly-raises-3b-at-a-30b-valuation/" rel="noopener noreferrer"&gt;TechCrunch reports&lt;/a&gt; that the round came together after Crusoe reportedly signed a roughly &lt;strong&gt;$13 billion, five-year cloud contract with Jane Street&lt;/strong&gt;, the quantitative trading firm.&lt;/p&gt;

&lt;p&gt;Eleven months ago the same company was valued at $10 billion. Its software did not get three times better in eleven months. What changed is who has the power contracts, the land, and the delivery slots.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔌 The valuation tripled because a customer signed, not because the tech improved
&lt;/h2&gt;

&lt;p&gt;Line the two rounds up next to each other and the mechanism is obvious.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;October 2025&lt;/th&gt;
&lt;th&gt;September 2026 (reported)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Amount raised&lt;/td&gt;
&lt;td&gt;$1.38 billion&lt;/td&gt;
&lt;td&gt;$3 billion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Valuation&lt;/td&gt;
&lt;td&gt;$10 billion&lt;/td&gt;
&lt;td&gt;$30 billion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Named investors&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Atreides Management, Valor Equity Partners (co-leads), Mubadala Capital&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trigger&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;~$13B, 5-year Jane Street cloud contract&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is a 3× valuation move in under a year, and the reported reason is a single anchor tenant. Crusoe already builds hyperscale campuses for &lt;strong&gt;Oracle, OpenAI, Meta and Microsoft&lt;/strong&gt;, and TechCrunch reports the company has met with Goldman Sachs and Morgan Stanley about a possible near-term IPO.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; In AI infrastructure right now, a signed multi-year offtake contract is worth more than any technical differentiation. The market is pricing contracted revenue, not cleverness.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I want to flag the hedge properly: every one of these numbers is reported, not confirmed by the company. Treat them as directionally useful, not as filings.&lt;/p&gt;




&lt;h2&gt;
  
  
  💰 Jane Street tells you who is actually paying for AI
&lt;/h2&gt;

&lt;p&gt;Everyone talks about chatbots. The biggest single AI compute contract in this story was signed by a proprietary trading firm.&lt;/p&gt;

&lt;p&gt;That is worth sitting with if you are choosing what to build or where to work:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The buyers with real budgets have a measurable loss function.&lt;/strong&gt; A trading firm knows exactly what a millisecond or a better signal is worth. It does not need to be convinced that AI is the future.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Boring verticals fund the boom.&lt;/strong&gt; Finance, logistics, insurance, energy. Not consumer apps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compute is being pre-sold in five-year blocks.&lt;/strong&gt; Capacity going to a customer like this is capacity that does not show up on a public cloud spot market at a discount.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For a Sri Lankan engineer, the practical read is that the demand side of AI is enterprise and quantitative, not consumer. If you are building a portfolio to get remote contracts, a project that shows you can cut an inference bill or wire a model into a real workflow beats another chat UI.&lt;/p&gt;




&lt;h2&gt;
  
  
  🌏 Sri Lanka is not winning the data centre race, and does not need to
&lt;/h2&gt;

&lt;p&gt;Crusoe started in 2018 mining crypto using &lt;strong&gt;flared natural gas&lt;/strong&gt; — energy that was being burned off and wasted. That is the whole trick. Find stranded, near-worthless power, put compute next to it, sell the compute. Then in 2024–2026, point the same physical asset at a customer paying AI prices instead of crypto prices.&lt;/p&gt;

&lt;p&gt;That arbitrage does not exist here. We do not have stranded gas fields, our grid is constrained, and industrial power is expensive. Anyone pitching you a "Sri Lanka AI data centre" play should be asked one question first: where is your cheap power coming from, and is it contracted?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The lesson from Crusoe is not "build data centres." It is "look at what asset you actually own, and find the customer who values it most." Crusoe's asset was never the mining rigs. It was proximity to wasted energy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The asset most of us own is different: low-cost, high-skill engineering time in a timezone that overlaps both Europe and Asia, billed in dollars. That is the arbitrage worth pressing. If you are working out what your rate needs to be after conversion and platform fees, our &lt;a href="https://induwara.lk/tools/freelancer-usd-lkr-calculator" rel="noopener noreferrer"&gt;freelancer USD-LKR earnings calculator&lt;/a&gt; does that maths.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ Build like compute stays expensive, because it will
&lt;/h2&gt;

&lt;p&gt;If the biggest buyers are locking capacity into five-year contracts, do not plan your side project around GPU prices collapsing. Plan around them staying annoying. Here is how I actually structure things:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;What it buys you&lt;/th&gt;
&lt;th&gt;When it fails&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Run inference in the browser (WASM / WebGPU)&lt;/td&gt;
&lt;td&gt;Zero server cost, zero upload, real privacy&lt;/td&gt;
&lt;td&gt;Big models, weak devices&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Use the smallest model that passes your eval&lt;/td&gt;
&lt;td&gt;Often 5–20× cheaper than the flagship&lt;/td&gt;
&lt;td&gt;Genuinely hard reasoning tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache aggressively on prompt and result&lt;/td&gt;
&lt;td&gt;Repeat traffic costs almost nothing&lt;/td&gt;
&lt;td&gt;High-variance user input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch overnight instead of real-time&lt;/td&gt;
&lt;td&gt;Cheaper tiers, no idle capacity&lt;/td&gt;
&lt;td&gt;Anything interactive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speculative decoding on self-hosted models&lt;/td&gt;
&lt;td&gt;Same output, fewer wall-clock seconds&lt;/td&gt;
&lt;td&gt;Poor draft-model match&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Most of the tools on this site are built on the first line of that table. They run in your browser, on your machine, with nothing uploaded, which is why they can be free with no signup. That design choice was originally about privacy. It turns out to also be the only cost structure that survives a compute market where a trading firm can outbid you by nine orders of magnitude.&lt;/p&gt;

&lt;p&gt;If you are self-hosting and want to sanity-check whether a draft model is worth the complexity, our &lt;a href="https://induwara.lk/tools/ai-speculative-decoding-calculator" rel="noopener noreferrer"&gt;speculative decoding speedup calculator&lt;/a&gt; will give you the expected speedup before you spend a weekend on it.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;Three things I would take from this, if I were you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stop competing on layers you cannot fund.&lt;/strong&gt; Foundation models and data centres are capital games measured in billions. Application and workflow layers are still open, and they are where the margin ends up once infrastructure commoditises.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Follow the money to the boring buyers.&lt;/strong&gt; A quant firm reportedly committing $13 billion over five years says more about where AI revenue is than any consumer launch this year.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Design for expensive compute.&lt;/strong&gt; Client-side execution, small models, caching and batching are not compromises. On this side of the world they are the design.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Crusoe spent 2018 to 2026 doing one consistent thing: standing next to cheap energy and selling what came out. Work out what you are standing next to, then find the customer who values it most. That part scales down to a one-person team perfectly well.&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>startup</category>
      <category>computecosts</category>
    </item>
    <item>
      <title>Belkin's semi-solid-state power banks: is 3x life worth $70?</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Fri, 04 Sep 2026 08:48:40 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/belkins-semi-solid-state-power-banks-is-3x-life-worth-70-3d1a</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/belkins-semi-solid-state-power-banks-is-3x-life-worth-70-3d1a</guid>
      <description>&lt;p&gt;Belkin's first &lt;strong&gt;semi-solid-state power bank&lt;/strong&gt; was announced at IFA this morning, and the headline spec isn't capacity or wattage. It's lifespan. Writing in &lt;a href="https://www.theverge.com/tech/988248/belkin-ultracharge-pro-boostsolid-cell-power-bank-battery-safer" rel="noopener noreferrer"&gt;The Verge&lt;/a&gt;, Andrew Liszewski reports that Belkin claims the new banks are "engineered to last up to 3x longer than other power banks."&lt;/p&gt;

&lt;p&gt;That's an unusual thing to sell. Nobody markets a power bank on how long it survives before it becomes a paperweight. For anyone who charges a phone through a full day in 30°C heat, that's the only spec that has ever really mattered.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔋 What Belkin actually announced
&lt;/h2&gt;

&lt;p&gt;Two products, both using what the company calls &lt;strong&gt;BoostSolid Cell&lt;/strong&gt;: a gel-like layer added alongside the liquid electrolyte you'd find in a normal lithium-ion cell.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Spec&lt;/th&gt;
&lt;th&gt;UltraCharge Pro Slim Magnetic 5K Qi2&lt;/th&gt;
&lt;th&gt;UltraCharge Pro 10K 60W + Display&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Price&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$69.99&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$89.99&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Capacity&lt;/td&gt;
&lt;td&gt;5,000mAh&lt;/td&gt;
&lt;td&gt;10,000mAh&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wireless output&lt;/td&gt;
&lt;td&gt;15W Qi2, magnetic&lt;/td&gt;
&lt;td&gt;Not stated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max wired output&lt;/td&gt;
&lt;td&gt;22.5W over USB-C&lt;/td&gt;
&lt;td&gt;60W over USB-C, single device&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Two devices at once&lt;/td&gt;
&lt;td&gt;5W wireless / 12W USB-C&lt;/td&gt;
&lt;td&gt;Output drops; figures not given&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ports&lt;/td&gt;
&lt;td&gt;USB-C&lt;/td&gt;
&lt;td&gt;2× USB-C, 1× USB-A&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Display&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Charge level, power output, internal temperature&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Size claim&lt;/td&gt;
&lt;td&gt;40% slimmer than other 5,000mAh banks&lt;/td&gt;
&lt;td&gt;Slimmer than similar-capacity rivals&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Belkin is late to this. The report names &lt;strong&gt;BMX&lt;/strong&gt;, &lt;strong&gt;Kuxiu&lt;/strong&gt; and &lt;strong&gt;Zens&lt;/strong&gt; as having shipped semi-solid-state banks already. What Belkin brings is distribution and a warranty department, which for a battery you'll keep in a laptop bag is not nothing.&lt;/p&gt;




&lt;h2&gt;
  
  
  🌡️ Heat is the spec that matters in a tropical climate
&lt;/h2&gt;

&lt;p&gt;The three claims Belkin makes for the gel layer are: longer usable life, reduced risk of overheating that leads to swelling or fire, and &lt;strong&gt;consistent power delivery across a wider temperature range&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That third one is the quiet winner if you live anywhere warm. Heat and a permanently full charge are the two things that age a lithium-ion cell fastest, and a power bank in Sri Lanka gets both. It sits in a bag on a bus. It sits on a desk beside a window. It gets topped up to 100% and left there for a week.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; The interesting product here isn't the battery, it's the display. A readout showing &lt;em&gt;internal temperature&lt;/em&gt; turns a sealed black box into an instrument. You can actually see when it's cooking and stop charging it. No spec sheet gives you that.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The safety angle is real too, but keep it in proportion. Semi-solid-state still contains liquid electrolyte. It's a mitigation, not a removal of the failure mode.&lt;/p&gt;




&lt;h2&gt;
  
  
  💰 Do the maths on cost per year, not cost per mAh
&lt;/h2&gt;

&lt;p&gt;Here's where I'd push back on the marketing. Take Belkin's 3x claim entirely at face value and it still doesn't obviously win on money.&lt;/p&gt;

&lt;p&gt;The figures below use Belkin's real prices and its own 3x multiplier. The comparison bank's price and life are illustrative — pick your own numbers, the method is the point.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Typical 5,000mAh bank &lt;em&gt;(illustrative)&lt;/em&gt;
&lt;/th&gt;
&lt;th&gt;Belkin 5K BoostSolid&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Purchase price&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;td&gt;$69.99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assumed usable life&lt;/td&gt;
&lt;td&gt;2 years&lt;/td&gt;
&lt;td&gt;6 years (3× the baseline)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cost per year&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$10.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$11.66&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Roughly break-even. At $69.99 against a $20 bank, three times the life buys you convenience and one less dead battery in a landfill, not savings. The premium only pays for itself if the thing you're replacing costs more than about $23, or if the 3x turns out to be conservative.&lt;/p&gt;

&lt;p&gt;And those are US dollar prices. By the time either of these reaches a shelf in Colombo you're paying import duty, VAT and a retailer margin on top. Before you compare a local listing to that $69.99, run the actual conversion — our &lt;a href="https://induwara.lk/tools/lkr-exchange-rate" rel="noopener noreferrer"&gt;LKR exchange rate page&lt;/a&gt; has the current mid-market rate, and the &lt;a href="https://induwara.lk/tools/currency-converter" rel="noopener noreferrer"&gt;currency converter&lt;/a&gt; will do the arithmetic.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 Four things the announcement doesn't answer
&lt;/h2&gt;

&lt;p&gt;I'd want these before paying a 3x premium for a 3x claim:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Three times longer than what?&lt;/strong&gt; No baseline is named. Against a well-made bank or against the cheapest thing on a shelf? The multiplier is meaningless without the denominator.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What's the cycle rating?&lt;/strong&gt; Battery life is normally quoted as capacity retention after N cycles. No such figure appears.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the warranty as long as the claim?&lt;/strong&gt; A six-year lifespan claim backed by a two-year warranty is a marketing number, not an engineering one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What happens to output as it ages?&lt;/strong&gt; A cell that holds capacity but loses its ability to sustain 60W is still a downgrade in year four.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;None of this makes the claim false. It makes it unverified, which is a different thing, and the honest position until someone runs these for a couple of years.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;If you're a student or a freelancer buying one power bank to carry every day, this is worth watching but probably not worth being first on.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Don't upgrade a working bank.&lt;/strong&gt; The greenest and cheapest battery is the one you already own. Belkin's own pitch is about replacement frequency, which is an argument for keeping things longer, not buying sooner.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Buy the display, not the chemistry.&lt;/strong&gt; If you do spend up, the 10K model's temperature readout is the feature you'll use. The gel layer is a claim; the screen is a fact.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you're replacing a swollen bank right now,&lt;/strong&gt; heat tolerance is a genuine reason to pay more here, more so than for a reader in a cold climate. Just don't expect the money back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Watch the second wave.&lt;/strong&gt; BMX, Kuxiu and Zens got here first, Belkin makes it mainstream, and mainstream is usually when prices start moving.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The broader signal matters more than either product. Solid-state batteries have been perpetually two years away for a decade. Semi-solid-state is the compromise that actually ships: keep the liquid electrolyte, add a gel, take the safety and longevity gains you can get today. That pattern, shipping the 70% version instead of waiting for the perfect one, is a reasonable way to build most things.&lt;/p&gt;

</description>
      <category>hardware</category>
      <category>batteries</category>
      <category>buyingadvice</category>
    </item>
    <item>
      <title>LLM inference tuning: which knobs are free, which cost you</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Wed, 02 Sep 2026 16:28:10 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/llm-inference-tuning-which-knobs-are-free-which-cost-you-1be6</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/llm-inference-tuning-which-knobs-are-free-which-cost-you-1be6</guid>
      <description>&lt;p&gt;&lt;strong&gt;LLM inference cost optimization&lt;/strong&gt; is usually sold to you as a pile of tricks: quantize this, batch that, turn on speculative decoding. Baseten's Philip Kiely published a piece on &lt;strong&gt;&lt;a href="https://www.baseten.co/blog/the-efficient-frontier-of-llm-inference/" rel="noopener noreferrer"&gt;the efficient frontier of LLM inference&lt;/a&gt;&lt;/strong&gt; (updated 1 September 2026) that does something more useful than adding another trick to the pile.&lt;/p&gt;

&lt;p&gt;It sorts the tricks into two buckets. That sorting is the part worth stealing, especially if your entire AI budget is smaller than one engineer's monthly salary.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 Two buckets: trades and gifts
&lt;/h2&gt;

&lt;p&gt;The article's framing is an &lt;strong&gt;efficient frontier&lt;/strong&gt; — "the range of optimal combinations when trading off between two valuable outcomes in a resource-constrained environment." For inference, that's usually latency against throughput.&lt;/p&gt;

&lt;p&gt;Some techniques move you &lt;em&gt;along&lt;/em&gt; that curve. You buy throughput by selling per-user speed. Others push the &lt;em&gt;whole curve&lt;/em&gt; outward, so you get more of both.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Technique&lt;/th&gt;
&lt;th&gt;Bucket&lt;/th&gt;
&lt;th&gt;What you give up&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Batch size&lt;/td&gt;
&lt;td&gt;Trade&lt;/td&gt;
&lt;td&gt;Bigger batches cut cost per token, worsen per-user latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tensor parallelism (TP)&lt;/td&gt;
&lt;td&gt;Trade&lt;/td&gt;
&lt;td&gt;Fast over NVLink, but all-to-all ops are expensive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Expert parallelism (EP)&lt;/td&gt;
&lt;td&gt;Trade&lt;/td&gt;
&lt;td&gt;Low EP is lower latency; wide EP across racks is higher throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Attention data parallelism (ADP)&lt;/td&gt;
&lt;td&gt;Trade&lt;/td&gt;
&lt;td&gt;Throughput up, per-request speed down&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quantization&lt;/td&gt;
&lt;td&gt;Mostly gift&lt;/td&gt;
&lt;td&gt;Latency &lt;em&gt;and&lt;/em&gt; throughput improve; quality is the second tradeoff&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kernel optimization&lt;/td&gt;
&lt;td&gt;Gift&lt;/td&gt;
&lt;td&gt;Nothing, if the kernel is correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speculative decoding&lt;/td&gt;
&lt;td&gt;Gift&lt;/td&gt;
&lt;td&gt;Draft model competes for the same GPU resources&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prefill/decode disaggregation&lt;/td&gt;
&lt;td&gt;Gift&lt;/td&gt;
&lt;td&gt;Complexity; mainly buys throughput at flat latency&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; Before you spend a week on an inference optimization, ask which bucket it's in. A trade needs a decision from you about what you're willing to lose. A gift needs engineering time and nothing else.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  📊 Most of those knobs are not yours to turn
&lt;/h2&gt;

&lt;p&gt;Here's the honest bit the article doesn't need to say, because it's a serving company writing for people who run their own serving. If you call an API, you own almost none of this.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;How you deploy&lt;/th&gt;
&lt;th&gt;Batch size&lt;/th&gt;
&lt;th&gt;Quantization&lt;/th&gt;
&lt;th&gt;Parallelism&lt;/th&gt;
&lt;th&gt;Spec decoding&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Closed API (Anthropic, OpenAI, Gemini)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open-weights via a hosted provider&lt;/td&gt;
&lt;td&gt;Sometimes&lt;/td&gt;
&lt;td&gt;Choose the endpoint&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Provider's choice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rented GPU (vLLM / SGLang / TensorRT-LLM)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Your own box&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;When you buy an API, you are buying somebody's &lt;strong&gt;chosen point on the frontier&lt;/strong&gt;. That reframes the shopping question. "Which provider is cheapest per million tokens?" is the wrong question if two providers serve the same open-weights model at wildly different batch sizes — the cheap one may be the one that decided your users can wait.&lt;/p&gt;

&lt;p&gt;The question to ask instead: &lt;em&gt;at my request shape, what tokens per second per user do I get, and at what price?&lt;/em&gt; Our &lt;a href="https://induwara.lk/tools/ai-inference-provider-comparison" rel="noopener noreferrer"&gt;inference provider comparison&lt;/a&gt; is a starting point, but your own workload beats any table.&lt;/p&gt;




&lt;h2&gt;
  
  
  🌐 The latency you cannot tune away from Colombo
&lt;/h2&gt;

&lt;p&gt;Serving-side tuning fights over milliseconds inside a data centre. If your users are in Sri Lanka and your inference runs in us-east, a chunk of your time-to-first-token is pure network round trip, and no kernel rewrite touches it.&lt;/p&gt;

&lt;p&gt;Two practical consequences:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Measure your baseline RTT first.&lt;/strong&gt; Run a &lt;code&gt;ping&lt;/code&gt; or a &lt;code&gt;curl -w '%{time_connect}'&lt;/code&gt; against your provider's endpoint from the machine your app actually runs on. If that number is large relative to your TTFT target, latency optimization on the model side is chasing the smaller half.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Move the app next to the model, not the model next to the user.&lt;/strong&gt; Streaming makes perceived latency mostly about the first token. Putting your Next.js server in the same region as your inference endpoint often beats every sampling parameter you were about to tweak.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;If your app server sits in Singapore and your LLM sits in Virginia, you have built a latency problem that quantization cannot solve.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  ⚡ "Jagged" is the most expensive word in that article
&lt;/h2&gt;

&lt;p&gt;Kiely describes the frontier as &lt;strong&gt;very jagged&lt;/strong&gt;, with cutoff points that are "unintuitive" and get "discovered empirically through sweeps." Quantization is called out as especially jagged: formats like &lt;strong&gt;MXFP4&lt;/strong&gt; and &lt;strong&gt;NVFP4&lt;/strong&gt; can deliver real serving gains with minimal quality loss.&lt;/p&gt;

&lt;p&gt;Jagged means intuition fails. It also means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Benchmarks from someone else's workload transfer badly to yours.&lt;/li&gt;
&lt;li&gt;The configuration that was optimal for &lt;strong&gt;GLM-5.3&lt;/strong&gt; is not automatically optimal for &lt;strong&gt;Kimi K3&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;A sweep of five configurations on one rented GPU-hour is cheaper than a month of guessing wrong.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point is the actionable one for a small team. A sweep is a scripted loop, a night of GPU time, and a CSV. Price it with the &lt;a href="https://induwara.lk/tools/ai-gpu-cloud-cost-calculator" rel="noopener noreferrer"&gt;GPU cloud cost calculator&lt;/a&gt; before assuming you can't afford it.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ Where a two-person team gets the most back
&lt;/h2&gt;

&lt;p&gt;If you are building a product rather than a serving platform, the ranking I'd use:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Batch your own requests.&lt;/strong&gt; This is the one trade you fully control even on a closed API. Anything not user-facing — classification, tagging, summarising a backlog — goes into a batch endpoint or an overnight queue. You are voluntarily selling latency you don't need.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cache aggressively.&lt;/strong&gt; The article assumes KV cache reuse and KV-aware routing as table stakes. On the API side, your equivalent is prompt caching and stable prompt prefixes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check whether speculative decoding is on.&lt;/strong&gt; The piece notes it works especially well for code generation, where token sequences are "relatively predictable" — techniques like &lt;strong&gt;EAGLE-3&lt;/strong&gt;, &lt;strong&gt;DSpark&lt;/strong&gt; and &lt;strong&gt;DFlash&lt;/strong&gt; skip forward passes outright. If you're serving a coding tool, this is the highest-leverage gift on the list. Our &lt;a href="https://induwara.lk/tools/ai-speculative-decoding-calculator" rel="noopener noreferrer"&gt;speculative decoding speedup calculator&lt;/a&gt; will tell you what acceptance rate you'd need before it pays off.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Only then touch parallelism.&lt;/strong&gt; TP, EP and ADP degrees matter when you own the hardware and have traffic to shape. Below that scale, they're a way to lose a weekend.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;The efficient frontier is a filter for your attention. Every inference optimization you read about this year is either a trade you must consciously choose, or a gift you should take. Most engineers get this backwards: they agonise over the gifts, which are free, and accept the trades by accident, usually by leaving a default in place.&lt;/p&gt;

&lt;p&gt;For an SL small team, the sequence is short:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Find out where your latency actually goes before optimizing the model.&lt;/li&gt;
&lt;li&gt;Take every gift your provider offers: caching, batch endpoints, spec decoding.&lt;/li&gt;
&lt;li&gt;Make the trades deliberately, and write down what you sold.&lt;/li&gt;
&lt;li&gt;If you're self-hosting, budget one GPU-hour for a sweep. The frontier is jagged, so the answer is empirical, and it's yours to find.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; You cannot beat a serving company at kernel engineering. You can absolutely beat them at knowing which requests of yours don't need to be fast.&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>llminference</category>
      <category>aiengineering</category>
      <category>costoptimization</category>
    </item>
    <item>
      <title>a16z's $8.5B growth fund: read it as a hardware signal</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Tue, 01 Sep 2026 06:11:13 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/a16zs-85b-growth-fund-read-it-as-a-hardware-signal-4332</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/a16zs-85b-growth-fund-read-it-as-a-hardware-signal-4332</guid>
      <description>&lt;p&gt;The &lt;strong&gt;a16z $8.5 billion growth fund&lt;/strong&gt; news landed on August 31, and the headline number is the least interesting part of it. &lt;a href="https://techcrunch.com/2026/08/31/a16z-brings-growth-fund-to-8-5b-days-after-launching-new-1-1b-fund/" rel="noopener noreferrer"&gt;TechCrunch reported&lt;/a&gt; that Andreessen Horowitz topped up its fifth growth fund by &lt;strong&gt;$1.75 billion&lt;/strong&gt;, three days after announcing a separate &lt;strong&gt;$1.1 billion&lt;/strong&gt; fund aimed at AI hardware.&lt;/p&gt;

&lt;p&gt;I care about the smaller fund. It tells you where the bottleneck is, and the bottleneck is what sets your costs.&lt;/p&gt;




&lt;h2&gt;
  
  
  💰 What actually got announced
&lt;/h2&gt;

&lt;p&gt;Two things, days apart, per TechCrunch:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Fund&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Launched&lt;/th&gt;
&lt;th&gt;Focus&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Growth Fund (fifth)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;$8.5B&lt;/strong&gt; (up from $6.75B)&lt;/td&gt;
&lt;td&gt;January 2026, topped up Aug 2026&lt;/td&gt;
&lt;td&gt;Enterprise + consumer AI, defense tech, robotics, infrastructure hardware and software, health tech&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Machine Age Fund&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$1.1B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;August 28, 2026&lt;/td&gt;
&lt;td&gt;AI hardware: chips, memory, networking, storage&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For scale: the growth fund has backed &lt;strong&gt;100+ companies over seven years&lt;/strong&gt;, and both of these sit on top of the &lt;strong&gt;$15 billion&lt;/strong&gt; a16z announced in January, when it reported &lt;strong&gt;$90 billion&lt;/strong&gt; in assets under management.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;David George&lt;/strong&gt;, the general partner leading the growth investing team, framed it in a blog post as companies "reaching the growth stage faster and gobbling up more cash at higher valuations than ever."&lt;/p&gt;

&lt;p&gt;That sentence is worth reading twice. It is not a boast about returns. It is a description of &lt;em&gt;burn&lt;/em&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔌 The $1.1B is the honest number
&lt;/h2&gt;

&lt;p&gt;A dedicated fund for &lt;strong&gt;chips, memory, networking, and storage&lt;/strong&gt; is an admission about where the constraint sits. Not models. Not apps. The physical layer underneath them.&lt;/p&gt;

&lt;p&gt;Three things follow from that, and none of them require you to believe any hype:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Serving capacity, not model quality, is the scarce thing.&lt;/strong&gt; You don't raise a billion dollars for memory and networking if the hard part is writing better prompts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory bandwidth is the quiet villain.&lt;/strong&gt; Storage and networking showing up alongside chips means the industry is spending on getting weights and KV cache &lt;em&gt;to&lt;/em&gt; the compute, not just on more compute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;This is a multi-year bet.&lt;/strong&gt; Hardware companies don't return capital in eighteen months. Someone thinks the demand curve holds long enough to justify a fund with a fund's lifetime.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; When the smart money moves from applications to memory and networking, it's telling you that inference economics — not model capability — will decide which AI products survive the next two years.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  📊 Why this touches your unit economics
&lt;/h2&gt;

&lt;p&gt;Here's the part that matters for anyone building a small product from Colombo, Kandy, or a bedroom with a decent connection.&lt;/p&gt;

&lt;p&gt;The API price you pay today is set in a market where enormous capital is subsidising capacity buildout. That is genuinely good for you right now. It is also not a law of physics.&lt;/p&gt;

&lt;p&gt;I don't know which way prices move, and I'd distrust anyone who claims they do. What I know is that my margin should not depend on the answer. So I run every AI feature I ship through one question: &lt;strong&gt;what happens at 3× the current token price?&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Product shape&lt;/th&gt;
&lt;th&gt;Cost per user action&lt;/th&gt;
&lt;th&gt;Survives a 3× price move?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One-off tool, short prompt, cached result&lt;/td&gt;
&lt;td&gt;Fractions of a cent&lt;/td&gt;
&lt;td&gt;Yes, comfortably&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chat feature, full history resent each turn&lt;/td&gt;
&lt;td&gt;Grows with every message&lt;/td&gt;
&lt;td&gt;Only with truncation or summarisation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Agent" that loops until confident&lt;/td&gt;
&lt;td&gt;Unbounded by design&lt;/td&gt;
&lt;td&gt;No, unless you cap the loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch job over a document corpus&lt;/td&gt;
&lt;td&gt;Predictable, measurable&lt;/td&gt;
&lt;td&gt;Yes, if you priced it upfront&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The middle two rows are where small teams quietly bleed. An agent loop with no hard iteration cap is a blank cheque written against a price you don't control.&lt;/p&gt;

&lt;p&gt;If you want the actual numbers for your own prompts rather than a feeling, our &lt;a href="https://induwara.lk/tools/ai-token-counter" rel="noopener noreferrer"&gt;AI token counter&lt;/a&gt; will give you the count, and the &lt;a href="https://induwara.lk/tools/ai-model-comparison" rel="noopener noreferrer"&gt;AI model comparison&lt;/a&gt; page is where I check whether a cheaper model would do the same job.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ What I'd change in how you build
&lt;/h2&gt;

&lt;p&gt;Concretely, from my own work on this site:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cap every loop.&lt;/strong&gt; Maximum iterations, maximum tokens, maximum wall clock. Log when you hit the cap instead of silently retrying.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure cost per successful action, not per call.&lt;/strong&gt; A cheap call that fails twice before working is not cheap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cache aggressively at the boundary.&lt;/strong&gt; Identical inputs should never reach a provider twice. This is the highest-leverage change most small products haven't made.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep one working fallback path.&lt;/strong&gt; A smaller model, a cheaper provider, or a plain deterministic version of the feature. Test it occasionally so it isn't fictional when you need it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write the provider layer behind your own interface.&lt;/strong&gt; Swapping models should be a config change, not a rewrite.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is exotic. It's the same discipline you'd apply to a database you don't own.&lt;/p&gt;




&lt;h2&gt;
  
  
  🌐 The career angle nobody mentions
&lt;/h2&gt;

&lt;p&gt;There's a second reading of the Machine Age Fund that's more useful than the funding story.&lt;/p&gt;

&lt;p&gt;A billion dollars aimed at &lt;strong&gt;chips, memory, networking, and storage&lt;/strong&gt; is a statement about which skills get scarce. The industry has spent two years hiring people who can write prompts. The money is now moving toward people who understand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Quantisation and what it actually costs you in accuracy&lt;/li&gt;
&lt;li&gt;Serving architecture, batching, and KV cache behaviour&lt;/li&gt;
&lt;li&gt;Memory bandwidth as a first-class constraint&lt;/li&gt;
&lt;li&gt;Networking between accelerators&lt;/li&gt;
&lt;li&gt;Plain systems engineering: profiling, measuring, not guessing&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;If you're a UCSC or Moratuwa undergrad deciding what to go deep on, that list is a better guide than any job board. Systems fundamentals were never obsolete. They just stopped being fashionable for a while.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The good news for a Sri Lankan engineer is that none of that requires capital. It requires a laptop, patience, and the willingness to read papers and profile things. The barrier is attention, not access.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;You will not raise from this fund. That's fine, and it isn't the point.&lt;/p&gt;

&lt;p&gt;The point is that &lt;strong&gt;$9.6 billion across two funds&lt;/strong&gt; is a very expensive signal about where the constraint sits, and the signal is free to read. Three takeaways:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Build as if inference prices are volatile.&lt;/strong&gt; Cap loops, cache, keep a fallback. Your margin should survive a price move in either direction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't compete on model access.&lt;/strong&gt; Everyone has the same models. Your edge is the thing a16z can't fund: specific knowledge of a specific market, and distribution into it. Sri Lanka is full of problems that no global product will ever bother to solve properly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Go deep on systems, not surfaces.&lt;/strong&gt; The scarce skill is moving toward the layer the money is buying.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The capital is chasing infrastructure because infrastructure is hard. What sits on top of it — a genuinely useful tool for a specific group of people — is still cheap to build and still mostly unbuilt.&lt;/p&gt;

&lt;p&gt;That gap is the opportunity, and it's the one part of this story you can act on tomorrow morning.&lt;/p&gt;

</description>
      <category>venturecapital</category>
      <category>aiinfrastructure</category>
      <category>startup</category>
    </item>
    <item>
      <title>Diffusion Language Models: Why Speed Beats Size Now</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Mon, 31 Aug 2026 10:03:48 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/diffusion-language-models-why-speed-beats-size-now-29nf</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/diffusion-language-models-why-speed-beats-size-now-29nf</guid>
      <description>&lt;p&gt;A &lt;strong&gt;diffusion language model&lt;/strong&gt; writes text like an image model paints: it starts from a mostly-masked draft and refines every position at once, rather than one token at a time. The Kuleshov Group has published &lt;a href="https://kuleshov-group.github.io/blog/blog/2026/how-to-build-a-diffusion-language-model/" rel="noopener noreferrer"&gt;How to build a diffusion language model&lt;/a&gt;, adapted from their ICLR 2026 and MLSS 2026 talks.&lt;/p&gt;

&lt;p&gt;I read it as an infrastructure post. The claim that matters isn't better text, it's &lt;em&gt;cheaper text per second of GPU time&lt;/em&gt;, and that changes who can afford to ship an AI feature.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 The actual break with how LLMs work today
&lt;/h2&gt;

&lt;p&gt;Every model most of us use in production is autoregressive: predict token, append, repeat. The post names three consequences of that design, and they're all structural rather than fixable with more parameters.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Property&lt;/th&gt;
&lt;th&gt;Autoregressive (GPT-style)&lt;/th&gt;
&lt;th&gt;Diffusion&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Generation order&lt;/td&gt;
&lt;td&gt;Strictly left to right&lt;/td&gt;
&lt;td&gt;All positions in parallel&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Attention over output&lt;/td&gt;
&lt;td&gt;Causal (backward only)&lt;/td&gt;
&lt;td&gt;Bidirectional&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Can it revise a token?&lt;/td&gt;
&lt;td&gt;No, once emitted it's final&lt;/td&gt;
&lt;td&gt;Yes, via remasking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost scaling&lt;/td&gt;
&lt;td&gt;Steps = sequence length&lt;/td&gt;
&lt;td&gt;Steps = a knob you set&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is the one I keep coming back to. In an autoregressive model, a 500-token answer costs 500 sequential forward passes. In a diffusion model, the number of denoising steps is a parameter you choose. Quality and latency stop being fixed by output length.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; Diffusion turns generation cost from a function of &lt;em&gt;how long the answer is&lt;/em&gt; into a function of &lt;em&gt;how good you need the answer to be&lt;/em&gt;. That is a budget dial, and small teams have never had one.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  ⚡ Tokens per second is the number that decides your bill
&lt;/h2&gt;

&lt;p&gt;Parameter counts are a vanity metric for anyone self-hosting. Throughput is what you actually pay for. Here's what the post reports, with the caveat that these are the authors' figures and I have not benchmarked any of them:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Scale&lt;/th&gt;
&lt;th&gt;Reported claim&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Mercury 2&lt;/strong&gt; (Inception Labs)&lt;/td&gt;
&lt;td&gt;Commercial&lt;/td&gt;
&lt;td&gt;~&lt;strong&gt;1,200 tokens/second&lt;/strong&gt; on standard GPUs; 5–10× faster than Claude Haiku or Gemini Flash&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Nemotron Diffusion&lt;/strong&gt; (NVIDIA)&lt;/td&gt;
&lt;td&gt;Up to 35B, open weights&lt;/td&gt;
&lt;td&gt;2–8× throughput of comparable autoregressive models, retaining up to 99% of quality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;LLaDA&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;8B, open weights&lt;/td&gt;
&lt;td&gt;Competitive with LLaMA2-7B and LLaMA3-8B on general, math and code benchmarks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Gemma Diffusion&lt;/strong&gt; (Google)&lt;/td&gt;
&lt;td&gt;Open weights&lt;/td&gt;
&lt;td&gt;Uniform-state diffusion with an encoder-decoder design&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Mercury 2 line is the one worth sitting with. The post claims it matches the quality of specialised inference chips like Groq and SambaNova while running on commodity GPUs. If that holds up under independent testing, the moat around fast inference was an architecture choice, not silicon.&lt;/p&gt;

&lt;p&gt;For a two-person team in Colombo renting a GPU by the hour, a 5× throughput gain is the difference between an AI feature that pays for itself and one you quietly turn off at the end of the month.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ The training recipe is less exotic than the name suggests
&lt;/h2&gt;

&lt;p&gt;This is where the post earns its title. The masked diffusion approach (MDLM) is described as, in effect, &lt;strong&gt;BERT with a randomised masking rate&lt;/strong&gt;. You mask a random fraction of tokens, train a bidirectional transformer to reconstruct them, and vary that fraction across training.&lt;/p&gt;

&lt;p&gt;The loss function is the part that should make every engineer sit up:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The training objective reduces to standard cross-entropy, averaged across all masking rates.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No new loss to derive. No exotic sampler to debug before you see a learning curve. If you have written a masked-language-model training loop, you have written most of a diffusion language model.&lt;/p&gt;

&lt;p&gt;The practical pieces the post lays out:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Block diffusion&lt;/strong&gt; — diffuse blocks of tokens (256, for example) conditioned on prior context, which gives you variable-length output &lt;em&gt;and&lt;/em&gt; KV caching, the same optimisation autoregressive serving depends on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Encoder-decoder split&lt;/strong&gt; — a heavy encoder processes the clean context once, a lightweight decoder does the iterative denoising. You pay for context understanding a single time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remasking&lt;/strong&gt; — re-mask a small subset of already-generated tokens so the model can correct itself. This is what makes inference-time scaling possible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uniform-state diffusion (UDLM)&lt;/strong&gt; — replace tokens with random vocabulary items rather than a mask token, which supports multi-step revision more naturally.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RL post-training&lt;/strong&gt; — Diffu-GRPO uses a mean-field approximation for policy gradients, with newer work deriving exact and approximate likelihood estimators.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Points 1 and 2 are what a small team should care about. They're the difference between a research curiosity and something you can put behind an API endpoint.&lt;/p&gt;




&lt;h2&gt;
  
  
  💰 What changes for a builder on a learning budget
&lt;/h2&gt;

&lt;p&gt;I run tools that call models constantly, and the cost model here is genuinely different. Three things I'd plan around:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Latency stops scaling with answer length.&lt;/strong&gt; Long-form generation, the thing that currently makes users stare at a spinner, is where diffusion gains the most.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality becomes a runtime decision.&lt;/strong&gt; You can serve a cheap 8-step sample to free users and a 64-step sample to paying ones, from the same weights. No second model, no second deployment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tuning a small diffusion LM is plausible on rented hardware.&lt;/strong&gt; Cross-entropy on a bidirectional transformer is not an exotic training run. A domain-specific model for Sinhala or Tamil text is not obviously out of reach for a university lab.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're modelling what any of this costs before you commit, our &lt;a href="https://induwara.lk/tools/ai-token-counter" rel="noopener noreferrer"&gt;AI token counter&lt;/a&gt; gives you the input side, and the &lt;a href="https://induwara.lk/tools/ai-speculative-decoding-calculator" rel="noopener noreferrer"&gt;speculative decoding calculator&lt;/a&gt; covers the closest thing the autoregressive world has to parallel generation. Both are free and need no signup.&lt;/p&gt;




&lt;h2&gt;
  
  
  🧪 Where I'd stay sceptical
&lt;/h2&gt;

&lt;p&gt;The post is an advocacy piece by researchers who work on diffusion. A few numbers deserve a second look before you rewrite your stack.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Claim&lt;/th&gt;
&lt;th&gt;The catch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"Up to 99% quality retained" (Nemotron)&lt;/td&gt;
&lt;td&gt;"Up to" is doing work. A 1% quality drop is not uniform across tasks.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Remasking improves MAUVE 0.40 → 0.66&lt;/td&gt;
&lt;td&gt;Autoregressive still scores &lt;strong&gt;0.76&lt;/strong&gt; in that comparison. The gap narrowed, it did not close.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLaDA competitive with LLaMA2-7B / LLaMA3-8B&lt;/td&gt;
&lt;td&gt;Those are the baselines being matched. Matching a 2023–2024 open model is a real result, not frontier parity.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The strongest evidence in the post is arguably from biology, not chat: &lt;strong&gt;NT-v3&lt;/strong&gt;, trained on over a trillion DNA tokens, generated regulatory sequences that outperformed native enhancers in wet-lab validation, and &lt;strong&gt;ESM3&lt;/strong&gt; reached 100B parameters for protein generation. Those are domains where bidirectional context and iterative revision obviously beat left-to-right, and where the win is measured in a lab rather than on a leaderboard.&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;Don't rewrite anything this week. Do three things instead:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Track throughput as a first-class metric&lt;/strong&gt; in whatever you're building. If diffusion serving matures, tokens/second is where the savings land, and you'll want a baseline to compare against.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pull LLaDA or Nemotron Diffusion weights and run one inference.&lt;/strong&gt; They're open. An afternoon of hands-on time now beats reading about it for a year.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you're a student picking a research topic&lt;/strong&gt;, this area is unusually accessible. The objective is cross-entropy, the architecture is a transformer you already understand, and the field is young enough that a careful paper from a Sri Lankan university would land.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The post frames diffusion as potentially being to inference-time scaling what the transformer was to pre-training scaling. That's a large claim and it's unproven. But the underlying shift is real and it's already shipping: the cost of a generated token is becoming something you tune rather than something you're handed.&lt;/p&gt;

&lt;p&gt;For those of us earning in rupees and paying for GPUs in dollars, a knob is worth a lot.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>llminfrastructure</category>
    </item>
    <item>
      <title>vLLM v0.28.0: the breaking change small GPU users must read</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:00:41 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/vllm-v0280-the-breaking-change-small-gpu-users-must-read-pkh</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/vllm-v0280-the-breaking-change-small-gpu-users-must-read-pkh</guid>
      <description>&lt;p&gt;&lt;strong&gt;vLLM v0.28.0&lt;/strong&gt; shipped on 26 August 2026 with &lt;strong&gt;584 commits from 270 contributors (76 new)&lt;/strong&gt;, and almost every headline in the &lt;a href="https://github.com/vllm-project/vllm/releases/tag/v0.28.0" rel="noopener noreferrer"&gt;release notes on GitHub&lt;/a&gt; is aimed at people running sixteen GPUs at once.&lt;/p&gt;

&lt;p&gt;I read it from the opposite end: what happens to someone serving one model on one rented card. The answer is that two lines buried under "Breaking Changes" matter far more to that person than the entire Kimi-K3 section.&lt;/p&gt;




&lt;h2&gt;
  
  
  🧨 bitsandbytes is no longer in the box
&lt;/h2&gt;

&lt;p&gt;The single most consequential line for a small-GPU setup: &lt;strong&gt;bitsandbytes support migrated to an out-of-tree plugin&lt;/strong&gt; (#43529). bitsandbytes is how a lot of people load 4-bit and 8-bit weights onto a consumer card without pre-quantising anything. If your launch command passes &lt;code&gt;--quantization bitsandbytes&lt;/code&gt;, a blind &lt;code&gt;pip install -U vllm&lt;/code&gt; will not do what it did last week.&lt;/p&gt;

&lt;p&gt;Here is the full breaking list, with who each one actually hits:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;th&gt;PR&lt;/th&gt;
&lt;th&gt;Who feels it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;bitsandbytes moved out-of-tree&lt;/td&gt;
&lt;td&gt;#43529&lt;/td&gt;
&lt;td&gt;Anyone loading BnB 4-bit/8-bit weights on a small GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transformers bumped to &lt;strong&gt;5.15.0&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;#51668&lt;/td&gt;
&lt;td&gt;Anyone pinning Transformers for another library in the same venv&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;calculate_kv_scales&lt;/code&gt; removed&lt;/td&gt;
&lt;td&gt;#49389&lt;/td&gt;
&lt;td&gt;FP8 KV cache users doing runtime scale calculation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;override_attention_dtype&lt;/code&gt; removed&lt;/td&gt;
&lt;td&gt;#48684&lt;/td&gt;
&lt;td&gt;People forcing an attention dtype to dodge a numerics bug&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;reasoning_content&lt;/code&gt; output removed&lt;/td&gt;
&lt;td&gt;#50624&lt;/td&gt;
&lt;td&gt;Clients parsing reasoning traces from the response&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV tiering metrics renamed &lt;code&gt;block&lt;/code&gt; → &lt;code&gt;chunk&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;#52812&lt;/td&gt;
&lt;td&gt;Anyone with a Grafana dashboard on offload metrics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MoE legacy code removed&lt;/td&gt;
&lt;td&gt;#51078&lt;/td&gt;
&lt;td&gt;Custom MoE kernels built against the old structure&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; on a single-GPU box, v0.28.0 is not a "pull and restart" upgrade. Pin your current version, read those seven lines, and only then move.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  📊 A doubled default that costs you memory
&lt;/h2&gt;

&lt;p&gt;Under "New defaults", &lt;code&gt;max_num_batched_tokens&lt;/code&gt; was &lt;strong&gt;raised from 8192 to 16384&lt;/strong&gt; (#51726). Prefix caching is now &lt;strong&gt;on by default for Mamba models&lt;/strong&gt; (#50991), and the Blackwell CUDA graph capture default went up to &lt;strong&gt;1024&lt;/strong&gt; (#49390).&lt;/p&gt;

&lt;p&gt;These are throughput wins on big hardware. On a 16GB or 24GB card they are the opposite: a larger batched-token budget means more activation memory reserved before your KV cache gets a look-in. If you upgrade and hit an out-of-memory error you did not hit before, that default is the first thing I would check.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Restore the old behaviour while you measure&lt;/span&gt;
vllm serve &amp;lt;model&amp;gt; &lt;span class="nt"&gt;--max-num-batched-tokens&lt;/span&gt; 8192
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you are sizing a card before you rent it, our &lt;a href="https://induwara.lk/tools/ai-llm-vram-calculator" rel="noopener noreferrer"&gt;LLM GPU Memory (VRAM) Calculator&lt;/a&gt; will tell you how much headroom you actually have to give away.&lt;/p&gt;




&lt;h2&gt;
  
  
  🖥️ The CPU and low-end story quietly got better
&lt;/h2&gt;

&lt;p&gt;This is the part of the release I did not expect, and it is the part that matters most if your budget is a VPS rather than an H100 hour:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An &lt;strong&gt;MLA backend for CPU&lt;/strong&gt;, so DeepSeek-V2/V3 can run there at all (#49453)&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;triton-cpu wheel&lt;/strong&gt; (#52092) and &lt;strong&gt;tcmalloc&lt;/strong&gt; in the CPU path (#50841)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPTQ and AWQ enabled on s390x&lt;/strong&gt; (#51148)&lt;/li&gt;
&lt;li&gt;An &lt;strong&gt;unquantized MoE backend for Power (VSX)&lt;/strong&gt; (#51624)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disk offloading&lt;/strong&gt; for the CPU offload connector (#49644), plus &lt;strong&gt;weight offloading&lt;/strong&gt; in Model Runner V2 (#51413)&lt;/li&gt;
&lt;li&gt;An &lt;strong&gt;XPU wheel added to the release pipeline&lt;/strong&gt; (#52108) for Intel cards&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The CPU image is published like any other: &lt;code&gt;docker pull vllm/vllm-openai-cpu:v0.28.0&lt;/code&gt;. None of this will be fast. But "slow and running on hardware I already pay for" beats "fast on a card I cannot afford" when you are learning, prototyping, or serving ten requests a day to a class project.&lt;/p&gt;




&lt;h2&gt;
  
  
  ⚡ Speculative decoding is where the real speed went
&lt;/h2&gt;

&lt;p&gt;Three of the release's headline items are speculative decoding: &lt;strong&gt;DFlash2&lt;/strong&gt; with local convolution and a candidate selector (#52816), &lt;strong&gt;DSpark&lt;/strong&gt; confidence-scheduled verification (#47808), and &lt;strong&gt;async scheduling auto-enabled for draft models&lt;/strong&gt; (#48341).&lt;/p&gt;

&lt;p&gt;The eye-catching numbers in the notes belong to frontier setups, and I want to be precise about that rather than let them read as general claims:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Claimed number&lt;/th&gt;
&lt;th&gt;What it actually applies to&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;~60% better TTFT&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Adaptive speculative token budget, DSpark, &lt;strong&gt;on Kimi-K3&lt;/strong&gt; (#51725)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;~17 GiB saved per GPU&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Optional shared-expert sharding, &lt;strong&gt;Kimi-K3 only&lt;/strong&gt; (#50912)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1.5~3x kernel-level speedup&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Combined all-gathers, kernel level, &lt;strong&gt;not end-to-end&lt;/strong&gt; (#51070)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;None of those figures transfer to a 7B model on one card. Kernel-level speedup is not request-level speedup, and a Kimi-K3 memory saving means nothing if you were never going to load Kimi-K3.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Speculative decoding itself does transfer, and it is one of the few free speedups available on modest hardware. Before you wire up a draft model, run the numbers on our &lt;a href="https://induwara.lk/tools/ai-speculative-decoding-calculator" rel="noopener noreferrer"&gt;Speculative Decoding Speedup Calculator&lt;/a&gt; — the payoff depends almost entirely on the draft acceptance rate, and a bad draft model makes things slower, not faster.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔐 The security note everyone exposing a port should read
&lt;/h2&gt;

&lt;p&gt;Buried in the Security section is a documentation change worth more than most of the features:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;vLLM's docs now warn that &lt;strong&gt;&lt;code&gt;--api-key&lt;/code&gt; does not gate all endpoints&lt;/strong&gt; (#51999).&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you are running vLLM on a cheap cloud box with a public IP and assuming that flag is your authentication layer, it is not. Put it behind a reverse proxy, or bind it to localhost and tunnel in. The same section also fixed a &lt;strong&gt;denial-of-service via sample-rate forgery&lt;/strong&gt; that bypassed the audio decode duration guard (#49948), which is exactly the class of bug that bites a publicly reachable multimodal endpoint.&lt;/p&gt;

&lt;p&gt;Two smaller hardening changes are the kind that break scripts quietly: &lt;code&gt;cache_salt&lt;/code&gt; must now be non-empty (#50816), and non-object JSON bodies return &lt;strong&gt;400 instead of 500&lt;/strong&gt; (#51654, #52528).&lt;/p&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;If you are a student, freelancer, or small team in Sri Lanka running open models on rented or borrowed hardware, my read is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Do not upgrade in place.&lt;/strong&gt; Check the seven breaking changes against your launch command first. bitsandbytes is the likely one to catch you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you upgrade, pass &lt;code&gt;--max-num-batched-tokens 8192&lt;/code&gt; on the first run&lt;/strong&gt; so you are comparing like with like, then raise it deliberately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ignore the Kimi-K3 and DeepSeek-V4 headlines.&lt;/strong&gt; They are genuinely impressive engineering for clusters, and irrelevant to a single card.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Take the CPU and offload work seriously.&lt;/strong&gt; Disk offloading and a CPU MLA backend widen what runs on hardware you can actually rent for a few dollars a month.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix your auth before your throughput.&lt;/strong&gt; A public endpoint with no real gate is a bigger problem than 20% fewer tokens per second.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The wider pattern is worth naming. vLLM's release notes now read like infrastructure documentation for labs, but the project keeps shipping the unglamorous portability work in the same release. That second half is the part that decides whether people outside a handful of well-funded labs can serve open models at all. Before you commit to renting anything, price the options with our &lt;a href="https://induwara.lk/tools/ai-gpu-cloud-cost-calculator" rel="noopener noreferrer"&gt;GPU Cloud Cost Calculator&lt;/a&gt; and check your expected throughput with the &lt;a href="https://induwara.lk/tools/ai-inference-speed-calculator" rel="noopener noreferrer"&gt;Inference Speed Calculator&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>vllm</category>
      <category>llminference</category>
      <category>opensource</category>
    </item>
    <item>
      <title>TurboKV: what a Rust key-value store's benchmarks really say</title>
      <dc:creator>Induwara Ashinsana</dc:creator>
      <pubDate>Sat, 29 Aug 2026 17:47:12 +0000</pubDate>
      <link>https://dev.to/induwara_ashinsana_9e4d5b/turbokv-what-a-rust-key-value-stores-benchmarks-really-say-5df4</link>
      <guid>https://dev.to/induwara_ashinsana_9e4d5b/turbokv-what-a-rust-key-value-stores-benchmarks-really-say-5df4</guid>
      <description>&lt;p&gt;&lt;strong&gt;TurboKV&lt;/strong&gt;, a new embedded &lt;strong&gt;Rust key-value store&lt;/strong&gt;, showed up on Hacker News this week with the kind of headline that usually makes me close the tab: "insanely fast." I opened the &lt;a href="https://github.com/kingroryg/turbokv" rel="noopener noreferrer"&gt;repository&lt;/a&gt; expecting a benchmark chart with no axis labels. I found the opposite.&lt;/p&gt;

&lt;p&gt;The speed claims are real enough. But the genuinely instructive part of this project is the paragraph &lt;em&gt;underneath&lt;/em&gt; the numbers, and that is what I want to talk about.&lt;/p&gt;




&lt;h2&gt;
  
  
  🔍 What TurboKV actually is
&lt;/h2&gt;

&lt;p&gt;Strip the adjectives and it is a small, specific thing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Embedded&lt;/strong&gt;, not networked. It runs inside your process as a library. There is no server, no port, no connection pool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Async&lt;/strong&gt;, built on &lt;strong&gt;Tokio&lt;/strong&gt;. &lt;code&gt;insert&lt;/code&gt;, &lt;code&gt;get&lt;/code&gt; and &lt;code&gt;remove&lt;/code&gt; are all &lt;code&gt;await&lt;/code&gt;ed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LSM-tree shaped&lt;/strong&gt;: memtables, SSTables, a write-ahead log, background compaction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apache 2.0&lt;/strong&gt;, requires &lt;strong&gt;Rust 1.85+&lt;/strong&gt;, currently at version &lt;strong&gt;0.6.0&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The API is about as small as this class of database gets:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;Db&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;open_with_options&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"./my-database"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nn"&gt;DbOptions&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;durable&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;&lt;span class="k"&gt;.await&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="nf"&gt;.insert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;b"user:1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;b"Ada"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="k"&gt;.await&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="nd"&gt;assert_eq!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="nf"&gt;.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;b"user:1"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="k"&gt;.await&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;b"Ada"&lt;/span&gt;&lt;span class="nf"&gt;.to_vec&lt;/span&gt;&lt;span class="p"&gt;()));&lt;/span&gt;
&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="nf"&gt;.close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="k"&gt;.await&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three durability presets ship in the box, and the README is blunt about what each one actually promises:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Preset&lt;/th&gt;
&lt;th&gt;When your write is acknowledged&lt;/th&gt;
&lt;th&gt;Honest reading&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fast()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;In-memory visibility, no WAL&lt;/td&gt;
&lt;td&gt;A cache. Crash means data loss.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;durable()&lt;/code&gt; (default)&lt;/td&gt;
&lt;td&gt;Appended to the WAL, no per-write sync&lt;/td&gt;
&lt;td&gt;Survives a process crash, not a power cut&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;paranoid()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;WAL group finished &lt;code&gt;sync_all&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Strongest, still bounded by your disk's honesty&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  📊 Read the benchmark table, not the headline
&lt;/h2&gt;

&lt;p&gt;Here is the published comparison against &lt;strong&gt;fjall 2.11.2&lt;/strong&gt; and &lt;strong&gt;redb 2.6.3&lt;/strong&gt;, measured on an Apple M4 with 32 GiB of RAM on 2026-08-28. Throughput is acknowledged keys per second, higher is better.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;TurboKV Recoverable&lt;/th&gt;
&lt;th&gt;fjall Buffer&lt;/th&gt;
&lt;th&gt;redb Eventual&lt;/th&gt;
&lt;th&gt;TurboKV / fjall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sequential fill (1 key/txn)&lt;/td&gt;
&lt;td&gt;1,407,678&lt;/td&gt;
&lt;td&gt;485,252&lt;/td&gt;
&lt;td&gt;1,397&lt;/td&gt;
&lt;td&gt;2.901×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Random fill (1 key/txn)&lt;/td&gt;
&lt;td&gt;834,137&lt;/td&gt;
&lt;td&gt;456,924&lt;/td&gt;
&lt;td&gt;1,549&lt;/td&gt;
&lt;td&gt;1.826×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Overwrite (1 key/txn)&lt;/td&gt;
&lt;td&gt;853,083&lt;/td&gt;
&lt;td&gt;446,733&lt;/td&gt;
&lt;td&gt;1,516&lt;/td&gt;
&lt;td&gt;1.910×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch (100 keys/txn)&lt;/td&gt;
&lt;td&gt;2,272,259&lt;/td&gt;
&lt;td&gt;511,600&lt;/td&gt;
&lt;td&gt;80,197&lt;/td&gt;
&lt;td&gt;4.441×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch (1,000 keys/txn)&lt;/td&gt;
&lt;td&gt;2,333,582&lt;/td&gt;
&lt;td&gt;572,671&lt;/td&gt;
&lt;td&gt;134,636&lt;/td&gt;
&lt;td&gt;4.075×&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Now look at what the author does with that redb column. Instead of banking the thousand-fold win, the README states that redb's &lt;code&gt;Durability::Eventual&lt;/code&gt; performs a macOS &lt;code&gt;F_BARRIERFSYNC&lt;/code&gt; on every transaction, while TurboKV and fjall stop at their process-crash-recoverable OS-cache boundaries. It then calls its own single-key rows "architectural context rather than a like-for-like durability claim."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; a benchmark that tells you which of its own rows are unfair is worth more than a benchmark that is 4× faster. The second one you have to re-run yourself; the first one you can reason about.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The project also renames its own default. The preset is called &lt;code&gt;durable()&lt;/code&gt; in code, but in the benchmark table it appears as &lt;strong&gt;"Recoverable"&lt;/strong&gt;, because it survives a crashed process and not a lost power supply. I have reviewed vendor benchmarks that would never make that distinction in public.&lt;/p&gt;




&lt;h2&gt;
  
  
  💰 Why "embedded" is the interesting word for a small team
&lt;/h2&gt;

&lt;p&gt;Most of us here are not sizing a fleet. We are running one box, and the monthly bill is in dollars against an LKR income. That changes which database property matters.&lt;/p&gt;

&lt;p&gt;A networked store like Redis or Postgres costs you a second process, its own memory floor, a socket, a supervisor entry and a backup story. An embedded store costs you a directory.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concern&lt;/th&gt;
&lt;th&gt;Networked store&lt;/th&gt;
&lt;th&gt;Embedded store like TurboKV&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Extra process to run&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Baseline RAM&lt;/td&gt;
&lt;td&gt;Its own allocation&lt;/td&gt;
&lt;td&gt;Yours, inside your app&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Network round trip per read&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shared across services&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No, single owner&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ops surface&lt;/td&gt;
&lt;td&gt;Config, auth, ports&lt;/td&gt;
&lt;td&gt;A folder on disk&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is the trade, and it is a real one. TurboKV states plainly that &lt;strong&gt;one open &lt;code&gt;Db&lt;/code&gt; handle exclusively owns its data directory&lt;/strong&gt;. If two of your services need the same data, this is the wrong tool and no benchmark number changes that.&lt;/p&gt;

&lt;p&gt;Watch the memory defaults if you are on a 1 GB or 2 GB VPS. Every preset starts with a &lt;strong&gt;64 MiB memtable&lt;/strong&gt; and a &lt;strong&gt;64 MiB block cache&lt;/strong&gt;, so roughly 128 MiB is spoken for before your application allocates anything. Both are public fields you can lower before opening. If your keys are actually embeddings rather than short strings, the arithmetic gets steeper fast, and our &lt;a href="https://induwara.lk/tools/ai-vector-storage-calculator" rel="noopener noreferrer"&gt;AI vector storage calculator&lt;/a&gt; will size that for you before you provision the box.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ Four things I would check before shipping it
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;It is 0.6.0.&lt;/strong&gt; That is a pre-1.0 version number on a storage engine. Storage bugs are the expensive kind, because they are discovered later than they happen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The build flags are not optional.&lt;/strong&gt; The persisted Bloom-filter format uses hardware AES. You are told to build with &lt;code&gt;RUSTFLAGS="-C target-feature=+aes,+sse2"&lt;/code&gt; on x86_64 and &lt;code&gt;+aes,+neon&lt;/code&gt; on ARM. Miss this in your CI image and you find out at runtime.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dropping the handle is not a clean shutdown.&lt;/strong&gt; Call &lt;code&gt;close()&lt;/code&gt; or &lt;code&gt;close_with_status()&lt;/code&gt;. This has to survive your panic paths and your container's stop signal, not just the happy path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The benchmark hardware is not your hardware.&lt;/strong&gt; Those numbers come from an Apple M4 with 32 GiB of RAM on APFS. Your $6 VPS has shared vCPUs and network-backed storage. Expect a different shape, not just a smaller multiple.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; treat the published throughput as evidence that the design is sound, not as a number you will reproduce.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  💡 What this means for you
&lt;/h2&gt;

&lt;p&gt;If you are building a Rust service that needs local ordered storage, a queue, a cache with range scans, or an index that must survive a restart, TurboKV is worth an afternoon. The API surface is small enough to learn in one sitting and the scan and batch semantics are documented in unusual detail.&lt;/p&gt;

&lt;p&gt;If you are not writing Rust, take the other thing instead. The habit worth stealing from this repo is the discipline of publishing what your numbers do &lt;em&gt;not&lt;/em&gt; prove: the hardware, the durability boundary, the row that is not a fair fight, and the raw artifact so someone can check you. That applies to your own benchmarks, your client-facing performance claims, and your final year project.&lt;/p&gt;

&lt;p&gt;Fast is a claim anyone can make. Legible is the harder one, and it is the one that survives review.&lt;/p&gt;

</description>
      <category>rust</category>
      <category>database</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
