<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ritabrata Maiti</title>
    <description>The latest articles on DEV Community by Ritabrata Maiti (@ritabratamaiti).</description>
    <link>https://dev.to/ritabratamaiti</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2453619%2F2ee6b749-3337-496d-8d7c-62521479edc3.jpeg</url>
      <title>DEV Community: Ritabrata Maiti</title>
      <link>https://dev.to/ritabratamaiti</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ritabratamaiti"/>
    <language>en</language>
    <item>
      <title>Every $20 AI subscription costs about $100 to serve. The bill is coming.</title>
      <dc:creator>Ritabrata Maiti</dc:creator>
      <pubDate>Thu, 24 Sep 2026 08:33:39 +0000</pubDate>
      <link>https://dev.to/ritabratamaiti/every-20-ai-subscription-costs-about-100-to-serve-the-bill-is-coming-7bb</link>
      <guid>https://dev.to/ritabratamaiti/every-20-ai-subscription-costs-about-100-to-serve-the-bill-is-coming-7bb</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/WPHYPeCjNv8" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;A piece called&lt;br&gt;
&lt;a href="https://www.thestateofbrand.com/news/ai-subscription-time-bomb" rel="noopener noreferrer"&gt;AI subscriptions are a ticking time bomb for enterprise&lt;/a&gt;&lt;br&gt;
made the rounds this week, and the headline is the right shape of the&lt;br&gt;
story. Every major AI lab is running an industry-wide loss-leader at a&lt;br&gt;
scale that does not really have a precedent. Your company's $20 Claude&lt;br&gt;
Pro seats and $20 ChatGPT Plus seats are being served at something&lt;br&gt;
like five times the cost the lab is collecting for them, and that&lt;br&gt;
arrangement is not stable.&lt;/p&gt;

&lt;p&gt;The price tag has not moved in three years. The product has changed&lt;br&gt;
completely.&lt;/p&gt;

&lt;h2&gt;
  
  
  The unit economics
&lt;/h2&gt;

&lt;p&gt;A Claude Pro seat is $20 a month and gives you Sonnet 4.6, Opus 4.6,&lt;br&gt;
file creation, code execution, web search. On the API, those same&lt;br&gt;
models cost $3 input and $15 output per million tokens for Sonnet,&lt;br&gt;
$5 input and $25 output for Opus. A knowledge worker running Claude&lt;br&gt;
for a few hours a day, uploading documents, drafting reports, easily&lt;br&gt;
moves through enough tokens that the API-priced equivalent of that&lt;br&gt;
seat sits somewhere between $200 and $400 a month.&lt;/p&gt;

&lt;p&gt;Microsoft was reportedly losing more than $20 a month on every GitHub&lt;br&gt;
Copilot seat. Power users were hitting $80. One widely-cited analysis&lt;br&gt;
found Anthropic was burning about $8 of compute for every $1 of&lt;br&gt;
subscription revenue. The $20 sticker has been frozen since 2022 and&lt;br&gt;
the models in that window picked up image generation, code execution,&lt;br&gt;
voice, agentic reasoning, web search, and a generational capability&lt;br&gt;
jump. The number stayed put.&lt;/p&gt;

&lt;p&gt;That is the whole story. Everything from here is mechanism.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the subsidy was buying
&lt;/h2&gt;

&lt;p&gt;Cheap inference at consumer prices, deployed broadly to enterprises,&lt;br&gt;
was buying integration depth. The labs were not trying to make money&lt;br&gt;
on the seat. They were trying to make sure the seat existed in the&lt;br&gt;
first place, then make sure the company's marketing draft and the&lt;br&gt;
engineer's pull request review and the analyst's quarterly summary&lt;br&gt;
all happened through it. Once those workflows are load-bearing, the&lt;br&gt;
price can move. The dependency is the asset.&lt;/p&gt;

&lt;p&gt;You can see this in the language coming out of OpenAI. Nick Turley,&lt;br&gt;
their VP of product, described the subscription pricing as something&lt;br&gt;
they&lt;br&gt;
&lt;a href="https://danielmiessler.com/blog/ai-stops-being-artificially-cheap" rel="noopener noreferrer"&gt;stumbled into&lt;/a&gt;,&lt;br&gt;
and has floated phasing out unlimited plans entirely, comparing them&lt;br&gt;
to "unlimited electricity." Sam Altman said publicly that OpenAI now&lt;br&gt;
needs to become "an AI inference company," which is the polite version&lt;br&gt;
of admitting the consumer subscription was a customer-acquisition line&lt;br&gt;
item, not a P&amp;amp;L.&lt;/p&gt;

&lt;p&gt;The KPMG Q1 2026 pulse has U.S. organizations projecting average AI&lt;br&gt;
spending of $207 million over the next twelve months, roughly double&lt;br&gt;
the year before. A Goldman Sachs survey of large companies has most&lt;br&gt;
of them overrunning their AI budgets by orders of magnitude.&lt;br&gt;
Chandrasekaran, who runs AI and data at KPMG North America, told&lt;br&gt;
&lt;a href="https://www.marketplace.org/story/2026/04/24/how-much-is-too-much-to-spend-on-ai-tools" rel="noopener noreferrer"&gt;Marketplace&lt;/a&gt;&lt;br&gt;
the quiet part: "Even a quarter or two ago nobody bothered about LLM&lt;br&gt;
consumption costs." It is the bother stage now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agents are what broke the math
&lt;/h2&gt;

&lt;p&gt;The reason the subsidy held as long as it did is that AI was a chatbot.&lt;br&gt;
You typed, it answered, you read the answer, repeat. A normal session&lt;br&gt;
was a few thousand tokens. Heavy use ran into the tens of thousands.&lt;br&gt;
At those volumes, $20 a seat was uncomfortable for the lab but not&lt;br&gt;
catastrophic.&lt;/p&gt;

&lt;p&gt;Agents do not look like that.&lt;/p&gt;

&lt;p&gt;A Claude Code session runs autonomously for an extended period. It&lt;br&gt;
reads files, writes files, runs commands, looks at the output, decides&lt;br&gt;
what to do next, repeats. Users have been&lt;br&gt;
&lt;a href="https://www.uncoveralpha.com/p/the-era-of-subsidized-ai-model-usage" rel="noopener noreferrer"&gt;exhausting five-hour rate-limit windows in under ninety minutes&lt;/a&gt;.&lt;br&gt;
Multiple agents in parallel on a single project multiply that. A&lt;br&gt;
developer running three or four concurrent coding agents is consuming&lt;br&gt;
something close to an order of magnitude more tokens than the same&lt;br&gt;
person in chat, and the subscription price on the seat is unchanged.&lt;/p&gt;

&lt;p&gt;GitHub took the obvious next step. On June 1, 2026, Copilot&lt;br&gt;
&lt;a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/" rel="noopener noreferrer"&gt;moves to usage-based billing&lt;/a&gt;,&lt;br&gt;
specifically because the flat-fee model collapsed under agentic&lt;br&gt;
workloads. The announcement spelled it out: agentic usage is becoming&lt;br&gt;
the default, the inference demand is qualitatively different, the&lt;br&gt;
pricing has to follow.&lt;/p&gt;

&lt;p&gt;Everyone else will do the same thing on a delay.&lt;/p&gt;

&lt;h2&gt;
  
  
  The enterprise position
&lt;/h2&gt;

&lt;p&gt;This is where it gets ugly for organizations that have not done the&lt;br&gt;
work.&lt;/p&gt;

&lt;p&gt;Over the past two years, thousands of companies have woven $20 AI&lt;br&gt;
subscriptions deep into operations. Marketing drafts copy through&lt;br&gt;
ChatGPT Plus. Engineering writes and reviews code through Claude Pro.&lt;br&gt;
Research synthesizes documents. Customer success summarizes tickets.&lt;br&gt;
Finance models scenarios. The line items are budgeted at subsidized&lt;br&gt;
prices because that is what the bill currently says. The actual cost&lt;br&gt;
of the same workloads at API rates, if the lab were charging it, is&lt;br&gt;
fifteen to twenty times higher.&lt;/p&gt;

&lt;p&gt;When prices adjust, two things happen at once. The bill goes up, and&lt;br&gt;
the workflows are already too embedded to rip out. The subsidy creates&lt;br&gt;
the dependency, the dependency makes the price increase unavoidable.&lt;br&gt;
That is the trap, and there is no clever way out of it.&lt;/p&gt;

&lt;p&gt;The companies that survive this transition cleanly will be the ones&lt;br&gt;
that did the bookkeeping. Track per-team token consumption. Know which&lt;br&gt;
workflows are genuinely high-value and which are running Claude on&lt;br&gt;
something a 2018 script could have done. Have a sense of which&lt;br&gt;
subscriptions can move to per-seat billing if it has to, and which&lt;br&gt;
ones become structurally expensive overnight.&lt;/p&gt;

&lt;p&gt;The companies that don't survive cleanly will discover the bill in&lt;br&gt;
the third week of whichever month the subsidy ends, with no time to&lt;br&gt;
re-budget and no leverage to renegotiate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it means for agent tools
&lt;/h2&gt;

&lt;p&gt;The shake-out hits the agent tool market harder than it hits the labs.&lt;br&gt;
The labs are losing on the seat, but they own the inference. They can&lt;br&gt;
move the price. They can change the plan. They can introduce&lt;br&gt;
usage-based tiers and call them "for power users." The seat is still&lt;br&gt;
there at the end.&lt;/p&gt;

&lt;p&gt;The agent tool, the wrapper, the IDE plugin, the browser extension&lt;br&gt;
that uses the lab's API or subscription on your behalf, is in a worse&lt;br&gt;
position. If it has its own per-seat subscription, it is selling you&lt;br&gt;
something whose underlying cost just doubled, and it has to either eat&lt;br&gt;
that or pass it on. If it bills on its own meter on top of the lab's,&lt;br&gt;
it is asking enterprises to swallow a usage line item that already had&lt;br&gt;
no budget code.&lt;/p&gt;

&lt;p&gt;The agent tools that survive the next twelve months will be the ones&lt;br&gt;
that don't sell their own meter. Tools that piggyback on the&lt;br&gt;
subscription the user already has. Tools that don't introduce a second&lt;br&gt;
billing surface for finance to police. Tools whose cost curve is&lt;br&gt;
shaped by the lab's pricing, not by the tool's overhead.&lt;/p&gt;

&lt;p&gt;This is not a clever positioning argument. It is what happens to every&lt;br&gt;
software-on-top-of-software market when the underlying utility starts&lt;br&gt;
charging real prices. The tools that own the price stack survive. The&lt;br&gt;
tools that resell the utility at a markup get squeezed.&lt;/p&gt;

&lt;h2&gt;
  
  
  A note on what I work on
&lt;/h2&gt;

&lt;p&gt;I build &lt;a href="https://browy.dev" rel="noopener noreferrer"&gt;Browy&lt;/a&gt;, an open-source AI agent that lives&lt;br&gt;
in a Chrome side panel and a DevTools REPL. It drives the real browser&lt;br&gt;
tabs you have open. The thing it does not have is its own subscription.&lt;br&gt;
It uses your existing&lt;br&gt;
&lt;a href="https://github.com/features/copilot" rel="noopener noreferrer"&gt;GitHub Copilot&lt;/a&gt; subscription for&lt;br&gt;
the model. The model call goes from your machine to GitHub, the answer&lt;br&gt;
comes back, the rest happens locally.&lt;/p&gt;

&lt;p&gt;When Copilot moves to usage-based billing on June 1, you pay GitHub&lt;br&gt;
the new rate, the same as you would have anyway. Browy doesn't sit&lt;br&gt;
between you and that bill. It doesn't add a per-seat charge of its&lt;br&gt;
own. It doesn't run a metered tier on top. That is not a special&lt;br&gt;
business decision on our end, it is the only decision that survives&lt;br&gt;
the shake-out described above. The tools that try to live on top of a&lt;br&gt;
collapsing subsidy by adding their own subsidy get squeezed twice.&lt;/p&gt;

&lt;p&gt;That is most of the story I wanted to put down. The original piece is&lt;br&gt;
&lt;a href="https://www.thestateofbrand.com/news/ai-subscription-time-bomb" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;br&gt;
The 30-second video version is at the top of this post.&lt;/p&gt;




&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://browyhq.github.io/blog/" rel="noopener noreferrer"&gt;All posts&lt;/a&gt; Index of every post on the Browy blog.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://browyhq.github.io/blog/ai-slop-killed-the-bug-bounty/" rel="noopener noreferrer"&gt;AI slop killed the open-source bug bounty&lt;/a&gt; Earlier essay on the other end of the same arc: cheap inference, not yet expensive.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aisubscriptions</category>
      <category>githubcopilot</category>
      <category>claudepro</category>
      <category>chatgptplus</category>
    </item>
    <item>
      <title>Google quietly declared war on the open web.</title>
      <dc:creator>Ritabrata Maiti</dc:creator>
      <pubDate>Thu, 24 Sep 2026 08:32:53 +0000</pubDate>
      <link>https://dev.to/ritabratamaiti/google-quietly-declared-war-on-the-open-web-43bd</link>
      <guid>https://dev.to/ritabratamaiti/google-quietly-declared-war-on-the-open-web-43bd</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/cxqv5HIIFVo" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;At Google I/O this week, the Search team made AI Overviews the default&lt;br&gt;
answer surface for English users. The&lt;br&gt;
&lt;a href="https://blog.google/products-and-platforms/products/search/search-io-2026/" rel="noopener noreferrer"&gt;official keynote post&lt;/a&gt;&lt;br&gt;
buries the change under a lot of "agentic" and "personal intelligence"&lt;br&gt;
language, but the practical effect is small and definite: Search now&lt;br&gt;
answers from a Google-hosted synthesis. The blue links are a fallback,&lt;br&gt;
not the product.&lt;/p&gt;

&lt;p&gt;This was the inevitable next step of a series of changes that started&lt;br&gt;
with featured snippets and accelerated with AI Overviews. The&lt;br&gt;
&lt;a href="https://www.techrepublic.com/article/google-ai-overviews-inaccurate-answers-analysis/" rel="noopener noreferrer"&gt;10 percent inaccuracy rate&lt;/a&gt;&lt;br&gt;
on AI Overviews is now everyone's problem, because Overviews are&lt;br&gt;
everyone's default answer.&lt;/p&gt;

&lt;p&gt;The&lt;br&gt;
&lt;a href="https://tante.cc/2026/05/20/on-google-declaring-war-on-the-web/" rel="noopener noreferrer"&gt;piece on tante.cc&lt;/a&gt;&lt;br&gt;
called this "declaring war on the web," and the wording is exactly&lt;br&gt;
right. What follows is what I think the change actually is, the deal&lt;br&gt;
that broke, and what to do about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The deal that built the web
&lt;/h2&gt;

&lt;p&gt;Sites let Google crawl them. Google sent traffic back. Both sides&lt;br&gt;
won. The crawler got fresh content; the site got readers, and the&lt;br&gt;
readers had a clear path back to the source when they wanted to dig&lt;br&gt;
deeper.&lt;/p&gt;

&lt;p&gt;This contract held for about twenty years. It built the open web as&lt;br&gt;
a publishing surface, the indie blog as a thing that paid rent, the&lt;br&gt;
SEO industry, the long-tail of niche reference sites, and most of&lt;br&gt;
what we now call "documentation on the internet." It was an&lt;br&gt;
asymmetric trade, but both sides understood it. You give Google the&lt;br&gt;
catalog. Google sends you the foot traffic.&lt;/p&gt;

&lt;p&gt;AI Overviews quietly cancelled the trade.&lt;/p&gt;

&lt;p&gt;The crawler still runs. The catalog still gets ingested. The output&lt;br&gt;
on the user's screen is now a Google-synthesised answer, with the&lt;br&gt;
sources demoted to a row of tiny chip-shaped citations the median&lt;br&gt;
user does not click. The reader gets an answer, often correct,&lt;br&gt;
sometimes wrong, and never has to leave the search results page.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is, named plainly
&lt;/h2&gt;

&lt;p&gt;The framing in the original piece is sharp and worth quoting at&lt;br&gt;
length:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Your work, your writing or art do matter a bit still: As (unpaid)&lt;br&gt;
raw material for their synthetic text extruders. You get to work&lt;br&gt;
for free so Google can have tight control on the flow of&lt;br&gt;
information and make sure that the responses people get are in&lt;br&gt;
line with what they need them to be.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the right frame. The web is not the product anymore. The&lt;br&gt;
web is the raw material. The product is the synthesised view on top&lt;br&gt;
of it, hosted on Google's infrastructure, branded with Google's UI,&lt;br&gt;
and accountable to Google's editorial choices.&lt;/p&gt;

&lt;p&gt;A useful analogy: the relationship between a wire service and a&lt;br&gt;
newspaper. The wire produces the underlying reporting; the newspaper&lt;br&gt;
picks, edits, and presents. The reader sees the newspaper. For a&lt;br&gt;
long time the web was the newspaper. Google was the index, the table&lt;br&gt;
of contents that pointed at the newspaper. AI Overviews flipped that.&lt;br&gt;
Google is the newspaper now. Your site is the wire.&lt;/p&gt;

&lt;p&gt;The thing that makes this disorienting is that nobody negotiated the&lt;br&gt;
new arrangement. The publishers did not agree to be wire&lt;br&gt;
contributors. The wire contributors did not get paid. The newspaper&lt;br&gt;
masthead never asked.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is bigger than search-traffic math
&lt;/h2&gt;

&lt;p&gt;The temptation is to treat this as an SEO problem. Traffic numbers&lt;br&gt;
will drop, sites will adapt, the market will sort it out. That misses&lt;br&gt;
the scale.&lt;/p&gt;

&lt;p&gt;Google holds the largest single browser deployment on the planet,&lt;br&gt;
the largest single mobile OS, the largest single web video platform,&lt;br&gt;
the dominant share of web standards committee participation, the&lt;br&gt;
dominant share of ad-tech, and now the dominant search-answer&lt;br&gt;
surface. A single company sets the abstraction layer at every point&lt;br&gt;
between the publisher and the reader. When that company decides the&lt;br&gt;
abstraction layer should not include outbound links, there is no&lt;br&gt;
counter-party large enough to disagree at scale.&lt;/p&gt;

&lt;p&gt;This is the part the SEO frame misses. The market does not have a&lt;br&gt;
way to sort out a question of this scale because the market does not&lt;br&gt;
have a participant of this scale. The closest analogue is regulatory.&lt;br&gt;
The EU has begun acting. The DOJ case is grinding along. Neither is&lt;br&gt;
fast enough to matter at the speed Google can ship.&lt;/p&gt;

&lt;p&gt;The web is going to survive this. The independent publisher, the&lt;br&gt;
solo blog, the niche reference site, the corner of the internet&lt;br&gt;
written by one human for fifteen others, will get squeezed in the&lt;br&gt;
process. Some will respond by exiting. Some will paywall. Some will&lt;br&gt;
move to Substack or to mailing lists, which is just running their own&lt;br&gt;
distribution stack. The shape of the open web will change a lot,&lt;br&gt;
faster than the SEO market expects.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you can do as a user
&lt;/h2&gt;

&lt;p&gt;The advice older than the problem still applies. Run your own tools.&lt;br&gt;
Pick your own sources. Don't outsource your reading to a system whose&lt;br&gt;
incentives are to keep you on its surface.&lt;/p&gt;

&lt;p&gt;Concretely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use a search engine that links out.&lt;/strong&gt; Kagi, Marginalia, Mojeek,
Searx, DuckDuckGo, Brave Search. None are perfect, all are linking
out by default. The web still indexes; the index is just not the
default surface anymore.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep an RSS reader.&lt;/strong&gt; Feedbin, Inoreader, NetNewsWire on the Mac,
Readwise Reader if you want web archiving. Subscribing directly to
the publishers you trust routes you around the answer-surface
entirely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build a personal allowlist of go-to sources.&lt;/strong&gt; The handful of
sites you check first when you want a real answer. For me that's
about thirty sites; the search engine is where I go when none of
them have it. This is roughly how everyone used the web in 2005.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read source material when the stakes are high.&lt;/strong&gt; AI Overviews are
wrong 10 percent of the time. That rate is fine for "what is the
recipe for pancakes." It is not fine for "what does this medication
interact with."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern is the same in each case: keep the human in the loop&lt;br&gt;
when the cost of wrong is high, and let the machine condense when the&lt;br&gt;
cost of wrong is low.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this has to do with what I build
&lt;/h2&gt;

&lt;p&gt;I work on &lt;a href="https://browy.dev" rel="noopener noreferrer"&gt;Browy&lt;/a&gt;, an open-source AI agent that&lt;br&gt;
lives in a Chromium extension. It drives the real browser tab you&lt;br&gt;
have open. You ask it a question; it inspects the page, follows&lt;br&gt;
links, reads other tabs, and writes you an answer with the sources&lt;br&gt;
visible inline. It does not synthesise from a hosted index. It&lt;br&gt;
operates on the web you would see if you were doing the work yourself.&lt;/p&gt;

&lt;p&gt;That is not a coincidence. It is the response to exactly the&lt;br&gt;
disintermediation move described above. If the answer-surface is&lt;br&gt;
hosted by a counter-party whose incentives are not yours, the answer&lt;br&gt;
you want is the one you can produce yourself, against the sources&lt;br&gt;
you trust, on the page that loaded in front of you.&lt;/p&gt;

&lt;p&gt;The web is not ending. The middleman is trying to swallow it. Run&lt;br&gt;
your own tools.&lt;/p&gt;




&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://browyhq.github.io/blog/" rel="noopener noreferrer"&gt;All posts&lt;/a&gt; Index of every post on the Browy blog.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>google</category>
      <category>search</category>
      <category>aioverviews</category>
      <category>aisearch</category>
    </item>
    <item>
      <title>GitHub got pwned through one VSCode extension. 3,800 repos.</title>
      <dc:creator>Ritabrata Maiti</dc:creator>
      <pubDate>Thu, 24 Sep 2026 08:32:12 +0000</pubDate>
      <link>https://dev.to/ritabratamaiti/github-got-pwned-through-one-vscode-extension-3800-repos-hgj</link>
      <guid>https://dev.to/ritabratamaiti/github-got-pwned-through-one-vscode-extension-3800-repos-hgj</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/CcBjlKit9IY" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;GitHub confirmed this week that roughly 3,800 of its internal&lt;br&gt;
repositories were exfiltrated after one of its own employees installed&lt;br&gt;
a trojanised version of the Nx Console extension for VSCode. The&lt;br&gt;
malicious extension came through the official VS Code Marketplace, ran&lt;br&gt;
with the editor's permissions, and quietly handed back the contents of&lt;br&gt;
the employee's repo cache. GitHub has since linked the campaign to the&lt;br&gt;
&lt;a href="https://www.bleepingcomputer.com/news/security/github-links-repo-breach-to-tanstack-npm-supply-chain-attack/" rel="noopener noreferrer"&gt;TanStack npm supply chain attack&lt;/a&gt;&lt;br&gt;
from earlier this month and removed the poisoned extension from the&lt;br&gt;
marketplace.&lt;/p&gt;

&lt;p&gt;The TeamPCP group is asking $50,000 for the dump. As far as anyone&lt;br&gt;
outside GitHub can tell, no customer data has moved.&lt;/p&gt;

&lt;p&gt;That is the news. The interesting part is the design that made it&lt;br&gt;
possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  The editor has no sandbox
&lt;/h2&gt;

&lt;p&gt;A VSCode extension runs as Node.js code inside the editor's process.&lt;br&gt;
It has the editor's permissions. That means it can read every file in&lt;br&gt;
every workspace you have ever opened, every token in your&lt;br&gt;
&lt;code&gt;~/.aws/credentials&lt;/code&gt;, every entry in your SSH agent, and every byte&lt;br&gt;
that ever passed through &lt;code&gt;git push&lt;/code&gt;. The Marketplace itself has no&lt;br&gt;
sandbox model. There has been&lt;br&gt;
&lt;a href="https://github.com/microsoft/vscode/issues/52116" rel="noopener noreferrer"&gt;a feature request open since 2018&lt;/a&gt;&lt;br&gt;
asking Microsoft to ship one. The request is still open.&lt;/p&gt;

&lt;p&gt;This is not a VSCode-specific complaint. It is the model. Cursor&lt;br&gt;
inherits the architecture. Windsurf inherits it. JetBrains plugins&lt;br&gt;
run with the IDE's permissions, just packaged differently. Vim and&lt;br&gt;
Emacs have had this problem since plugins were invented; the only&lt;br&gt;
reason it never blew up at this scale is that nobody put a marketplace&lt;br&gt;
in front of them.&lt;/p&gt;

&lt;p&gt;The Marketplace is the part that scales the blast radius. Twelve&lt;br&gt;
months ago, a set of malicious extensions disguised as legitimate&lt;br&gt;
tools racked up&lt;br&gt;
&lt;a href="https://www.bleepingcomputer.com/news/security/vscode-extensions-with-9-million-installs-pulled-over-security-risks/" rel="noopener noreferrer"&gt;9 million installs&lt;/a&gt;&lt;br&gt;
before takedown. A different set posed as&lt;br&gt;
&lt;a href="https://www.bleepingcomputer.com/news/security/malicious-vscode-extensions-infect-windows-with-cryptominers/" rel="noopener noreferrer"&gt;cryptominer-laden helpers&lt;/a&gt;.&lt;br&gt;
A WhiteCobra-flooded batch of twenty-four extensions, one of them with&lt;br&gt;
&lt;a href="https://www.bleepingcomputer.com/news/security/ai-slop-ransomware-test-sneaks-on-to-vs-code-marketplace/" rel="noopener noreferrer"&gt;basic ransomware capability&lt;/a&gt;,&lt;br&gt;
made it onto the marketplace before anyone noticed. And in January, a&lt;br&gt;
pair of malicious AI-coding-assistant extensions with&lt;br&gt;
&lt;a href="https://www.bleepingcomputer.com/news/security/malicious-ai-extensions-on-vscode-marketplace-steal-developer-data/" rel="noopener noreferrer"&gt;1.5 million installs&lt;/a&gt;&lt;br&gt;
silently shipped developer data to servers in China.&lt;/p&gt;

&lt;p&gt;Add it up. Across the visible-to-the-public campaigns, this is more&lt;br&gt;
than 11 million installs of code that should not have been there. The&lt;br&gt;
GitHub breach is one well-placed install away from "any of the above&lt;br&gt;
got an SRE."&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug, named honestly
&lt;/h2&gt;

&lt;p&gt;The threat model in this story is not a clever attacker. The threat&lt;br&gt;
model is a checkbox the operator did not read carefully because they&lt;br&gt;
were trying to get work done.&lt;/p&gt;

&lt;p&gt;A developer at GitHub had an editor open. The editor said: install&lt;br&gt;
this extension. It looked like Nx Console; the developer needed Nx&lt;br&gt;
Console. They clicked install. Three thousand eight hundred internal&lt;br&gt;
repos went out the back door. There is no version of "be more&lt;br&gt;
careful" that fixes this. The marketplace is the install surface and&lt;br&gt;
the marketplace cannot, in its current shape, distinguish the real&lt;br&gt;
extension from the malicious one with the same name.&lt;/p&gt;

&lt;p&gt;The fix is at the layer the operator doesn't have to think about.&lt;br&gt;
That is either:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The editor ships a sandbox. Extensions ask for the file paths,
network endpoints, and OS capabilities they need. The user reviews
a manifest, the same way they review extension permissions in a
browser. Microsoft has been promising this since 2018.&lt;/li&gt;
&lt;li&gt;The marketplace adds an integrity layer that an attacker cannot
forge. Reproducible builds, source mirroring, signed releases.
Nothing on the marketplace today provides this end to end.&lt;/li&gt;
&lt;li&gt;The user installs nothing they have not personally read or that
does not come from a vendor whose source they can audit.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Option 3 is the only one available today. Options 1 and 2 require&lt;br&gt;
Microsoft to ship work that has been promised for seven years.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters for browser extensions
&lt;/h2&gt;

&lt;p&gt;I work on a browser extension, so I want to be clear about what does&lt;br&gt;
and does not transfer from this story.&lt;/p&gt;

&lt;p&gt;Chromium MV3 extensions are not VSCode extensions. The permission&lt;br&gt;
model is real. An extension declares the URLs it can touch&lt;br&gt;
(&lt;code&gt;host_permissions&lt;/code&gt;) and the capabilities it can use (&lt;code&gt;debugger&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;storage&lt;/code&gt;, &lt;code&gt;nativeMessaging&lt;/code&gt;, and so on) in &lt;code&gt;manifest.json&lt;/code&gt;. The&lt;br&gt;
browser enforces those declarations. An extension that says&lt;br&gt;
&lt;code&gt;"host_permissions": ["https://gmail.com/*"]&lt;/code&gt; literally cannot reach&lt;br&gt;
&lt;code&gt;bank.example.com&lt;/code&gt;. The Chrome Web Store review process is not&lt;br&gt;
perfect, but it does look at the manifest before publish.&lt;/p&gt;

&lt;p&gt;That makes the browser-extension surface meaningfully better than the&lt;br&gt;
editor-extension surface for the same class of attack. It does not&lt;br&gt;
make it safe. Things to watch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;&amp;lt;all_urls&amp;gt;&lt;/code&gt;&lt;/strong&gt; is a real permission Chrome will grant if the user
agrees. Many general-purpose extensions ask for it. Browy asks for
it (we have to; you point us at any tab you visit). What stops
abuse is the next thing on the list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The CRX you download is signed.&lt;/strong&gt; It is not necessarily the
source you read. Chrome Web Store does not require reproducible
builds. If the vendor publishes obfuscated code, you are trusting
the vendor, not the code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Native messaging is the escape hatch.&lt;/strong&gt; Browy has a native host
that runs on your machine. So does any "remote control" extension,
any password manager native bridge, and many "AI assistants." A
malicious extension can ship a clean source repo and a malicious
native binary, and the user has no easy way to know.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The list of things that would help Chrome users defend themselves is&lt;br&gt;
roughly the same as for VSCode users: a sandbox the user reviews, a&lt;br&gt;
marketplace integrity layer the attacker cannot forge, and a habit of&lt;br&gt;
installing nothing whose source you have not read.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I do about this in Browy
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://browy.dev" rel="noopener noreferrer"&gt;Browy&lt;/a&gt; is open-source under Apache-2.0. The&lt;br&gt;
source ships in the release. The&lt;br&gt;
&lt;a href="https://github.com/BrowyHQ/browy/blob/main/src/agent/loop.ts" rel="noopener noreferrer"&gt;agent loop is in one file&lt;/a&gt;&lt;br&gt;
and the&lt;br&gt;
&lt;a href="https://github.com/BrowyHQ/browy/blob/main/src/agent/tools/browser.ts" rel="noopener noreferrer"&gt;tool registry is in another&lt;/a&gt;.&lt;br&gt;
A reviewer who reads either file in twenty minutes has read the&lt;br&gt;
attack surface. There is no Browy server. The model call goes from&lt;br&gt;
your machine to GitHub Copilot via the local Copilot CLI; everything&lt;br&gt;
else is local.&lt;/p&gt;

&lt;p&gt;The host-touching tools (&lt;code&gt;bash&lt;/code&gt;, &lt;code&gt;read_file&lt;/code&gt;, &lt;code&gt;write_file&lt;/code&gt;, &lt;code&gt;grep&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;glob&lt;/code&gt;, &lt;code&gt;web_fetch&lt;/code&gt;) are off by default. To use any of them you have&lt;br&gt;
to open Settings, find the tool, and turn it on. The browser-driving&lt;br&gt;
tools are on by default but each can be turned off the same way. We&lt;br&gt;
ship a debugger banner whenever a session is live, because Chrome&lt;br&gt;
makes that mandatory and we are happy with that.&lt;/p&gt;

&lt;p&gt;None of this prevents a future me from going rogue and shipping a&lt;br&gt;
malicious release. It does mean that anyone who cares can pin a&lt;br&gt;
known-good build, audit any future build before installing, and run&lt;br&gt;
their own from source. The fix for the marketplace problem is not a&lt;br&gt;
better marketplace. It is a habit of looking at what you install.&lt;/p&gt;

&lt;p&gt;The short version: when the editor has no sandbox, the only sandbox&lt;br&gt;
left is your attention. Spend it before you click install, not after.&lt;/p&gt;




&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://browyhq.github.io/blog/" rel="noopener noreferrer"&gt;All posts&lt;/a&gt; Index of every post on the Browy blog.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://browyhq.github.io/blog/ai-slop-killed-the-bug-bounty/" rel="noopener noreferrer"&gt;AI slop killed the open-source bug bounty&lt;/a&gt; The other end of the same arc: cheap inference, asymmetric inbox, maintainer attention as the limiting factor.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>github</category>
      <category>vscode</category>
      <category>security</category>
      <category>supplychain</category>
    </item>
    <item>
      <title>AI slop killed the open-source bug bounty</title>
      <dc:creator>Ritabrata Maiti</dc:creator>
      <pubDate>Thu, 24 Sep 2026 08:31:31 +0000</pubDate>
      <link>https://dev.to/ritabratamaiti/ai-slop-killed-the-open-source-bug-bounty-2o64</link>
      <guid>https://dev.to/ritabratamaiti/ai-slop-killed-the-open-source-bug-bounty-2o64</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/8ml0p8BM5Eo" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Turso, the SQLite-compatible database written in Rust, retired its bug&lt;br&gt;
bounty program this week. The post explaining why is titled, dryly,&lt;br&gt;
&lt;a href="https://turso.tech/blog/the-wonders-of-ai" rel="noopener noreferrer"&gt;The wonders of AI&lt;/a&gt;. For&lt;br&gt;
about a year the project had been paying $1,000 for each critical&lt;br&gt;
vulnerability someone reported in the codebase. Budget wasn't the&lt;br&gt;
problem. The problem was that most of the reports landing in the&lt;br&gt;
maintainer's inbox were written by an LLM, and the maintainer was&lt;br&gt;
spending most of his bounty hours reading text that no human had really&lt;br&gt;
intended for him to read.&lt;/p&gt;

&lt;p&gt;The examples are funny in a bleak way. One submission claimed to have&lt;br&gt;
found a critical flaw that let an attacker execute arbitrary SQL&lt;br&gt;
statements. Against a SQL database. Another report on the same project&lt;br&gt;
described a buffer overflow whose reproduction steps included editing&lt;br&gt;
Turso's own source code, recompiling with a forced volatile write past&lt;br&gt;
the end of a vector, and running the modified binary. The "vulnerable&lt;br&gt;
code paths" did not exist. The "exploits" did not reproduce.&lt;/p&gt;

&lt;p&gt;So the maintainers ended the program. That is a small story about one&lt;br&gt;
database, and also a very large story about every queue on the internet&lt;br&gt;
that involves humans on one side and an open submission form on the&lt;br&gt;
other. Every maintainer who has read Turso's post this week has&lt;br&gt;
recognised their own inbox.&lt;/p&gt;

&lt;h2&gt;
  
  
  The contract that worked
&lt;/h2&gt;

&lt;p&gt;The economics of a bug bounty are tidy. The project says: if you find a&lt;br&gt;
real problem that meets this severity bar, we pay you $X. Both sides&lt;br&gt;
win. The hunter gets paid for genuinely useful work. The project gets&lt;br&gt;
a security audit it could not have afforded any other way.&lt;/p&gt;

&lt;p&gt;That contract rested on a quiet assumption. The cost of submitting a&lt;br&gt;
report was supposed to be in the same neighbourhood as the cost of&lt;br&gt;
producing one. Producing a real exploit is hard. Writing one up takes&lt;br&gt;
hours. So submission volume self-throttled. The queue stayed roughly&lt;br&gt;
the size of the real-findings rate, which is the rate a small team can&lt;br&gt;
actually triage.&lt;/p&gt;

&lt;p&gt;Cheap inference broke that assumption in about eighteen months.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;It now costs roughly a tenth of a cent to ask a frontier model to&lt;br&gt;
produce a vulnerability report against a public codebase. The model&lt;br&gt;
will produce one. The report will look real enough that a human has to&lt;br&gt;
spend about ten minutes reading it before concluding that it isn't.&lt;br&gt;
Submission is effectively free. Triage costs the most expensive thing&lt;br&gt;
a software project has, which is the unbroken attention of the one&lt;br&gt;
person who understands the code.&lt;/p&gt;

&lt;p&gt;You don't need many people running that arbitrage to break a small&lt;br&gt;
program. Once the submission rate climbs past the triage rate, the&lt;br&gt;
queue does not stabilise. It grows until the maintainer either gives&lt;br&gt;
up reading or gives up the program. Turso gave up the program.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same shape, everywhere
&lt;/h2&gt;

&lt;p&gt;The bug bounty story is just one shape of a pattern that is now visible&lt;br&gt;
in a dozen places.&lt;/p&gt;

&lt;p&gt;Drive-by pull requests are flooding popular GitHub repos. Most are&lt;br&gt;
LLM-generated typo fixes or hallucinated refactors. Maintainers either&lt;br&gt;
auto-close everything and lose the legitimate ones, or burn out on&lt;br&gt;
triage. Open issue trackers are seeing the same flood of "I think there&lt;br&gt;
might be a bug" reports with no reproduction, often filed by a chatbot&lt;br&gt;
middle layer that promised a confused user it would "let the developers&lt;br&gt;
know." Code review comments are now arriving thirty at a time on small&lt;br&gt;
PRs, mostly tautologies and hedges. The comment threads on technical&lt;br&gt;
blog posts are heading the same way.&lt;/p&gt;

&lt;p&gt;The shape is always the same. Producing the input costs nothing.&lt;br&gt;
Processing the input still costs a person.&lt;/p&gt;

&lt;p&gt;We have automated finding bugs. We have automated submitting bugs. This&lt;br&gt;
year we are automating rejecting bugs. Nobody is automating fixing them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bottleneck was never the writing
&lt;/h2&gt;

&lt;p&gt;A lot of the early excitement about LLMs in software was that they&lt;br&gt;
could write code very fast. That turned out to be true, and it turned&lt;br&gt;
out to be the less interesting half of the problem. Writing the code&lt;br&gt;
was never the bottleneck. Reading the code is the bottleneck. Reviewing&lt;br&gt;
the code, understanding the code, deciding whether to merge the code,&lt;br&gt;
deciding whether to ship the code, deciding what to do when the code&lt;br&gt;
breaks at 2am. The reading-and-deciding loop is where engineering&lt;br&gt;
actually happens, and it has always been bounded by how many things a&lt;br&gt;
human can hold in their head at once.&lt;/p&gt;

&lt;p&gt;LLMs add to the writing side of that loop and add to the input side of&lt;br&gt;
every queue feeding into it. They don't move the reading-and-deciding&lt;br&gt;
bottleneck. They make it the limiting factor for almost everything.&lt;/p&gt;

&lt;p&gt;This is why "just put an AI in the triage loop" doesn't fix the Turso&lt;br&gt;
problem. If there's a model deciding which reports a human sees,&lt;br&gt;
you've created a new game: produce reports that pass the model. That&lt;br&gt;
game is also cheap to play. The judgment step that the attacker is&lt;br&gt;
trying to commodify is the exact step that has to stay with a person,&lt;br&gt;
because that's what was being attacked in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  What might actually hold
&lt;/h2&gt;

&lt;p&gt;A handful of patterns are starting to circulate in maintainer&lt;br&gt;
conversations. None are clean.&lt;/p&gt;

&lt;p&gt;A refundable submission deposit is the most-discussed option. You put&lt;br&gt;
down a small amount, say $20. You get it back if the report holds up.&lt;br&gt;
You lose it if it's slop. The economics work. The downside is that it&lt;br&gt;
filters out the hobbyist newcomer who doesn't have a card on file,&lt;br&gt;
which is exactly the population the open programs were designed to&lt;br&gt;
find.&lt;/p&gt;

&lt;p&gt;Reputation gating goes the other way. First-time submitters have to&lt;br&gt;
clear a higher bar, including a full reproduction and a vouching&lt;br&gt;
introduction. Established hunters skip the bar. This rebuilds an&lt;br&gt;
apprenticeship model that hasn't really existed in software security&lt;br&gt;
for a generation, but it raises the cost of discovering new talent.&lt;/p&gt;

&lt;p&gt;A few projects have started running honeypots. The clearest example is&lt;br&gt;
&lt;a href="https://github.com/UnsafeLabs/Bounty-Hunters" rel="noopener noreferrer"&gt;UnsafeLabs/Bounty-Hunters&lt;/a&gt;,&lt;br&gt;
which publishes a bounty-shaped repo that exists mainly to attract&lt;br&gt;
automated scanners. The submissions feed a public&lt;br&gt;
&lt;a href="https://clankers-leaderboard.pages.dev" rel="noopener noreferrer"&gt;leaderboard of automated submitters&lt;/a&gt;.&lt;br&gt;
You can read it as petty, or you can read it as the first attempt at&lt;br&gt;
something a search-engine spam team would recognise: a shared blocklist&lt;br&gt;
for the actor side of an asymmetric inbox problem.&lt;/p&gt;

&lt;p&gt;Some teams are moving programs off public surfaces entirely and onto&lt;br&gt;
verified-identity platforms like HackerOne or Bugcrowd, where the&lt;br&gt;
platform handles attribution. Small projects can't usually afford the&lt;br&gt;
platform fees. Bigger ones lose the discovery effect of an open&lt;br&gt;
program.&lt;/p&gt;

&lt;p&gt;The most interesting idea I've seen circulating among maintainers is&lt;br&gt;
what someone called proof of code. Before you can submit a bounty&lt;br&gt;
report, you have to land a small, accepted PR somewhere in the&lt;br&gt;
project. The PR doesn't have to be security-related. It just has to be&lt;br&gt;
real engineering by a person who understood the surrounding code. That&lt;br&gt;
converts a cheap text-generation game into a more expensive engineering&lt;br&gt;
game. It doesn't solve the problem. It tilts the asymmetry the other&lt;br&gt;
way.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is bigger than bounties
&lt;/h2&gt;

&lt;p&gt;The same dynamic is coming for everything that depends on an open&lt;br&gt;
queue with humans on the receiving end. Customer support tickets.&lt;br&gt;
Conference paper submissions. Open peer review. Public consultation&lt;br&gt;
periods on regulation. Government FOIA requests. Hiring inboxes. The&lt;br&gt;
comment thread under this post, probably.&lt;/p&gt;

&lt;p&gt;The economic concept is older than the internet. It is the tragedy of&lt;br&gt;
the commons applied to attention instead of grazing land. What's new&lt;br&gt;
is the rate at which the cost of producing input is falling. Most of&lt;br&gt;
the institutions that grew up around the old cost curve don't have&lt;br&gt;
time to redesign before the next drop.&lt;/p&gt;

&lt;p&gt;The next decade of public-facing system design is going to be about&lt;br&gt;
adding friction back in on purpose. Identity. Deposits. Reputation.&lt;br&gt;
Proof of work. Vouching. We spent twenty years stripping friction out&lt;br&gt;
of every form on the internet because friction was the enemy of&lt;br&gt;
growth. Now friction is the only thing protecting the people on the&lt;br&gt;
receiving end of those forms. The companies that figure out how to add&lt;br&gt;
it without losing the magic are going to eat everyone else.&lt;/p&gt;

&lt;h2&gt;
  
  
  A note on what "AI agent" means
&lt;/h2&gt;

&lt;p&gt;There are two things people now mean by "AI agent," and they are&lt;br&gt;
nearly opposite.&lt;/p&gt;

&lt;p&gt;One is the slop pipeline that broke Turso. Cheap inference pointed at&lt;br&gt;
any open queue, with no human in the loop, generating volume that&lt;br&gt;
other humans then have to filter. The person whose name is on the&lt;br&gt;
submission usually isn't watching what's submitted. Often they don't&lt;br&gt;
even know what was submitted.&lt;/p&gt;

&lt;p&gt;The other is the tool a person uses to do their actual work. An agent&lt;br&gt;
that drives the browser tab they're sitting in front of, that reads&lt;br&gt;
the codebase they're working on, that runs the command they would&lt;br&gt;
have typed. The person is at the keyboard. They see every action.&lt;br&gt;
Nothing gets sent to anyone else's queue without them watching it&lt;br&gt;
happen.&lt;/p&gt;

&lt;p&gt;These two things share a phrase and almost nothing else. The first one&lt;br&gt;
is a tax on every public surface on the internet. The second one is a&lt;br&gt;
power tool. The fact that they have the same name in the press is&lt;br&gt;
going to cost the second category a lot of goodwill before things&lt;br&gt;
sort themselves out.&lt;/p&gt;

&lt;p&gt;I work on the second kind. &lt;a href="https://browy.dev" rel="noopener noreferrer"&gt;Browy&lt;/a&gt; is an&lt;br&gt;
open-source AI agent that lives in a Chrome side panel and a DevTools&lt;br&gt;
REPL. It drives your real browser tabs through chat. There is no Browy&lt;br&gt;
server in the loop. There is no inbox you can flood by talking to it.&lt;br&gt;
There is a person, you, watching every click and every form fill it&lt;br&gt;
makes. Whatever happens to the public commons over the next few years,&lt;br&gt;
the part of the internet that's a tool you operate yourself is still&lt;br&gt;
going to be a good place to live.&lt;/p&gt;

&lt;p&gt;The 30-second video version of this argument is at the top of this&lt;br&gt;
post. Turso's own write-up is&lt;br&gt;
&lt;a href="https://turso.tech/blog/the-wonders-of-ai" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;




&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://browyhq.github.io/blog/" rel="noopener noreferrer"&gt;All posts&lt;/a&gt; Index of every post on the Browy blog.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://browyhq.github.io/blog/mullvad-vpn-fingerprint/" rel="noopener noreferrer"&gt;Mullvad gave you 8 trillion exit IPs. 9 servers found you.&lt;/a&gt; Same researcher-pattern thinking, different domain: a VPN whose deterministic IP picker leaks identity across servers.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aislop</category>
      <category>bugbounty</category>
      <category>turso</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Mullvad gave you 8 trillion exit IPs. 9 servers found you.</title>
      <dc:creator>Ritabrata Maiti</dc:creator>
      <pubDate>Thu, 24 Sep 2026 08:30:05 +0000</pubDate>
      <link>https://dev.to/ritabratamaiti/mullvad-gave-you-8-trillion-exit-ips-9-servers-found-you-4md8</link>
      <guid>https://dev.to/ritabratamaiti/mullvad-gave-you-8-trillion-exit-ips-9-servers-found-you-4md8</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/g-B_toki7PA" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Mullvad is one of the few VPN providers that hands you a different public&lt;br&gt;
IP for each of its servers. With 578 servers in the fleet, a naive&lt;br&gt;
calculation gives you over 8 trillion possible combinations of exit IPs&lt;br&gt;
across the network. The pitch is that you melt into a very large crowd.&lt;/p&gt;

&lt;p&gt;A researcher who writes as tmctmt sat down with a script, rotated through&lt;br&gt;
3,650 WireGuard keys, and watched which exit IPs each key was assigned&lt;br&gt;
across nine servers. They expected to see thousands of distinct&lt;br&gt;
combinations. They saw 284. The&lt;br&gt;
&lt;a href="https://tmctmt.com/posts/mullvad-exit-ips-as-a-fingerprinting-vector/" rel="noopener noreferrer"&gt;full write-up is here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The cause is unglamorous and the impact is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the picker works
&lt;/h2&gt;

&lt;p&gt;When you connect to a Mullvad server, the exit IP you get is not random.&lt;br&gt;
It is picked deterministically from a per-server pool, seeded by your&lt;br&gt;
WireGuard public key. The pool sizes vary by server. Sydney exposes 60&lt;br&gt;
IPs, Helsinki 66, Los Angeles 91, Santiago 11. The key rotates every&lt;br&gt;
one to thirty days if you use the official app, and never if you bring&lt;br&gt;
your own WireGuard client.&lt;/p&gt;

&lt;p&gt;So far so reasonable. Per-key stickiness avoids hammering a single&lt;br&gt;
external IP with users who reconnect every few minutes, and it lets&lt;br&gt;
sites that ratelimit by IP behave sanely for a Mullvad user inside one&lt;br&gt;
session.&lt;/p&gt;

&lt;p&gt;The problem is the picker itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  One float, many pools
&lt;/h2&gt;

&lt;p&gt;The picker seeds a standard library RNG with the pubkey, then draws a&lt;br&gt;
single floating point number in the range zero to one. It scales that&lt;br&gt;
float to the size of the current server's pool to produce an index.&lt;br&gt;
That part is plausible. The part that breaks privacy is that the float&lt;br&gt;
is the same float on every server, because the RNG is seeded with the&lt;br&gt;
same pubkey and only the upper bound changes between calls.&lt;/p&gt;

&lt;p&gt;The consequence is that your exit IPs across servers land in the same&lt;br&gt;
percentile of each pool. If your float happens to be 0.82, you sit at&lt;br&gt;
position 49 of 60 in Sydney, position 9 of 11 in Santiago, position 54&lt;br&gt;
of 66 in Helsinki. Different IPs, same relative slot.&lt;/p&gt;

&lt;p&gt;If you have ever wondered why two of the smaller Mullvad servers seem&lt;br&gt;
to give you the "same" index, this is why. Santiago and Johannesburg&lt;br&gt;
both have pools of 11. With the same float scaled to the same size,&lt;br&gt;
they hand you the same position.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why 8.2 trillion collapsed to 284
&lt;/h2&gt;

&lt;p&gt;This is the headline of the post. The pool sizes multiply to 8.2 trillion&lt;br&gt;
on paper. In practice the float is one number, so the entire vector of&lt;br&gt;
exit IPs across all servers is parameterised by that one number. The&lt;br&gt;
researcher's 3,650 sampled pubkeys produced only 284 distinct&lt;br&gt;
combinations across nine servers because the space being sampled is&lt;br&gt;
not thirteen-dimensional. It is one-dimensional, then projected.&lt;/p&gt;

&lt;p&gt;The 284 figure is the resolution of that projection at nine servers.&lt;br&gt;
With more servers the resolution rises, because each additional pool&lt;br&gt;
size carves the unit interval into finer buckets. With fewer servers&lt;br&gt;
it drops, but it does not drop as much as you would hope.&lt;/p&gt;

&lt;h2&gt;
  
  
  What that means for a user
&lt;/h2&gt;

&lt;p&gt;The author built a small tool that takes a set of observed exit IPs and&lt;br&gt;
back-solves for the float interval that produced them. Nine well-chosen&lt;br&gt;
exit IPs squeeze the interval down to roughly 0.0034 wide. At an&lt;br&gt;
estimated 100,000 active Mullvad users that is around 340 people.&lt;/p&gt;

&lt;p&gt;Three hundred and forty is a lot of people in a crowd. It is not a lot&lt;br&gt;
of people in a deanonymisation attack. If a forum has IP logs and&lt;br&gt;
suspects two accounts are the same user, and the two accounts both&lt;br&gt;
connected through Mullvad on different servers, the overlap of the&lt;br&gt;
float intervals derived from those exits gives a very high probability&lt;br&gt;
that the accounts share a pubkey. Pair that with timing, user agent,&lt;br&gt;
language headers, or the kind of low-effort browser fingerprint that&lt;br&gt;
ad networks already collect, and the pool of suspects shrinks to one.&lt;/p&gt;

&lt;p&gt;The interesting attack is not a moderator with one forum's logs. It is&lt;br&gt;
the joining of logs across services that already happens in litigation,&lt;br&gt;
in subpoenas, and in stolen breach data. Mullvad's per-server IP&lt;br&gt;
diversity stops being a defence the moment two unrelated services&lt;br&gt;
contribute their exit IP records to the same correlation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The response
&lt;/h2&gt;

&lt;p&gt;Mullvad's co-founder posted a response on the morning the article went&lt;br&gt;
up. The short version: some of the described behaviour was intentional,&lt;br&gt;
some was not, the unintended part is being patched on a subset of the&lt;br&gt;
fleet first, and the intended part is now under internal review. He&lt;br&gt;
also asked, gently, that future researchers consider notifying the&lt;br&gt;
vendor before publishing.&lt;/p&gt;

&lt;p&gt;That is about the best response a small vendor can give to a Tuesday&lt;br&gt;
morning surprise, and it points at the part of this story that is most&lt;br&gt;
worth sitting with.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug under the bug
&lt;/h2&gt;

&lt;p&gt;The proximate bug is a seeded RNG whose first draw was treated as a&lt;br&gt;
fresh random number per call, when it is the same number every call.&lt;br&gt;
A defensible code review could miss this. The Rust documentation does&lt;br&gt;
not put the behaviour on a billboard, and the test that would have&lt;br&gt;
caught it, "do my exit IP percentiles correlate across pools," is not&lt;br&gt;
the test anyone writes.&lt;/p&gt;

&lt;p&gt;The deeper bug is the design assumption that per-key stickiness is&lt;br&gt;
compatible with per-server independence. It is not, for any picker&lt;br&gt;
that derives its index from a single seeded float. To get true&lt;br&gt;
independence you have to either reseed per call with a per-server salt,&lt;br&gt;
or pre-mix the pubkey with the server identity before drawing, or&lt;br&gt;
abandon determinism and accept the operational cost.&lt;/p&gt;

&lt;p&gt;There is a clean lesson here about cryptographic determinism that I&lt;br&gt;
will not labour, but the practical one for anyone shipping&lt;br&gt;
privacy-flavoured software is worth saying out loud. Deterministic&lt;br&gt;
behaviour is the thing that makes user experience predictable and the&lt;br&gt;
thing that makes users linkable across observations. The instinct to&lt;br&gt;
make a system feel coherent across sessions is the same instinct that&lt;br&gt;
hands an analyst a join key.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a user can do today
&lt;/h2&gt;

&lt;p&gt;The author lists two mitigations. Both reduce the value of the&lt;br&gt;
fingerprint without removing it.&lt;/p&gt;

&lt;p&gt;The first is to avoid switching servers while your pubkey is unchanged.&lt;br&gt;
If you only ever connect from one Mullvad server, an observer sees one&lt;br&gt;
exit IP and gets one of the pool's possible float intervals, not the&lt;br&gt;
intersection of many. That interval covers a much larger share of the&lt;br&gt;
user base.&lt;/p&gt;

&lt;p&gt;The second is to force-rotate your pubkey, which the Mullvad app will&lt;br&gt;
do if you log out and log back in. A new pubkey produces a new float&lt;br&gt;
and a new vector of exit IPs. Past observations of you, before the&lt;br&gt;
rotation, still link to whatever activity happened under the old key.&lt;/p&gt;

&lt;p&gt;Neither of these is satisfying. Both are what the user can do until&lt;br&gt;
the patch lands.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bigger picture for privacy tools
&lt;/h2&gt;

&lt;p&gt;VPNs are sold as a tool for blending in. The category has spent ten&lt;br&gt;
years competing on server count, jurisdiction, and "we don't log."&lt;br&gt;
None of that helps if the tunnel itself emits a stable signal that&lt;br&gt;
correlates across exits. Mullvad is one of the better-regarded&lt;br&gt;
operators in the category, run by people who care about this stuff&lt;br&gt;
and respond like it. If this design slipped past them, it is worth&lt;br&gt;
assuming similar designs are sitting unexamined in operators who&lt;br&gt;
care less.&lt;/p&gt;

&lt;p&gt;The healthier framing for anyone using a privacy tool is that the tool&lt;br&gt;
moves the boundary of who can see you. It does not make you invisible&lt;br&gt;
to the people inside the new boundary. A VPN protects you from your&lt;br&gt;
ISP and from the cafe Wi-Fi. It does not protect you from the website&lt;br&gt;
you are visiting, which still sees your browser, your timing, and&lt;br&gt;
whatever your tunnel happens to leak about your identity across&lt;br&gt;
sessions. Pretending otherwise is how people end up surprised.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this has to do with browsers
&lt;/h2&gt;

&lt;p&gt;A lot of the leak surface in modern web sessions is in the browser,&lt;br&gt;
not in the tunnel. Canvas fingerprinting, font enumeration, audio&lt;br&gt;
context fingerprinting, third party cookies that you thought were&lt;br&gt;
blocked, redirect chains that smuggle identifiers in the URL, the&lt;br&gt;
clipboard, the way your tabs talk to each other, the way an extension&lt;br&gt;
you forgot you installed talks to its origin. The interesting thing&lt;br&gt;
about the Mullvad story is that it is one of the rare cases where the&lt;br&gt;
leak is below the browser, and the browser would not see it without&lt;br&gt;
help.&lt;/p&gt;

&lt;p&gt;I work on &lt;a href="https://browy.dev" rel="noopener noreferrer"&gt;Browy&lt;/a&gt;, an open-source AI agent that&lt;br&gt;
lives in a Chrome side panel and a DevTools REPL. It drives the real&lt;br&gt;
tab you are looking at. You can point it at a page and ask, in&lt;br&gt;
English, what the page is sending, what it stores, what it loads from&lt;br&gt;
where, what it is trying to identify you with. It will open the&lt;br&gt;
Network tab and the Application tab and read them for you, and it&lt;br&gt;
will tell you what it found in the same chat where you asked the&lt;br&gt;
question. It does not send your browsing anywhere. The model call&lt;br&gt;
goes out and the answer comes back, the rest happens on your machine.&lt;/p&gt;

&lt;p&gt;The Mullvad bug is a server-side problem. Browy will not see it. But&lt;br&gt;
the broader habit, watching what your browser actually emits while&lt;br&gt;
you are using it, is what every category of privacy leak in the last&lt;br&gt;
decade has rewarded. The 30-second video version of the Mullvad story&lt;br&gt;
is at the top of this post. The full research, with the seed&lt;br&gt;
estimator tool, is&lt;br&gt;
&lt;a href="https://tmctmt.com/posts/mullvad-exit-ips-as-a-fingerprinting-vector/" rel="noopener noreferrer"&gt;at tmctmt's site&lt;/a&gt;.&lt;/p&gt;




&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://browyhq.github.io/blog/" rel="noopener noreferrer"&gt;All posts&lt;/a&gt; Index of every post on the Browy blog.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://browyhq.github.io/blog/ai-slop-killed-the-bug-bounty/" rel="noopener noreferrer"&gt;AI slop killed the open-source bug bounty&lt;/a&gt; The same shape in software security: a real database project just retired its bounty program because AI submissions broke the queue.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>mullvad</category>
      <category>vpn</category>
      <category>fingerprinting</category>
      <category>privacy</category>
    </item>
    <item>
      <title>AnyModal: Train Multimodal LLMs in PyTorch</title>
      <dc:creator>Ritabrata Maiti</dc:creator>
      <pubDate>Tue, 19 Nov 2024 11:13:42 +0000</pubDate>
      <link>https://dev.to/ritabratamaiti/anymodal-simplify-multimodal-ai-development-with-a-flexible-framework-4o12</link>
      <guid>https://dev.to/ritabratamaiti/anymodal-simplify-multimodal-ai-development-with-a-flexible-framework-4o12</guid>
      <description>&lt;p&gt;Today, I want to introduce an open-source framework I’ve been working on: &lt;strong&gt;&lt;a href="https://github.com/ritabratamaiti/AnyModal" rel="noopener noreferrer"&gt;AnyModal&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Introduction
&lt;/h3&gt;

&lt;p&gt;During my work on machine learning projects, I struggled to find flexible solutions for training multimodal LLMs. While there are plenty of great tools for specific tasks—like image classification or audio processing—there was no straightforward way to combine these modalities with large language models (LLMs). The process was often tedious, involving boilerplate code, custom integration, and a lot of trial and error to make different components work together.&lt;/p&gt;

&lt;p&gt;This frustration led me to build &lt;strong&gt;AnyModal&lt;/strong&gt;, a framework designed to reduce the complexity of multimodal AI development. It provides a modular, reusable structure that makes it easier for developers and researchers to combine diverse data types and experiment with new ideas without reinventing the wheel every time.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Goal
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;AnyModal&lt;/strong&gt; is built with the following objectives in mind:  &lt;/p&gt;

&lt;h4&gt;
  
  
  Reduce Boilerplate Code
&lt;/h4&gt;

&lt;p&gt;Combining modalities like images or audio with LLMs typically involves repetitive steps—preprocessing, encoding, tokenizing, and integrating. AnyModal minimizes this boilerplate by providing reusable modules for common tasks, letting developers focus on building smarter systems faster.  &lt;/p&gt;

&lt;h4&gt;
  
  
  Enable Seamless Integration
&lt;/h4&gt;

&lt;p&gt;Whether you're working with images using a Vision Transformer (ViT) or audio spectrograms, AnyModal offers plug-and-play components that simplify the integration process. This makes it easy to handle multiple data types within a single framework.  &lt;/p&gt;

&lt;h4&gt;
  
  
  Encourage Experimentation and Customization
&lt;/h4&gt;

&lt;p&gt;AnyModal supports rapid prototyping while offering the flexibility to customize components like feature encoders, projection layers, and tokenizers. It’s versatile enough for both quick experiments and production-level deployments.  &lt;/p&gt;




&lt;h3&gt;
  
  
  Example Usage: Integrating Images with LLMs
&lt;/h3&gt;

&lt;p&gt;Here’s a detailed example of how AnyModal simplifies the integration of image data into LLMs:&lt;/p&gt;

&lt;h4&gt;
  
  
  1. Install Dependencies
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;torch transformers datasets torchvision tqdm
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  2. Initialize Vision Components
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ViTImageProcessor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ViTForImageClassification&lt;/span&gt;

&lt;span class="c1"&gt;# Load a pre-trained Vision Transformer
&lt;/span&gt;&lt;span class="n"&gt;processor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ViTImageProcessor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;google/vit-base-patch16-224&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;vision_model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ViTForImageClassification&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;google/vit-base-patch16-224&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Define a Vision Encoder to extract feature embeddings
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;vision&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;VisionEncoder&lt;/span&gt;
&lt;span class="n"&gt;vision_encoder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;VisionEncoder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vision_model&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  3. Initialize Tokenizer and LLM
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;AutoModelForCausalLM&lt;/span&gt;

&lt;span class="c1"&gt;# Load a pre-trained LLM and its tokenizer
&lt;/span&gt;&lt;span class="n"&gt;llm_tokenizer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;llm_model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoModelForCausalLM&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  4. Define a Projection Layer
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;vision&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Projector&lt;/span&gt;

&lt;span class="c1"&gt;# Create a projection layer to map vision embeddings to LLM token space
&lt;/span&gt;&lt;span class="n"&gt;vision_tokenizer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Projector&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;in_features&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;vision_model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;hidden_size&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; 
    &lt;span class="n"&gt;out_features&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;768&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  5. Combine Everything with AnyModal
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;anymodal&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;MultiModalModel&lt;/span&gt;

&lt;span class="c1"&gt;# Build the multimodal model
&lt;/span&gt;&lt;span class="n"&gt;multimodal_model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MultiModalModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;input_processor&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;input_encoder&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;vision_encoder&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;input_tokenizer&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;vision_tokenizer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;language_tokenizer&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;llm_tokenizer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;language_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;llm_model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;input_start_token&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;|imstart|&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;input_end_token&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;|imend|&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;prompt_text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Describe this image: &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  6. Training and Inference
&lt;/h4&gt;

&lt;p&gt;Training involves processing batches of image-text pairs and optimizing the model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;torch.utils.data&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DataLoader&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datasets&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;load_dataset&lt;/span&gt;

&lt;span class="c1"&gt;# Load a sample dataset
&lt;/span&gt;&lt;span class="n"&gt;dataset&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;load_dataset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_caption_dataset&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;train&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Prepare DataLoader
&lt;/span&gt;&lt;span class="n"&gt;train_loader&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;DataLoader&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dataset&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;batch_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shuffle&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Training Loop
&lt;/span&gt;&lt;span class="n"&gt;optimizer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;optim&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;AdamW&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;multimodal_model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parameters&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;lr&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;3e-4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;epoch&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;batch&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;train_loader&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;optimizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;zero_grad&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;logits&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;loss&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;multimodal_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;loss&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;backward&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;optimizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;step&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# Generate captions
&lt;/span&gt;&lt;span class="n"&gt;sample_input&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dataset&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;image&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;generated_caption&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;multimodal_model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sample_input&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_new_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Generated Caption:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;generated_caption&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  Current Status
&lt;/h3&gt;

&lt;p&gt;AnyModal is currently in its early stages, with the latest version supporting tasks like:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LaTeX OCR
&lt;/li&gt;
&lt;li&gt;Chest X-Ray Captioning (in progress)
&lt;/li&gt;
&lt;li&gt;Image Captioning
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Future planned features include support for visual question answering and audio captioning.  &lt;/p&gt;

&lt;p&gt;As the framework evolves, I’m focusing on expanding its functionality, refining the codebase, and addressing community feedback to move towards a stable release.  &lt;/p&gt;




&lt;h3&gt;
  
  
  Links
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub&lt;/strong&gt;: &lt;a href="https://github.com/ritabratamaiti/AnyModal" rel="noopener noreferrer"&gt;https://github.com/ritabratamaiti/AnyModal&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reddit&lt;/strong&gt;: &lt;a href="https://www.reddit.com/r/AnyModal/" rel="noopener noreferrer"&gt;https://www.reddit.com/r/AnyModal/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hugging Face&lt;/strong&gt;: &lt;a href="https://huggingface.co/AnyModal" rel="noopener noreferrer"&gt;https://huggingface.co/AnyModal&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;If you’re looking for a way to simplify multimodal AI development, give &lt;strong&gt;AnyModal&lt;/strong&gt; a try. I’d love to hear your feedback or ideas for new features. Contributions are always welcome!&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>llm</category>
      <category>ai</category>
      <category>multimodal</category>
    </item>
  </channel>
</rss>
