<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dragos Roua</title>
    <description>The latest articles on DEV Community by Dragos Roua (@dragos_roua).</description>
    <link>https://dev.to/dragos_roua</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3626747%2Fa8e48f6d-1290-4b50-9369-7be234c9f1ec.jpg</url>
      <title>DEV Community: Dragos Roua</title>
      <link>https://dev.to/dragos_roua</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dragos_roua"/>
    <language>en</language>
    <item>
      <title>AI API Token Prices Are Not Real - They're Aspirational</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Sun, 06 Sep 2026 07:50:04 +0000</pubDate>
      <link>https://dev.to/dragos_roua/ai-api-token-prices-are-not-real-theyre-aspirational-4npg</link>
      <guid>https://dev.to/dragos_roua/ai-api-token-prices-are-not-real-theyre-aspirational-4npg</guid>
      <description>&lt;p&gt;&lt;em&gt;They are what neolabs like Anthropic and OpenAI need you to believe so they can make back their investors’ money.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  But You Just Paid the Invoice
&lt;/h2&gt;

&lt;p&gt;Yes, I hear you. You just paid a Claude / Codex / Grok invoice and it was probably in the hundreds or thousands of dollars. I’m not saying those tokens you just paid don’t exist. I’m saying they are not based on a real market mechanism. They’re an expectation, an aspiration, they’re what neolabs hope to receive. How is that even possible?&lt;/p&gt;

&lt;p&gt;I tried to explains everything in the video above.&lt;/p&gt;

&lt;p&gt;Here’s a short summary, so you know what to expect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Token Price Even Exists
&lt;/h2&gt;

&lt;p&gt;A SOTA model is made of 2 things: data and compute. Both are extremely expensive. Neolabs took a lot of money from investors to train those models, and they came up with something plausible, and lately, something that can even produce production level code.&lt;/p&gt;

&lt;p&gt;But this is horrendously expensive. Now, as a normal user, you wouldn’t pay the real price, simply because it’s prohibitive and it’s still more effective to hire a developer. So neolabs started to sell subscriptions, which have a number of tokens included. It’s like tasting the product. If you want more, you get the price per token.&lt;/p&gt;

&lt;p&gt;The best way to understand this is to think at a Ferrari: it’s an extremely expensive car, you cannot afford it probably. So you do not buy the car. You rent ten hours. Renting is the price by the token. The gap between those two prices is unusually large because in the case of AI the product is new and demand is still low.&lt;/p&gt;

&lt;h2&gt;
  
  
  We Have 3 Layers of Price
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Real cost — billions. Nobody can afford this.&lt;/li&gt;
&lt;li&gt;Lab token price — close to what they need to look solvent in front of investors, not necessarily what the work is really worth.&lt;/li&gt;
&lt;li&gt;Open source / local — often 10× cheaper, 95–98% good enough for coding, analysis, and admin.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;a href="https://dragosroua.com/running-local-private-ai-models-how-and-why/" rel="noopener noreferrer"&gt;local, private AI&lt;/a&gt; is starting to make sense financially, too. Six months ago a local DeepSeek-class box was $25–50k. Now ~$10k (Mac Studio or two DGX Sparks) gets you Qwen / GLM territory near last-gen Opus / GPT.&lt;/p&gt;

&lt;h2&gt;
  
  
  Something Changed in the Last 6 Months
&lt;/h2&gt;

&lt;p&gt;Now we know what are the real use cases for LLMs: coding, data work, email/meeting cleanup. These are impressive but in and by themselves do not automatically support the rates neolabs are asking for frontier models.&lt;/p&gt;

&lt;p&gt;The reality will probably kick in in the next 3–6 months:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Either neoabs cut prices to defend their market share, or&lt;/li&gt;
&lt;li&gt;Open-source (China + US) takes over the market entirely&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The IPO is the key moment and we are a couple of months away from this.&lt;/p&gt;

</description>
      <category>api</category>
      <category>ai</category>
      <category>programming</category>
      <category>ipo</category>
    </item>
    <item>
      <title>Anthropic and OpenAI Are Both Expecting IPOs at Trillion Dollar Valuations. But Will They Win the AI Race?</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Fri, 04 Sep 2026 02:23:58 +0000</pubDate>
      <link>https://dev.to/dragos_roua/anthropic-and-openai-are-both-expecting-ipos-at-trillion-dollar-valuations-but-will-they-win-the-5505</link>
      <guid>https://dev.to/dragos_roua/anthropic-and-openai-are-both-expecting-ipos-at-trillion-dollar-valuations-but-will-they-win-the-5505</guid>
      <description>&lt;p&gt;A one-trillion-dollar valuation is no joke. Anthropic and OpenAI are both at this level, and they are both preparing for IPOs.&lt;/p&gt;

&lt;p&gt;Of course, expectations are high. Many people will bet insane amounts of money on these two companies. But are they even in a good position to win the AI race?&lt;/p&gt;

&lt;p&gt;I try to find the answers in the video above. Here’s a brief description of what’s in it.&lt;/p&gt;

&lt;p&gt;Both Anthropic and OpenAI started as chatbots, then quickly pivoted to coding. LLMs are extremely performant pattern matching machines, so coding, which is a general language domain with very small vocabulary and easy grammar, was a low-hanging fruit. Both companies are relying extensively on this and the majority of their users range from vibe coders to enterprises looking to cut their costs.&lt;/p&gt;

&lt;p&gt;But there’s life outside the tech / digital vertical. Too many people get stuck in their own opinions and end up in echo chambers. While outside these echo chambers, established companies like Google, Meta or SpaceXAI are already finding use cases outside coding. Google is embedding AI snippets in search results. Meta uses the Ask Meta megabot to increase engaging, and SpaceXAI is tweaking its algorithm using Grok, while at the same time uses Grok as an affordable conversation partner. All these 3 companies are quiet, but there’s a fundamental difference between them and the 2 main contenders.&lt;/p&gt;

&lt;p&gt;And that difference is users. They already have billions of users, acquired through years of marketing. Whereas Anthropic and OpenAI are either resorting to fear mongering marketing, or spend insane amounts of cash on user acquisition.&lt;/p&gt;

&lt;p&gt;So, while being way more quiet, the 3 established companies have actually a better chance at winning the AI race. And that boils down to 3 parameters: users (already talked about that, but worth mentioning another key detail), compute and data.&lt;/p&gt;

&lt;p&gt;So, not only they have users, but they have paying users and they know their spending habits. Converting those users to AI-powered services will be a no-brainer.&lt;/p&gt;

&lt;p&gt;They also have compute. Some of them, like Google and Meta are even building their own chips, while SpaceXAI built 2 mega data centers, Colossus 1 and Colossus 2, which are already renting compute to Anthropic.&lt;/p&gt;

&lt;p&gt;And of course, they have a lot of data, gathered during years of operation. They can train way better models and at more competitive prices.&lt;/p&gt;

&lt;p&gt;Among these 3 established companies, SpaceXAI deserves a special mention. Google and Meta are essentially digital companies, but SpaceXAI can also distribute AI in electric cars, in Optimus robots, and, soon, via Neuralink brain-to-computer interfaces.&lt;/p&gt;

&lt;p&gt;If you like the video, like and share, and if you have any other topics you want to me touch on, just live a comment on YouTube, I read them all.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>ipo</category>
      <category>claude</category>
      <category>githubcopilot</category>
    </item>
    <item>
      <title>You Can Now Share Your Grok Bots (And Soon You Could Sell Them)</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Sat, 29 Aug 2026 03:37:30 +0000</pubDate>
      <link>https://dev.to/dragos_roua/you-can-now-share-your-grok-bots-and-soon-you-could-sell-them-2oec</link>
      <guid>https://dev.to/dragos_roua/you-can-now-share-your-grok-bots-and-soon-you-could-sell-them-2oec</guid>
      <description>&lt;p&gt;SpaceXAI just made Grok Bot templates shareable. I had early access to the feature, meaning I published / shared my first Grok Bot before the feature went live. In this post I’ll walk you trough the process of creating and sharing your own. But before that, be aware the SpaceXAI confirmed a marketplace for bots will be up and running as early as next week.&lt;/p&gt;

&lt;p&gt;I already &lt;a href="https://dragosroua.com/grok-bot-is-the-iphone-moment-for-ai-agents-quick-review/" rel="noopener noreferrer"&gt;called Grok Bot the iPhone moment for agents&lt;/a&gt;. Not because any of its parts were brand new. In the case of the iPhone: the music player, the computer and the phone — they all existed before. But the way they were put together created a product with its own class. The same thing was true for Grok Bot, but now we have even more evidence. What App Store was to iPhone, the upcoming bot marketplace could be for Grok Bot.&lt;/p&gt;

&lt;p&gt;Shareable templates it’s what makes this possible.&lt;/p&gt;

&lt;p&gt;Anything you build inside the app — give it a job, let it grow a routine — can now be packaged, shared and it gets its own link. Someone else opens that link and drops the bot into their own Grok Bot. It’s that simple.&lt;/p&gt;

&lt;p&gt;I walked through it with something I named Outbid Mania. There’s nothing fancy about the name, I was just curious to observe the phenomenon behind it. The entire specification was just the prompt: look at outbid.lol, look at the clones, get revenue for both, and then — this is the part that matters — tell me if the virality is real or just noise. I asked for raw data but also a personal opinion. I wanted to see how Grok bot will change its opinion if something had happened.&lt;/p&gt;

&lt;p&gt;Then I waited.&lt;/p&gt;

&lt;p&gt;It quickly figured out the concept. It went online, it fetched the data, then it sets up a routine by itself: a dashboard every morning around 6:30. That routine is “embedded”, it travels with the shared bot. That’s why routines are not a cute extra. They’re part of the product you hand someone.&lt;/p&gt;

&lt;p&gt;After its first pass, it thought the virality was big. Second pass, the data moved and it changed its mind. That’s the behavior I want from an agent. And that’s exactly what I expect in general, a little bit of opinion and judgement.&lt;/p&gt;

&lt;p&gt;Publishing is the part that should have been complicated but, surprisingly, it isn’t. You just tap the agent name in the corner, then scroll, then publish / share your agent. Deceptively simple. You get a unique link and that’s where the bot lives. If you paste the link in a browser you get a product page: title, what it does, a button to add it to your own install. It’s almost like a stripped down App Store.&lt;/p&gt;

&lt;p&gt;Claude’s ecosystem tried to do this with artifacts. They didn’t quite pick up. Maybe too complicate? A Grok Bot template is a URL and a preview. That’s closer to how people actually send things to each other.&lt;/p&gt;

&lt;p&gt;Now, my agent was fairly simple, just a prompt and nothing else. But try to think what happens when you put the template next to plugins and connectivity and the usage surface explodes. You can ask the bot to watch Tesla or Apple or Nvidia shares and then get an alert when price crosses a certain threshold. You can even give the bot a budget and tell it to act certain levels.&lt;/p&gt;

&lt;p&gt;I don’t see this stopping anywhere soon: the limit is whatever people will invent.&lt;/p&gt;

&lt;p&gt;So invent. If you find a template worth existing — a watcher, a researcher, a morning briefing, a thing that files PRs while you sleep — put the idea in the comments on the video. Share the ones that work. When the marketplace is up I’ll drop a short tutorial here.&lt;/p&gt;

&lt;p&gt;The iPhone moment was never “brand new stuff.” It was “this just works, and now other people can have the same object.” Templates are how Grok Bot evolves to its own App Store.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>automation</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Grok Bot is the "iPhone Moment" for AI Agents</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Fri, 28 Aug 2026 04:02:08 +0000</pubDate>
      <link>https://dev.to/dragos_roua/grok-bot-is-the-iphone-moment-for-ai-agents-30ce</link>
      <guid>https://dev.to/dragos_roua/grok-bot-is-the-iphone-moment-for-ai-agents-30ce</guid>
      <description>&lt;p&gt;I finally got into Grok Bot. Yesterday, actually. Two false starts before that — SpaceXAI rolled this out slowly and I thought I was in when I actually wasn't.&lt;/p&gt;

&lt;p&gt;This is a very quick rundown of what I was able to see in the last 24 hours.&lt;/p&gt;

&lt;p&gt;It comes with its own computer. That’s a bigger deal than it sounds. If you’ve been running your own agent stack (OpenClaw, Hermes, a VPS just to keep the thing alive), you already know what I'm talking about. Here you get the machine with the bot. And of course you still get your local computer, so it can edit files and ship code on the Mac too.&lt;/p&gt;

&lt;p&gt;It draws tokens from a separate bucket. If you’re on Grok or Cursor, Bot doesn’t just pull from the same quota. SpaceXAI is usually cautious with tokens — they drain fast. This time it seems reasonable.&lt;/p&gt;

&lt;p&gt;The UI is suspiciously thin. You look at it and think nothing lives there. Then you set up three agents and it starts working. I did:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one for iOS / mobile apps&lt;/li&gt;
&lt;li&gt;one for web, the 20-year-old blog, a few content sites, a couple of SaaS&lt;/li&gt;
&lt;li&gt;Analytics Maverick, which talks to the other two&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Setup was dumb simple. It went to the local machine, asked for separate Git credentials so it could open PRs and push, and found the repos even though I hadn’t given it a lot of context. Not spectacular, but it just worked.&lt;/p&gt;

&lt;p&gt;Then the magic ensued: Analytics Maverick started talking to the web and mobile agents. I supervised here and there, but everything was 90% autonomous. That was the part that doesn’t feel like traditional software. It feels like a coworker hiding in the silicon of the MacBook Pro.&lt;/p&gt;

&lt;p&gt;Which gets me back to the "iPhone moment".&lt;/p&gt;

&lt;p&gt;Steve Jobs didn’t invent the phone, the music player, or the computer. They were there before him. But he put them in one object and it became a new class of product. Agents on your machine existed before. same fort Inter-agent chat and VPS babysitting. But this is the first version that feels like a product instead of an Ikea assembly kit.&lt;/p&gt;

&lt;p&gt;Now about the cots, because this matters too. If you’re on nothing, it starts around $60/month. I pay about $50 total — $30 SuperGrok, $20 Cursor entry — and Bot came along with its own usage. So it's cheaper to run 2 minimal subscription for now. And since we're there: please don’t do the “I have 10 Claude subs and 20 ChatGPT plans and I’m at $400–$1,000/month” thing unless the return is actually there. Easy to spend, but way harder to get it back.&lt;/p&gt;

&lt;p&gt;If you get access to Grok Bot, tell me what you think, I'm curious about your use cases.&lt;/p&gt;

&lt;p&gt;If you prefer watching this as a video, click the YouTube thing (I'm testing a new setup, recording from coffee shops).&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How I Use Grok Build to Run Free Models Alongside Grok-4.6</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Tue, 25 Aug 2026 03:26:28 +0000</pubDate>
      <link>https://dev.to/dragos_roua/how-i-use-grok-build-to-run-free-models-alongside-grok-46-1106</link>
      <guid>https://dev.to/dragos_roua/how-i-use-grok-build-to-run-free-models-alongside-grok-46-1106</guid>
      <description>&lt;p&gt;I really like Grok Build as a harness. The terminal UI is polished without too much glitter, the tools are decent, and there’s a certain way it sticks to a project. What I like less is having access only to the Grok models lineup. What if I can have the same rich layer, but talking to open weights models?&lt;/p&gt;

&lt;p&gt;So I did what every AI obsessed person does these days: started to dig into config files, squeezing every ounce of juice from everywhere I can. It turned out that Grok Build talks to anything that speaks an OpenAI-compatible chat API. It can be a model on your machine, or a free model behind OpenRouter. You use the same interface and you flip models with /model and keep working.&lt;/p&gt;

&lt;p&gt;The whole customization is in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;~/.grok/config.toml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Below is my actual setup. There are two different options, you can pick just one, or keep both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Option 1: Local models
&lt;/h2&gt;

&lt;p&gt;For this you need something serving the model on localhost. llama.cpp, LM Studio, Ollama — whatever you already like, as long as it exposes &lt;code&gt;/v1/chat/completions&lt;/code&gt;. I configured my server to listen on port 8080, but you can choose whatever you want, it's local anyway.&lt;/p&gt;

&lt;p&gt;Then you tell Grok Build about it in the config.toml file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="nn"&gt;[model.gemma-4-12b-local]&lt;/span&gt; 
&lt;span class="py"&gt;model&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"gemma-4-12b-it-qat-q4_0"&lt;/span&gt; 
&lt;span class="py"&gt;base_url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"http://127.0.0.1:8080/v1"&lt;/span&gt; 
&lt;span class="py"&gt;name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"Gemma 4 12B QAT Q4_0 (local)"&lt;/span&gt; 
&lt;span class="py"&gt;api_backend&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"chat_completions"&lt;/span&gt; 
&lt;span class="py"&gt;context_window&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;16384&lt;/span&gt; 
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="err"&gt;model.qwen&lt;/span&gt;&lt;span class="mi"&gt;-3-8-27&lt;/span&gt;&lt;span class="err"&gt;b-local&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; 
&lt;span class="py"&gt;model&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"Qwen3.8-27B-UD-IQ2_S"&lt;/span&gt; 
&lt;span class="py"&gt;base_url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"http://127.0.0.1:8080/v1"&lt;/span&gt; 
&lt;span class="py"&gt;name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"Qwen 3.8 27b QT 2 (local)"&lt;/span&gt; 
&lt;span class="py"&gt;api_backend&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"chat_completions"&lt;/span&gt; 
&lt;span class="py"&gt;context_window&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;16384&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On my M1 16GB MacBook Pro I can run heavily quantized models in the Gemma / Qwen layer, nothing above that. Performance isn’t great, but it’s local. A few notes that can save you some time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;model must match the id your local server expects, not the actual marketing name.&lt;/li&gt;
&lt;li&gt;base_url ends at /v1. Grok Build appends the rest.&lt;/li&gt;
&lt;li&gt;context_window should be kept minimal if you're low on RAM (like I am). If you aim for 128k and the quantized model only holds 16k, you get compaction at nearly every prompt and long sessions are completely amnesic.&lt;/li&gt;
&lt;li&gt;No API key needed for localhost.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Start the server, which loads the weights, restart Grok Build, then in the new session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/model gemma-4-12b-local
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or you can start grok directly with the model as an argument:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;grok &lt;span class="nt"&gt;-m&lt;/span&gt; gemma-4-12b-local
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And that’s the whole local option. It’s offline, you don’t pay anything, and it’s private by default. Like I said, quality depends on your hardware, specifically RAM, and how aggressively you quantized. As a rule of thumb, meaningful work can be done if you have over 32GB of RAM, 16GB, like I do now, is mainly for experiments.&lt;/p&gt;

&lt;h2&gt;
  
  
  Option 2: OpenRouter free models
&lt;/h2&gt;

&lt;p&gt;Local is really great, but most of the time you want a bigger model than your machine can hold. OpenRouter has a free tier for a bunch of open models. We will use the same harness, but with a different endpoint.&lt;/p&gt;

&lt;p&gt;You will need an OpenRouter API key for this. Generate one in your dashboard, then add it to your environment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENROUTER_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-or-..."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then define the provider once, so you don’t repeat yourself for every model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="nn"&gt;[model_providers.openrouter]&lt;/span&gt; 
&lt;span class="py"&gt;base_url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"https://openrouter.ai/api/v1"&lt;/span&gt; 
&lt;span class="py"&gt;env_key&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"OPENROUTER_API_KEY"&lt;/span&gt; 
&lt;span class="py"&gt;api_backend&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"chat_completions"&lt;/span&gt; 
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="err"&gt;model_providers.openrouter.extra_headers&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; 
&lt;span class="py"&gt;HTTP-Referer&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"https://x.ai"&lt;/span&gt; 
&lt;span class="py"&gt;X-Title&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"Grok Build"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then add the free models you care about. Here’s my non-exhaustive list:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="nn"&gt;[model.glm-5-2-free]&lt;/span&gt; 
&lt;span class="py"&gt;model&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"z-ai/glm-5.2:free"&lt;/span&gt; 
&lt;span class="py"&gt;name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"GLM 5.2 (OpenRouter free)"&lt;/span&gt; 
&lt;span class="py"&gt;model_provider&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"openrouter"&lt;/span&gt; 
&lt;span class="py"&gt;base_url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"https://openrouter.ai/api/v1"&lt;/span&gt; 
&lt;span class="py"&gt;env_key&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"OPENROUTER_API_KEY"&lt;/span&gt; 
&lt;span class="py"&gt;api_backend&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"chat_completions"&lt;/span&gt; 
&lt;span class="py"&gt;context_window&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;256000&lt;/span&gt;&lt;span class="err"&gt;`&lt;/span&gt;

&lt;span class="nn"&gt;[model.openrouter-free]&lt;/span&gt; 
&lt;span class="py"&gt;model&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"openrouter/free"&lt;/span&gt; 
&lt;span class="py"&gt;name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"OpenRouter free router"&lt;/span&gt; 
&lt;span class="py"&gt;description&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"Routes to whatever free OpenRouter model is available"&lt;/span&gt; 
&lt;span class="py"&gt;model_provider&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"openrouter"&lt;/span&gt; 
&lt;span class="py"&gt;base_url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"https://openrouter.ai/api/v1"&lt;/span&gt; 
&lt;span class="py"&gt;env_key&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"OPENROUTER_API_KEY"&lt;/span&gt; 
&lt;span class="py"&gt;api_backend&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"chat_completions"&lt;/span&gt; 
&lt;span class="py"&gt;context_window&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;200000&lt;/span&gt; 
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="err"&gt;model.gemma&lt;/span&gt;&lt;span class="mi"&gt;-4-31&lt;/span&gt;&lt;span class="err"&gt;b-free&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="py"&gt;model&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"google/gemma-4-31b-it:free"&lt;/span&gt; 
&lt;span class="py"&gt;name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"Gemma 4 31B (OpenRouter free)"&lt;/span&gt; 
&lt;span class="py"&gt;model_provider&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"openrouter"&lt;/span&gt; 
&lt;span class="py"&gt;base_url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"https://openrouter.ai/api/v1"&lt;/span&gt; 
&lt;span class="py"&gt;env_key&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"OPENROUTER_API_KEY"&lt;/span&gt; 
&lt;span class="py"&gt;api_backend&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"chat_completions"&lt;/span&gt; 
&lt;span class="py"&gt;context_window&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;262144&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;:free suffix&lt;/code&gt; is very important. Without it you hit the paid route. &lt;code&gt;openrouter/free&lt;/code&gt; is also an interesting option - it picks whatever free model is available that day. I find it really cool for experiments. But for real work I pin a specific one, usually GLM 5.2 free or Gemma 4 31B free, so behavior stays somewhat consistent.&lt;/p&gt;

&lt;p&gt;In the Grok Build harness you switch the same way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/model glm-5-2-free
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And you can also check what Grok Build can see:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;grok models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How I Actually Use This
&lt;/h2&gt;

&lt;p&gt;Almost 90% of my work sessions are on a paid Grok model, I maximize the full harness and its tools. I switch to local for private or offline sessions. And I choose OpenRouter free when I want a bigger open model without drawing from my Grok usage. Sometimes I go for models like Nemotron or DeepSeek Flash. I didn’t include the configs for those in this article, I just leave this as a little bit of homework for you.&lt;/p&gt;

&lt;p&gt;The important part is the config file. Once ~/.grok/config.toml knows about a model, the rest of Grok Build treats it like any other: tools, sessions, /model, Ctrl+M picker. You're never starting a second app onto your workflow. You're pointing the same app at a different brain.&lt;/p&gt;

&lt;p&gt;If something fails to connect, curl the endpoint first. For local:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://127.0.0.1:8080/v1/models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For OpenRouter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://openrouter.ai/api/v1/models &lt;span class="se"&gt;\ &lt;/span&gt;&lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OPENROUTER_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If curl is happy and Grok Build isn’t, it’s almost always a typo in model, base_url, or env_key.&lt;/p&gt;

&lt;p&gt;That’s my entire setup: one harness, two free paths and endless choices. All you have to do is edit the toml, restart or switch with &lt;code&gt;/model&lt;/code&gt;, and keep building.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How To Choose an Open Weights Model - Episode 4</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Sun, 23 Aug 2026 04:38:27 +0000</pubDate>
      <link>https://dev.to/dragos_roua/how-to-choose-an-open-weights-model-episode-4-354m</link>
      <guid>https://dev.to/dragos_roua/how-to-choose-an-open-weights-model-episode-4-354m</guid>
      <description>&lt;p&gt;The previous three episodes covered how to evaluate open-weights models, how to access them through providers, and what it takes to run them on your own hardware.&lt;/p&gt;

&lt;p&gt;The last piece is the harness — the software that sits between you and the model.&lt;/p&gt;

&lt;p&gt;A harness can be very simple or quite sophisticated. In practice people move through three main layers.&lt;/p&gt;

&lt;p&gt;The first is a minimal chat interface such as llama.cpp. You point it at a model file and start a chat. There is no conversation history beyond the current session, no tool calling, and no isolation from the rest of the system. It is useful for a quick test, especially after downloading a quantized model, because setup is minimal and feedback is immediate.&lt;/p&gt;

&lt;p&gt;The second layer is the more complete command-line interfaces. Examples include the CLIs from Anthropic, OpenAI’s Codex-style CLI, and Grok’s own build tooling. These add three practical features: sandboxing (the model does not operate directly on your files), tool use (the model can call external functions such as web search or file searching), and session management (you can pause and resume work, sometimes even from a phone). For people who already live in the terminal (guilty as charged), this is often the daily driver.&lt;/p&gt;

&lt;p&gt;The third layer consists of much more complex interfaces such as Pi or Hermes. These provide graphical or conversational UIs, better support for spoken interaction, and in some cases even connections to messaging apps. They feel closer to a personal assistant. Non-technical users often prefer this level once they decide they want something more polished.&lt;/p&gt;

&lt;p&gt;Most people do not need to start at the top. A basic tool is enough to decide whether a particular model is worth keeping. Once daily use begins, a proper CLI usually becomes the practical choice. The more elaborate interfaces are optional and mainly about comfort and reach.&lt;/p&gt;

&lt;p&gt;That completes the short series. The four episodes together give an easy way to evaluate: the metrics that describe a model, the two ways of accessing it, the hardware choice for local inference, and the software layer that makes the model usable.&lt;/p&gt;

&lt;p&gt;The full playlist is here: &lt;a href="https://youtube.com/playlist?list=PLKxL2crAoLBU&amp;amp;si=3nDQeNgCw51rFDuG" rel="noopener noreferrer"&gt;https://youtube.com/playlist?list=PLKxL2crAoLBU&amp;amp;si=3nDQeNgCw51rFDuG&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How To Choose an Open Weights Model - Episode 3</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Sat, 22 Aug 2026 03:19:21 +0000</pubDate>
      <link>https://dev.to/dragos_roua/how-to-choose-an-open-weights-model-episode-3-514b</link>
      <guid>https://dev.to/dragos_roua/how-to-choose-an-open-weights-model-episode-3-514b</guid>
      <description>&lt;p&gt;In the first two episodes I covered the four metrics that matter when you evaluate an open-weights model, and how to access those models through third-party providers.&lt;/p&gt;

&lt;p&gt;This episode is about running the open weights model local, on your own hardware.&lt;/p&gt;

&lt;p&gt;The main constraint is basically RAM. The weights file has to fit in the memory of the machine you are using. If the file is larger than the available RAM (or VRAM), performance simply vanishes, it just doesn't work. Quantization is the main metric that you need to watch, to make local use practical. Models are commonly released or converted at 2-bit, 4-bit, 6-bit or 8-bit precision. Lower bit depth produces a smaller file and usually significant loss in quality.&lt;/p&gt;

&lt;p&gt;Two hardware lines dominate local inference right now.&lt;/p&gt;

&lt;p&gt;Apple Silicon covers the MacBook Air, MacBook Pro, Mac Mini and Mac Studio. Memory ranges from 16 GB on the low end up to 512 GB on the highest configurations. The unified memory architecture is unusually fast for this kind of workload, which is why even mid-range Macs can feel responsive with well-quantized models.&lt;/p&gt;

&lt;p&gt;NVIDIA is the other main option. Consumer GPUs typically offer between 8 GB and 24 GB of VRAM. Larger systems built on the Blackwell platform (DGX Spark and boards from other manufacturers that follow the same specification) provide more capacity. In some setups it is possible to connect multiple machines and pool their memory.&lt;/p&gt;

&lt;p&gt;A rough practical map looks like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;16–32 GB: 2-bit to 4-bit quantized models in the smaller-to-mid size range. Inference is usable for many everyday tasks, but very slow&lt;/li&gt;
&lt;li&gt;64–128 GB: 70-billion-parameter models become comfortable. Quality can approach what people currently get from strong closed models.&lt;/li&gt;
&lt;li&gt;256 GB and above: large models such as the bigger DeepSeek or Kimi releases become realistic.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is cheap at the high end, and the very large configurations remain out of reach for most individuals. The mid-range, however, is already practical for a lot of real work.&lt;/p&gt;

&lt;p&gt;Running the model is only part of the task. You still need a way to talk to it, manage sessions, and give it tools. That layer is called the harness, and it is the subject of the final episode.&lt;/p&gt;

&lt;p&gt;Episode 3 is live. The full series playlist is here: &lt;a href="https://youtube.com/playlist?list=PLKxL2crAoLBU&amp;amp;si=3nDQeNgCw51rFDuG" rel="noopener noreferrer"&gt;https://youtube.com/playlist?list=PLKxL2crAoLBU&amp;amp;si=3nDQeNgCw51rFDuG&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How To Choose an Open Weights Model - Episode 2</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Fri, 21 Aug 2026 03:00:35 +0000</pubDate>
      <link>https://dev.to/dragos_roua/how-to-choose-an-open-weights-model-episode-2-4nca</link>
      <guid>https://dev.to/dragos_roua/how-to-choose-an-open-weights-model-episode-2-4nca</guid>
      <description>&lt;p&gt;In &lt;a href="https://dragosroua.com/video-how-to-chose-an-open-weights-model-episode-1/" rel="noopener noreferrer"&gt;Episode 1&lt;/a&gt; I covered the four metrics that actually matter when you evaluate an open-weights model: parameter count, architecture (dense vs MoE), quantization, and whether the model is text-only or multimodal.&lt;/p&gt;

&lt;p&gt;Once you know what to look for, the next practical question is: how do you actually use the model?&lt;/p&gt;

&lt;p&gt;There are two main ways.&lt;/p&gt;

&lt;p&gt;You can download the weights and run the model yourself, or you can use a third-party provider that already hosts it for you. This episode is about the second option — the providers.&lt;/p&gt;

&lt;p&gt;When a new open-weights model is released, a bunch of companies that already have data centers and spare compute quickly make it available. You don’t have to download anything. They host the model and expose it through an API, usually in the OpenAI style. That means almost every tool, library, or agent framework you already use can talk to it without major changes.&lt;/p&gt;

&lt;p&gt;The biggest practical advantage is convenience. With a single API key from one provider you usually get access to a whole set of models: DeepSeek, Kimi, MiniMax, GLM, Qwen, and others. You can switch between them depending on the task instead of managing separate accounts and keys.&lt;/p&gt;

&lt;p&gt;Pricing is almost always pay-as-you-go and split into two parts: you pay for the input tokens (what you send to the model) and for the output tokens (what the model generates). This becomes important with models that “think” a lot or produce long responses, because the output side can add up quickly. Still, even with that, the cost is usually a fraction of what you pay for the big closed models.&lt;/p&gt;

&lt;p&gt;Most providers also do some form of intelligent routing in the background. They automatically send your request to the instance that has the best availability or the lowest latency at that moment. You don’t have to manage that yourself.&lt;/p&gt;

&lt;p&gt;This approach works especially well when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You have reliable internet&lt;/li&gt;
&lt;li&gt;You’re working from a thin laptop and move around a lot&lt;/li&gt;
&lt;li&gt;You don’t want to carry (or buy) a machine with 64 or 128 GB of RAM just to run models&lt;/li&gt;
&lt;li&gt;Your work spans different categories — coding one hour, writing the next, maybe some audio or video tasks later — and you want to switch models easily&lt;/li&gt;
&lt;li&gt;In short, third-party providers give you the convenience of closed models while still using open weights, usually at a much lower cost.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Of course this is only one side of the story. The other side is running the models yourself on your own hardware — either on a laptop or on a small server under the desk. That’s what Episode 3 will cover, including the realistic hardware requirements and the trade-offs involved.&lt;/p&gt;

&lt;p&gt;And if you haven’t seen Episode 1 yet, start there: &lt;a href="https://youtu.be/iBOh7atzUZY" rel="noopener noreferrer"&gt;https://youtu.be/iBOh7atzUZY&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I’ll post the next ones as they go live.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How To Choose an Open Weights Model - Episode 1</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Wed, 19 Aug 2026 13:25:10 +0000</pubDate>
      <link>https://dev.to/dragos_roua/how-to-choose-an-open-weights-model-episode-1-2pbh</link>
      <guid>https://dev.to/dragos_roua/how-to-choose-an-open-weights-model-episode-1-2pbh</guid>
      <description>&lt;p&gt;State of the art AI models like Claude, Grok or Codex are very powerful, but they get expensive fast. If you use them seriously, you’re looking at $100–300 a month, and sometimes more. The cheaper plans, which are all $20, burn through their limits quickly, and the higher tiers start to feel like too much spending when you’re using them every day for coding, writing, or agent work. Just by doing normal stuff you’re between $50-$100.&lt;/p&gt;

&lt;p&gt;Open weights models have improved a lot in the last year. In many cases you can now get similar quality to the closed models, but at a fraction of the price — or completely free if you run them on your own hardware.&lt;/p&gt;

&lt;p&gt;The catch is that choosing the right one is still confusing – even for me, if I’m being honest. There is no single company telling you “this one is good for coding, this one is good for agents, this one is too heavy for your machine.” You have to figure it out yourself.&lt;/p&gt;

&lt;p&gt;I got tired of that, so I decided to make a short 4-episode mini-series that walks through the entire decision process. No theory for the sake of theory, just the basic things that actually help you pick a model and use it.&lt;/p&gt;

&lt;p&gt;Episode 1 is already live and you can watch it from here. In it I cover:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why open weights models are finally worth serious attention right now&lt;/li&gt;
&lt;li&gt;The real difference between open source and open weights (most people mix these two up)&lt;/li&gt;
&lt;li&gt;The only four metrics that matter when you evaluate a model&lt;/li&gt;
&lt;li&gt;Parameter count and what the different size ranges are actually good for (1–8B, 14–70B, 100B+)&lt;/li&gt;
&lt;li&gt;Dense architecture versus Mixture of Experts (MoE), and why this changes local use requirements&lt;/li&gt;
&lt;li&gt;Quantization levels (2-bit, 4-bit, 6-bit, 8-bit) and the quality versus size trade-offs&lt;/li&gt;
&lt;li&gt;Text-only models versus multimodal models, and when each one makes sense
The episode is short and focused. I kept it under ten minutes so you can watch it once and come away with a clear mental checklist.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rest of the series will be in the same ballpark and they will cover:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Episode 2: accessing models through third-party providers and APIs&lt;/li&gt;
&lt;li&gt;Episode 3: running models locally on your own hardware (and what the these hardware requirements look like)&lt;/li&gt;
&lt;li&gt;Episode 4: harnesses — the software layer that sits on top of the model and makes it usable, from simple command-line tools up to full agent interfaces&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you’re already paying for closed models and wondering whether open weights are good enough for your workflow, start with Episode 1. It gives you the basic language and the four criteria you need before you start downloading anything.&lt;/p&gt;

&lt;p&gt;I’ll post the next ones as they go live.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How AI Watermarking Works? Can It Be Removed?</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Tue, 18 Aug 2026 08:03:05 +0000</pubDate>
      <link>https://dev.to/dragos_roua/how-ai-watermarking-works-can-it-be-removed-379g</link>
      <guid>https://dev.to/dragos_roua/how-ai-watermarking-works-can-it-be-removed-379g</guid>
      <description>&lt;p&gt;Short answer: yes, it can be removed.&lt;/p&gt;

&lt;p&gt;Long answer, it’s a bit more complicated. To get the full picture, just watch the video, it’s roughly 5 minutes.&lt;/p&gt;

&lt;p&gt;Key points:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Watermarking is implemented because EU regulations&lt;/li&gt;
&lt;li&gt;It survives copy and paste, because it’s based on probabilistic token choices, not on invisible characters&lt;/li&gt;
&lt;li&gt;Heavily editing the text basically removes this type of watermarking&lt;/li&gt;
&lt;li&gt;If you want to avoid it entirely, using local, open source models is the safest way&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In this video: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How Anthropic’s implementation (using Google’s SynthID) actually works &lt;/li&gt;
&lt;li&gt;Why longer text is much easier to detect than short replies &lt;/li&gt;
&lt;li&gt;The surprising side-effect: even human text you paste into Claude for proofreading can get watermarked &lt;/li&gt;
&lt;li&gt;Two practical ways to remove or avoid the watermark &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Timestamps: &lt;/p&gt;

&lt;p&gt;0:00 – Intro: What AI watermarking is &lt;/p&gt;

&lt;p&gt;0:55 – How the watermark is applied (token-level changes) &lt;/p&gt;

&lt;p&gt;2:28 – Why short chats are safer than long-form content &lt;/p&gt;

&lt;p&gt;2:48 – The proofreading trap &lt;/p&gt;

&lt;p&gt;3:13 – Method 1: Heavy human editing (recommended) &lt;/p&gt;

&lt;p&gt;3:48 – Method 2: Open-source models (with an important caveat) &lt;/p&gt;

&lt;p&gt;5:14 – EU regulation context + final thoughts &lt;/p&gt;

&lt;p&gt;Key takeaway: Watermarking is not baked into the model weights — it is applied after generation. &lt;/p&gt;

&lt;p&gt;Local open-source models on your own hardware currently give the strongest protection. &lt;/p&gt;

&lt;p&gt;If you found this kind of content useful, like and subscribe to my channel, it really makes me wanna publish more.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>programming</category>
    </item>
    <item>
      <title>Why China Will Not Win the AI Race (Spoiler: It Already Did)</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Sun, 16 Aug 2026 09:42:49 +0000</pubDate>
      <link>https://dev.to/dragos_roua/why-china-will-not-win-the-ai-race-spoiler-it-already-did-ko4</link>
      <guid>https://dev.to/dragos_roua/why-china-will-not-win-the-ai-race-spoiler-it-already-did-ko4</guid>
      <description>&lt;p&gt;For the last week I’ve been on a trip to China. Although I’m quite used to Asia (see below) I’ve never been to mainland China. My only previous contact was a trip to Hong Kong more than 15 years ago. I just returned yesterday, so the experience is still fresh. If you’re in a rush, here’s the TLDR:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;China is way more relevant than Western media portrays it&lt;/li&gt;
&lt;li&gt;It has a big advantage in AI, and the current context is favorable for an even faster evolution&lt;/li&gt;
&lt;li&gt;The “communism” vs “democracy” narrative has limited application here
Now, let’s dive in.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Context
&lt;/h2&gt;

&lt;p&gt;You may wonder why and how this article may be relevant to you, so here are a few things that clarify the context:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I know communism first hand, I lived in a communist country (Romania) until I was 19&lt;/li&gt;
&lt;li&gt;I’ve lived in South East Asia for about 3 years now, in Vietnam, and on top of that I visited extensively Japan, Korea, Thailand and Bali.&lt;/li&gt;
&lt;li&gt;I’ve studied AI since before ChatGPT 3, back then it was called “machine learning”&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A Very Big Place
&lt;/h2&gt;

&lt;p&gt;First thing that hits you in China is that it’s big. It has a certain comfort of dimensions, so to speak, that makes it seem endless. You don’t even have to travel for days, you get it in the first few hours.&lt;/p&gt;

&lt;p&gt;The infrastructure — I mean roads and general transportation — is second to none. I never experienced any sort of traffic jam, there are spaghetti junctions even in the smaller cities and that makes transportation feel solved.&lt;/p&gt;

&lt;p&gt;The level of development seems to even out. We traveled from Hangzhou to Shanghai by bus and I never noticed any significant difference. By comparison, once you get out of Saigon to travel, let’s say, to Da Nang, you will notice a completely different kind of city, more crowded and significantly less developed.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Being big and more or less uniform makes it a very good market. One that can absorb new products very fast.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  From Proletariat to Middle Class
&lt;/h2&gt;

&lt;p&gt;I had, for about 3 years, the biggest car portal in Romania. I still know the brands and understand the economy behind owning and using a car. Chinese cars are on the top tier: plenty of Tesla and BYD (I just learned it means “Build Your Dreams”), or less known brands like AION or Roewe. Your regular BMW and Mercedes are common place. Almost all bikes I’ve seen, especially in Shanghai, are electric.&lt;/p&gt;

&lt;p&gt;Based on what I’ve seen on the roads, China doesn’t have a proletariat class anymore — I guess this may be true for a few decades already, I just had my own confirmation.&lt;/p&gt;

&lt;p&gt;This is a very important detail, from a business perspective.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A strong middle class will consume better and more expensive products. Being also evened out, without too many differences between rich and poor, will make it absorb the products faster.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Size Matters
&lt;/h2&gt;

&lt;p&gt;So, putting these together, you get a very relevant player, one that is probably sharing the first place with USA as the leader. The only difference is that China doesn’t have 850 military bases across the world, but this is a political view that I will not touch. Let’s stick to the economics.&lt;/p&gt;

&lt;p&gt;Very big country + functioning middle class + great infrastructure = way more relevant than the Western media portrays.&lt;/p&gt;

&lt;p&gt;If you’ve never been there and you only consume your information from Western media, you are biased. To make an informed decision, you really have to be there.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Obstacle Is The Way — The AI Version
&lt;/h2&gt;

&lt;p&gt;It still baffles me that one of the most famous books in the West, The Obstacle Is The Way, was not yet read by the decision makers. If they did read it, they will never impose any ban on China.&lt;/p&gt;

&lt;p&gt;Because of all this pressure they applied, now Chinese companies have to find new ways to compete. And they did find new ways. It was the restrictions that made them better.&lt;/p&gt;

&lt;p&gt;Here are just a few examples: DeepSeek, Kimi, MiniMax, Qwen (backed by Alibaba, which I think it has strong ties with the government, the only player with really big media exposure).&lt;/p&gt;

&lt;p&gt;So, trying to mechanically stopping China is backfiring big time now.&lt;/p&gt;

&lt;h2&gt;
  
  
  But They Are Communist, right? Right?
&lt;/h2&gt;

&lt;p&gt;During my trip we made on average 300km per day on the roads. I never once saw any communist propaganda. For what it’s worth, I’ve never seen any kind of advertising either.&lt;/p&gt;

&lt;p&gt;This is in stark contrast with Vietnam (another single party country) where you see ads everywhere on the roads and maybe 2–3% of them are communist propaganda.&lt;/p&gt;

&lt;p&gt;At least at the visual layer, communism is somehow pushed in the background. I know for sure they have limited free speech, they have to obey more rules than in the West, but I do not see any visual enforcement on the outside. At the social fabric level, this is enforced, but you don’t have first contact with it if you just travel.&lt;/p&gt;

&lt;p&gt;My point is that purely from an economic point of view, politics seems not to stand in the way, or at least not in an exaggerated, dictatorial way as it seems from the outside. Of course, there may be businesses benefitting from government favors, but isn’t the United States owning some shares in Intel? How is that different?&lt;/p&gt;

&lt;h2&gt;
  
  
  The Takeaway: It Already Happened
&lt;/h2&gt;

&lt;p&gt;China is such a big country that a star analogy may be useful. The light from a very distant star may take years to reach us. Similarly, the impact of the progress that China has already made may take a few years to materialize in the West.&lt;/p&gt;

&lt;p&gt;They already have everything in place: great infrastructure, functioning middle class, top tier intelligence. On top of that, the West unknowingly added a few obstacles that made them even more effective.&lt;/p&gt;

&lt;p&gt;From a benchmarking perspective, the open source models from China are 6–9 months behind the SOTA Western models. But seeing things on the ground, I’m now sure the next generation that is trained by Chinese neolabs is already ahead.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>deepseek</category>
      <category>travel</category>
    </item>
    <item>
      <title>AI Can Follow You Everywhere Now. Unless You Outsmart It</title>
      <dc:creator>Dragos Roua</dc:creator>
      <pubDate>Thu, 13 Aug 2026 06:36:50 +0000</pubDate>
      <link>https://dev.to/dragos_roua/ai-can-follow-you-everywhere-now-unless-you-outsmart-it-5bgh</link>
      <guid>https://dev.to/dragos_roua/ai-can-follow-you-everywhere-now-unless-you-outsmart-it-5bgh</guid>
      <description>&lt;p&gt;Recently, Anthropic announced that all text produced by its models will be watermarked, making your entire interaction not only recognizable as coming from Claude, but also traceable to your identity. According to their press release, the watermarking will survive any copy and paste, and it will be “invisible” to the user.&lt;/p&gt;

&lt;p&gt;Two things worth mentioning here.&lt;/p&gt;

&lt;p&gt;First, this is a regulatory measure, it’s a EU law that is enforced starting August 2. FWIW, Anthropic did more than the law asked for, but let’s leave this for another time. For now, just understand that this is coming from the top to the bottom, it’s not one AI player going rogue. Everybody will follow suit (Google already did it, even open sourced their SynthId SDK for this in 2024).&lt;/p&gt;

&lt;p&gt;Second, the so called watermark is a steganography technique used to hide a message in plain sight, by encrypting token choices. In other words, Anthropic recognizes the text because it choses the next token according to a proprietary algorithm.&lt;/p&gt;

&lt;p&gt;So, the output of the model is now permanently “styled” in such a way that it will always be traced back to the model.&lt;/p&gt;

&lt;p&gt;Unless the text changes, that is. If you edit the output in a very meaningful way, the token choices are now broken, and the text is “free floating”.&lt;/p&gt;

&lt;h2&gt;
  
  
  Proprietary Models Are Locking You In, Open Source Models Not So Much
&lt;/h2&gt;

&lt;p&gt;If you still use proprietary models for text generation (docs, blog posts, emails, website copy, etc) keep in mind that the model you used will follow you everywhere. Something you wrote today will still be tracked back to you 5 years from now.&lt;/p&gt;

&lt;p&gt;In a (not so) dystopian scenario, by proving that they were part in the generation, AI neolabs can even claim ownership and ask you money for it (even though you already paid for the generation itself). This is not happening yet, to be clear, but nobody says it won’t, either.&lt;/p&gt;

&lt;p&gt;There are two ways out of this permanent white surveillance:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;always edit your texts thoroughly – this will break the algorithm and make it unrecognizable&lt;/li&gt;
&lt;li&gt;use open source models which are clearly not watermarking their output&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now, for the first part, the editing. I think it’s worth mentioning that the longer you do this, the more confused AI neolabs will be about who is the text generator, and, in time, it will rend the entire watermarking thing obsolete. No one will rely on something with so many false positives.&lt;/p&gt;

&lt;p&gt;As for the second, part, we may be facing a very urgent choice: collectively train models using distributed software, to build clean, open source AI. Everything will be out in the open and we will know the number of params, the training patterns and whether or not is there any watermarking or other stupid surveillance going on.&lt;/p&gt;

&lt;p&gt;It won’t be easy, I reckon. But we already did this with money: we started to mine Bitcoin 16 years ago, and look how far we’ve come.&lt;/p&gt;

&lt;p&gt;The clock is ticking.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>watermarking</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
