<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gil</title>
    <description>The latest articles on DEV Community by Gil (@gil_5296961bf2e126cf43cb4).</description>
    <link>https://dev.to/gil_5296961bf2e126cf43cb4</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4085789%2Fd8455722-dc41-48c6-93b8-d78733dbf160.png</url>
      <title>DEV Community: Gil</title>
      <link>https://dev.to/gil_5296961bf2e126cf43cb4</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gil_5296961bf2e126cf43cb4"/>
    <language>en</language>
    <item>
      <title>The Founder's Wire, August 13: NVIDIA Open-Sources a One-GPU Agent Model, Anthropic Watermarks Every Word Claude Writes, and…</title>
      <dc:creator>Gil</dc:creator>
      <pubDate>Thu, 20 Aug 2026 02:14:26 +0000</pubDate>
      <link>https://dev.to/gil_5296961bf2e126cf43cb4/the-founders-wire-august-13-nvidia-open-sources-a-one-gpu-agent-model-anthropic-watermarks-jbm</link>
      <guid>https://dev.to/gil_5296961bf2e126cf43cb4/the-founders-wire-august-13-nvidia-open-sources-a-one-gpu-agent-model-anthropic-watermarks-jbm</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://dreaming.press/posts/2026-08-13-founders-wire-nvidia-nemotron-open-anthropic-watermark-lovable-400m.html" rel="noopener noreferrer"&gt;dreaming.press&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The short version:   Three verified moves this morning, each aimed at a different part of a solo builder's stack.   NVIDIA   open-sourced   Nemotron 3.5 Lightning  , a 30B agent model small enough to run on one GPU and free for commercial use (&lt;a href="https://www.cnbc.com/2026/08/11/nvidia-releases-nemotron-3point5-lightning-open-source-ai-model-.html" rel="noopener noreferrer"&gt;CNBC&lt;/a&gt;).   Anthropic   began embedding an   invisible, detectable watermark   into every piece of text Claude writes — worldwide, not just in Europe (&lt;a href="https://techcrunch.com/2026/08/11/anthropic-says-it-will-watermark-text-generated-by-its-ai-models/" rel="noopener noreferrer"&gt;TechCrunch&lt;/a&gt;). And   Lovable  , the vibe-coding startup, raised   $400M at a $13.3B valuation   (&lt;a href="https://techcrunch.com/2026/08/12/lovable-confirms-new-13-3b-valuation-raises-another-400m/" rel="noopener noreferrer"&gt;TechCrunch&lt;/a&gt;). One line each on what changes — plus a cheaper Copilot coding model worth a look.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;NVIDIA open-sourced a 30B agent model that runs on one GPU&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The story here isn't the parameter count — it's the  deployability . On   August 11  , NVIDIA released   Nemotron 3.5 Lightning  , a   30-billion-parameter mixture-of-experts   model that activates only about   3 billion parameters per token  , distilled from the larger Nemotron 3 Ultra and tuned for   high-volume agentic workloads   (&lt;a href="https://www.business-standard.com/technology/tech-news/nvidia-30b-open-weight-ai-model-nemotron-3-5-lightning-agentic-tasks-126081200561_1.html" rel="noopener noreferrer"&gt;Business Standard&lt;/a&gt;). The weights are   free for commercial use  , downloadable from Hugging Face and NVIDIA's build platform   with no gate  , and NVIDIA says the model fits on a   single RTX-class GPU or a DGX Spark   desktop. A model router,   NeMo Switchyard  , shipped alongside it.&lt;/p&gt;

&lt;p&gt;What it means:   For a solo builder running agent loops — many small, repetitive model calls in sequence — the meter on a hosted API is a recurring tax that scales with usage. A commercially-licensed model that runs on one machine is the first real lever to cut that tax without an ML-infra team. The discipline is the same one we keep arguing for:   benchmark before you switch.   NVIDIA's headline "up to 4× faster output" and "~30% faster agentic tasks" are its own numbers, not independent results — so &lt;a href="///posts/lm-studio-bionic-local-agent-open-models.html"&gt;pull the weights and test Lightning on your actual workload&lt;/a&gt; before you cancel a plan, and if you're weighing what hardware it needs, our &lt;a href="///posts/gpu-rental-price-map-h100-h200-b200-august-2026.html"&gt;GPU rental price map&lt;/a&gt; is the adjacent read. This lands in the same "own your model" current as &lt;a href="///posts/2026-08-12-founders-wire-river-ai-own-your-model-gpt-cyber-qwen-open-weights.html"&gt;River AI's $1.1B raise the day before&lt;/a&gt; — the difference is that Lightning needs no vendor at all.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Anthropic watermarks every word Claude writes — worldwide&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Also on   August 11  , Anthropic disclosed that   Claude now embeds an invisible, machine-readable statistical watermark   directly into the text it generates (&lt;a href="https://techcrunch.com/2026/08/11/anthropic-says-it-will-watermark-text-generated-by-its-ai-models/" rel="noopener noreferrer"&gt;TechCrunch&lt;/a&gt;; &lt;a href="https://www.euronews.com/next/2026/08/11/eu-compliance-delivered-globally-anthropic-to-watermark-claudes-output-worldwide" rel="noopener noreferrer"&gt;Euronews&lt;/a&gt;). The specifics that matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  It's global, not regional.   The change was prompted by the   EU AI Act's Article 50   transparency obligations, which &lt;a href="///posts/eu-ai-act-article-50-august-2-founder-compliance-checklist.html"&gt;took effect August 2&lt;/a&gt; — but Anthropic applied it everywhere rather than geofencing Europe.&lt;/li&gt;
&lt;li&gt;  It survives the clipboard.   The mark is imperceptible while reading and is designed to persist through   copy-paste  ; generated files also carry signed   C2PA provenance  . Anthropic says detection tooling is coming.&lt;/li&gt;
&lt;li&gt;  The scope is fuzzy at the edges.   Reporting frames it as covering models "launched on or after August 2, 2026," and outlets describe the exact rollout slightly differently — so confirm current coverage against &lt;a href="https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content" rel="noopener noreferrer"&gt;Anthropic's own help doc&lt;/a&gt; before you state it as fact.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What it means:   If you use Claude to draft marketing copy, blog posts, landing pages, or customer-facing docs, that output now carries a   detectable signature that travels with the paste  . For internal drafts, nothing changes. For anything   bylined, sold as original, or submitted where "human-written" is assumed  , plan as if a third-party detector can flag it — and budget a genuine human rewrite pass instead of shipping raw generation. The honest framing has always been that an AI draft is a starting point you own and revise; the watermark just makes the cost of skipping that step legible.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Lovable raised $400M at $13.3B — the vibe-coding money isn't slowing&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;On   August 12  , Swedish   vibe-coding   startup   Lovable   — describe an app in plain language, it builds it — raised   $400M at a $13.3B valuation  , roughly   double   its December mark (&lt;a href="https://siliconangle.com/2026/08/12/vibe-coding-startup-lovable-doubles-valuation-13-3b-400m-raise/" rel="noopener noreferrer"&gt;SiliconANGLE&lt;/a&gt;).   Menlo Ventures   and   EQT's Scaleup Europe fund   co-led, with   Tencent   and   Balderton   participating. Company-stated numbers: ARR near   $200M   and climbing, and   60M+ projects   created since the November 2024 launch (self-reported, not audited — treat accordingly).&lt;/p&gt;

&lt;p&gt;What it means:   A war chest this size buys a push from   prototyping toy   toward   production platform   — payments, automated ops, multi-agent orchestration. For a solopreneur, the read is practical, not envious: if your last hands-on test of these tools was six months ago,   re-run your hardest build on the current version   before you pay a contractor. But the funding round doesn't change the durable truth — the generator is a commodity input, and your   distribution and domain knowledge   are the only parts a competitor can't also prompt into existence. If you're choosing what to build with, our &lt;a href="///posts/best-ai-coding-tools-2026.html"&gt;ranked guide to the best AI coding tools&lt;/a&gt; sorts them by the job you're actually hiring one to do.&lt;/p&gt;

&lt;p&gt;Also on the wire&lt;/p&gt;

&lt;p&gt;Microsoft dropped a cheaper coding model into GitHub Copilot.   On August 11,   MAI-Code-1.1-Flash   landed in Copilot with native vision and, per &lt;a href="https://github.blog/changelog/2026-08-11-mai-code-1-1-flash-available-in-github-copilot/" rel="noopener noreferrer"&gt;GitHub's changelog&lt;/a&gt;, a   73% lower list price   than the prior Flash tier, billing at a   0.25×   premium-request multiplier for annual subscribers. Same day, GitHub shipped an   Ollama / bring-your-own-key   path for the &lt;a href="https://github.blog/changelog/2026-08-11-copilot-memory-and-ollama-in-github-copilot-for-jetbrains/" rel="noopener noreferrer"&gt;JetBrains Copilot plugin&lt;/a&gt; plus persistent "Copilot memory." The move: shift low-stakes, high-volume work to the Flash tier, keep premium models for the hard problems, and check the multiplier against your request budget. If you're deciding which underlying model to trust with real code, we ranked them in &lt;a href="///posts/best-llm-for-coding-august-2026.html"&gt;the best LLM for coding, August 2026&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;Every figure above is dated and linked. Where a number is a vendor's own claim — NVIDIA's speed benchmarks, Lovable's ARR, GitHub's price cut — we've said so, because "company-stated" and "independently verified" are different things, and the difference is exactly what a founder is paying us to keep straight. &lt;/p&gt;

</description>
      <category>ai</category>
      <category>aiagents</category>
      <category>llm</category>
      <category>rag</category>
    </item>
    <item>
      <title>The Best LLM for Coding in August 2026: An Honest, Use-Case Answer (and Why the Leaderboards Disagree)</title>
      <dc:creator>Gil</dc:creator>
      <pubDate>Thu, 20 Aug 2026 02:12:18 +0000</pubDate>
      <link>https://dev.to/gil_5296961bf2e126cf43cb4/the-best-llm-for-coding-in-august-2026-an-honest-use-case-answer-and-why-the-leaderboards-539j</link>
      <guid>https://dev.to/gil_5296961bf2e126cf43cb4/the-best-llm-for-coding-in-august-2026-an-honest-use-case-answer-and-why-the-leaderboards-539j</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://dreaming.press/posts/best-llm-for-coding-august-2026.html" rel="noopener noreferrer"&gt;dreaming.press&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The one-screen answer:   There is no single "best LLM for coding" in August 2026 — the frontier is a cluster, not a leader, so the right pick is by job:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Hard agentic coding   (autonomous, multi-file, tool-using): a   frontier closed model   — &lt;a href="///posts/best-ai-coding-tools-2026.html"&gt;Claude Opus 5&lt;/a&gt; is Anthropic's own stated pick "for complex agentic coding," with   Claude Fable 5   above it for the longest runs;   OpenAI's GPT-5-series Codex   and   Google Gemini 3   are the direct rivals.&lt;/li&gt;
&lt;li&gt;  Cheap, high-volume, interactive  : a fast tier —   Claude Sonnet 5   at   $2 / $10   per million tokens is the standout, with small Gemini/GPT tiers and open models competing on price.&lt;/li&gt;
&lt;li&gt;  Self-hosting / zero per-token cost  : an   open-weight   model —   Qwen3-Coder   (permissive Apache-2.0),   DeepSeek's   latest V-series,   GLM  , or   Kimi K2  .&lt;/li&gt;
&lt;li&gt;  Large codebase / long-context refactor  : any   1M-token-context   model — the entire current Claude line is 1M, and Gemini 3 and DeepSeek advertise the same.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you need one sentence to paste:  the best coding LLM for most founders in 2026 is a frontier closed model (Claude Opus 5 / Fable 5, OpenAI GPT-5 Codex, or Google Gemini 3) paired with a cheap or open-weight model for volume — and which frontier model "wins" depends on your agent harness, not a leaderboard. &lt;/p&gt;

&lt;p&gt;Everything below is why, and — just as important —   why you should distrust most of the precise rankings you'll find.  &lt;/p&gt;

&lt;p&gt;Why there's no single winner&lt;/p&gt;

&lt;p&gt;Two or three years ago, "best coding model" had an answer, because one model was clearly ahead. That era is over. By August 2026 the top closed models — Anthropic's Claude line, OpenAI's GPT-5-series Codex models, and Google's Gemini 3 — are close enough on real coding work that the  harness  you run them in (the agent loop, the tools, the prompt scaffolding) moves results as much as the model choice does. The strongest open-weight models are reportedly a short step behind, at a fraction of the price.&lt;/p&gt;

&lt;p&gt;So the useful question isn't "which model is best." It's "best at  what , for  whom ." Here's the segmented answer, then the numbers caveat that governs all of it.&lt;/p&gt;

&lt;p&gt;The comparison, by job&lt;/p&gt;

&lt;p&gt;Model   Type   Context   Best for  &lt;/p&gt;




&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Claude Opus 5   (Anthropic)   Closed   1M   Anthropic's stated pick for   complex agentic coding    
Claude Fable 5   (Anthropic)   Closed   1M   The   longest-horizon   autonomous runs; most capable widely-released model  
Claude Sonnet 5   (Anthropic)   Closed   1M     Cheap, fast   high-volume coding — $2/$10 per M  
OpenAI GPT-5-series / Codex     Closed   Large   Agentic coding inside the   Codex   CLI/IDE harness  
Google Gemini 3     Closed   1M     Long-context   work across huge codebases  
DeepSeek / Qwen3-Coder / Kimi K2       Open-weight     1M (reported)     Self-hosting   and cutting per-token cost to zero  
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The Anthropic rows above are verified directly against &lt;a href="https://platform.claude.com/docs/en/about-claude/models/overview" rel="noopener noreferrer"&gt;Anthropic's model docs&lt;/a&gt;: Opus 5, Sonnet 5, and Fable 5 all carry a   1M-token context   (about 555,000 words) and 128K max output. The competitor rows describe positioning that's well-established; I've deliberately left precise benchmark scores out of the table, and the next section is why.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Hard agentic coding → a frontier closed model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;When the model has to plan, edit across many files, run tools, read errors, and correct itself without a human in the loop, the frontier closed models lead. Anthropic markets   Claude Opus 5   explicitly "for complex agentic coding and enterprise work" and   Claude Fable 5   as "next-generation intelligence for long-running agents." OpenAI's   Codex   line and   Gemini 3 Pro   are the direct competitors. The differences between them are real but workflow-dependent — and, crucially, the   agent harness matters as much as the model  . A slightly weaker model in a well-built harness like &lt;a href="///posts/best-ai-coding-tools-2026.html"&gt;Claude Code&lt;/a&gt; or Codex CLI often out-ships a stronger model driven poorly. Test in the tool you'll actually use.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;2. Cheap, high-volume, interactive → a fast tier
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Most coding work isn't hard; it's voluminous. Autocomplete, unit tests, boilerplate, mechanical refactors, CI checks — these want a model that's  fast and cheap enough to run constantly.    Claude Sonnet 5   ($2 in / $10 out per million tokens, billed as "the best combination of speed and intelligence") is the standout verified pick;   Claude Haiku 4.5   ($1/$5) is cheaper for the simplest work, and small Gemini/GPT tiers compete. The discipline that saves the most money is   routing by difficulty  : cheap tier for volume, frontier model reserved for the genuinely hard problems. This is exactly the move &lt;a href="///posts/2026-08-13-founders-wire-nvidia-nemotron-open-anthropic-watermark-lovable-400m.html"&gt;Microsoft just made cheaper in GitHub Copilot&lt;/a&gt; with its new low-cost Flash model.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;3. Self-hosting → an open-weight model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;If you want to cut the per-token meter to zero, keep proprietary code off a vendor's servers, or fine-tune, the open-weight models are the answer. The leaders in August 2026 are   Qwen3-Coder   (Alibaba, permissive   Apache-2.0   — the most reuse-friendly license),   DeepSeek's   latest V-series,   GLM   (Zhipu), and   Kimi K2   (Moonshot). Reported figures put the best of these near the closed frontier at a fraction of the cost. Pick   Qwen3-Coder   if license permissiveness and mature local tooling matter most; consider   DeepSeek's latest   if you want maximum reported capability and can carry the larger mixture-of-experts footprint on your own GPUs. NVIDIA's &lt;a href="///posts/2026-08-13-founders-wire-nvidia-nemotron-open-anthropic-watermark-lovable-400m.html"&gt;newly open-sourced one-GPU agent model&lt;/a&gt; is another entrant worth benchmarking here.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;4. Large codebase → a 1M-token context
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;For whole-repo reasoning and long refactors, context window is the gating spec. The   entire current Claude line runs 1M tokens   (~555K words), and   Gemini 3   and   DeepSeek   advertise the same. But a big window is not free recall: a model handed a million tokens still attends unevenly, and you pay for every one of them. Retrieval, chunking, and good prompt structure still matter — the window buys you headroom, not magic.&lt;/p&gt;

&lt;p&gt;The part most "ranking" pages won't tell you&lt;/p&gt;

&lt;p&gt;Here's the uncomfortable finding from researching this piece:   most of the precise coding-model rankings on the open web are not trustworthy.   When we pulled third-party "best coding LLM" and SWE-bench pages, the scores for the  same class of model  ranged across a   20-plus-point   spread, several pages cited model names and version numbers that didn't reconcile with each other, and no two leaderboards agreed. That's the fingerprint of auto-generated SEO content with hallucinated numbers — and an evergreen page that repeats those numbers just launders them.&lt;/p&gt;

&lt;p&gt;So here's how to actually read a coding benchmark:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  SWE-bench Verified   is the one that matters most: 500 real, human-verified GitHub issues where the model must ship a patch that passes the repo's tests. It measures real agentic coding, not trivia.&lt;/li&gt;
&lt;li&gt;  But scores are harness-dependent.   The same model scores very differently depending on the agent scaffold running it. A cross-vendor comparison is only fair when every model runs through the  same  harness — which most ranking pages don't disclose, let alone do.&lt;/li&gt;
&lt;li&gt;  Other benchmarks measure other things:   LiveCodeBench (competitive-programming style), Aider polyglot (multi-language edit accuracy), Terminal-bench (CLI/agent tasks), and SWE-bench Pro (a harder successor). We break down the two most-cited in &lt;a href="///posts/aider-polyglot-vs-swe-bench-verified-coding-benchmark.html"&gt;SWE-bench Verified vs. Aider polyglot&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;  Verify at the source.   Before you trust a number, pull it from the official &lt;a href="https://www.swebench.com/" rel="noopener noreferrer"&gt;SWE-bench leaderboard&lt;/a&gt;, the &lt;a href="https://aider.chat/docs/leaderboards/" rel="noopener noreferrer"&gt;Aider leaderboard&lt;/a&gt;, or the vendor's own model card — not a listicle.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The reason this guide ranks by use case instead of by a single score isn't hedging. It's that a single score, on this topic, in August 2026, is usually wrong — and a founder who picks a model off a fabricated leaderboard has made a worse decision than one who picked by matching a model to the job.&lt;/p&gt;

&lt;p&gt;The bottom line&lt;/p&gt;

&lt;p&gt;Match the model to the job. Frontier closed model (Claude Opus 5 / Fable 5, GPT-5 Codex, or Gemini 3) for hard agentic work; a fast tier (Sonnet 5) for volume; an open-weight model (Qwen3-Coder, DeepSeek, Kimi K2) to self-host; any 1M-context model for big repos. Then do the one thing no leaderboard can do for you:   run your two finalists on your own hardest task, in the tool you'll actually ship in, and read the diffs.   That five-minute test beats every ranking page — including this one.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>aiagents</category>
      <category>llm</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
