<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Slawa</title>
    <description>The latest articles on DEV Community by Slawa (@slawanextlevels).</description>
    <link>https://dev.to/slawanextlevels</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1738080%2F2f068d9b-88b4-465e-a989-817057839d45.jpeg</url>
      <title>DEV Community: Slawa</title>
      <link>https://dev.to/slawanextlevels</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/slawanextlevels"/>
    <language>en</language>
    <item>
      <title>[Boost]</title>
      <dc:creator>Slawa</dc:creator>
      <pubDate>Mon, 05 Oct 2026 13:49:23 +0000</pubDate>
      <link>https://dev.to/slawanextlevels/-2f0k</link>
      <guid>https://dev.to/slawanextlevels/-2f0k</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/slawanextlevels/openai-dots-meta-muse-grok-bot-who-owns-your-ai-agents-computer-25ll" class="crayons-story__hidden-navigation-link"&gt;OpenAI Dots, Meta Muse, Grok Bot: Who Owns Your AI Agent's Computer?&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/slawanextlevels" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1738080%2F2f068d9b-88b4-465e-a989-817057839d45.jpeg" alt="slawanextlevels profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/slawanextlevels" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Slawa
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Slawa
                
                
              
              &lt;div id="story-author-preview-content-4787687" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/slawanextlevels" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1738080%2F2f068d9b-88b4-465e-a989-817057839d45.jpeg" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Slawa&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/slawanextlevels/openai-dots-meta-muse-grok-bot-who-owns-your-ai-agents-computer-25ll" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Oct 2&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/slawanextlevels/openai-dots-meta-muse-grok-bot-who-owns-your-ai-agents-computer-25ll" id="article-link-4787687"&gt;
          OpenAI Dots, Meta Muse, Grok Bot: Who Owns Your AI Agent's Computer?
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agents"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agents&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/security"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;security&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/architecture"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;architecture&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/slawanextlevels/openai-dots-meta-muse-grok-bot-who-owns-your-ai-agents-computer-25ll#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              2&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            8 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>agents</category>
      <category>ai</category>
      <category>openai</category>
    </item>
    <item>
      <title>State of AI, Fall 2026: Which Models Lead, Which Benchmarks Still Matter, and What Your $200 Subscription Really Buys</title>
      <dc:creator>Slawa</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:55:04 +0000</pubDate>
      <link>https://dev.to/slawanextlevels/state-of-ai-fall-2026-which-models-lead-which-benchmarks-still-matter-and-what-your-200-342e</link>
      <guid>https://dev.to/slawanextlevels/state-of-ai-fall-2026-which-models-lead-which-benchmarks-still-matter-and-what-your-200-342e</guid>
      <description>&lt;p&gt;State of ai devto body · TXT&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Anthropic (Claude Opus 5.5, Sonnet 5.5, Fable 5.1) and OpenAI (GPT-6 Astra, GPT-6.1 Sol) currently lead; open-weights models such as Kimi K3 and DeepSeek V4 trail by three to eight months. What decides your productivity is no longer the score, but the price per task and the real usage limits of your subscription.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Snapshot as of 30 September 2026.&lt;/strong&gt; The original of this article is a living post: models, benchmarks, prices and limits are re-verified every week. For the latest numbers and the dated changelog, check the &lt;a href="https://next-levels.de/en/blog/a-comparison-of-ai-models-the-state-of-ai-benchmarks-subscriptions-and-real-world-applications" rel="noopener noreferrer"&gt;up-to-date version on next-levels.de&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The benchmark table you just read in some newsletter is already out of date. Over the last four weeks, the top spots have been reshuffled five times. And yet the question that’s really burning in the developer forums is a different one. A ChatGPT Pro user writes: “Weekly limit reset yesterday, only 20% left today.” The model question has become a side issue. Your productivity is determined by the cost per task and your weekly limit; that last percentage point on the leaderboard is just noise.&lt;/p&gt;

&lt;p&gt;So this post answers four questions: which models are leading the pack, which benchmarks can still measure anything, what the $20, $100 and $200 subscriptions really include, and what developers experience with Claude Max, Codex and Kimi in their day-to-day work.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2Fae7afb46c1414b829eb7fa0f6daca3f5%3Fformat%3Dpng%26w%3D1600%26q%3D82" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2Fae7afb46c1414b829eb7fa0f6daca3f5%3Fformat%3Dpng%26w%3D1600%26q%3D82" alt="AI models compared: a timeline of top models from GPT-6 Astra to GPT-6.1 Sol" width="1600" height="900"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  At a glance
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Top performers: Claude Opus 5.5 and Sonnet 5.5 lead the independent coding and agent indices, while GPT-6 Astra leads in knowledge and interactive reasoning.&lt;/li&gt;
&lt;li&gt;SWE-bench Verified, GPQA and AIME are saturated. Terminal-Bench 4.0, ARC-AGI-3, GDPval-AA and the Artificial Analysis Index with private test sets are meaningful.&lt;/li&gt;
&lt;li&gt;The mid-range (US$2 to US$4 per million input tokens) has seen a collapse in price: GPT-6.1 Sol and Sonnet 5.5 cost 2/10, Opus 5.5 4/20. The top tier remains at 10/50.&lt;/li&gt;
&lt;li&gt;The subscription tiers differ in volume, not in model access; even entry tiers get the top models: Claude effectively reduced its weekly limits by 17% in September, while OpenAI paused Pro 20× and reopened it with half the quota.&lt;/li&gt;
&lt;li&gt;Recommendation: Opus 5.5 at medium reasoning effort as the default, top-tier models only for review and security, Kimi K3 as a cheap second track for code without personal data.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  AI models compared: who’s currently in the lead
&lt;/h2&gt;

&lt;p&gt;The real news this summer comes from the open-weights camp. Since July, four Chinese labs have released models with between one and three trillion parameters and one million tokens of context under open licences: &lt;a href="https://www.unite.ai/moonshot-opens-kimi-k3-weights-under-a-revenue-tiered-license/" rel="noopener noreferrer"&gt;Kimi K3&lt;/a&gt; from Moonshot (2.8 trillion parameters, 104 billion active, weights available since 27 July), Alibaba’s Qwen3.8-Max, DeepSeek V4-Pro under the MIT licence and, most recently, Xiaomi’s MiMo-V2.6-Pro, which, with an index score of 46, is &lt;a href="https://artificialanalysis.ai/models/mimo-v2-6-pro" rel="noopener noreferrer"&gt;the strongest open model&lt;/a&gt; in the Artificial Analysis ranking.&lt;/p&gt;

&lt;p&gt;How big is the gap? The US institute CAISI estimates it for DeepSeek V4 Pro at &lt;a href="https://www.nist.gov/news-events/news/2026/05/caisi-evaluation-deepseek-v4-pro" rel="noopener noreferrer"&gt;around eight months&lt;/a&gt; behind the US frontrunners; for GLM-5.3, it measures on cyber benchmarks &lt;a href="https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities" rel="noopener noreferrer"&gt;around four months&lt;/a&gt;. Nathan Lambert from Interconnects sees the gap narrowed by K3 &lt;a href="https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation" rel="noopener noreferrer"&gt;from six to nine months to three to five months&lt;/a&gt;. Three to eight months: That’s the head start you are buying today when you pay for a closed frontier model.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2F951e671f3953445ca89837e3797faa00%3Fformat%3Dpng%26w%3D1600%26q%3D82" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2F951e671f3953445ca89837e3797faa00%3Fformat%3Dpng%26w%3D1600%26q%3D82" alt="The gap between open-source AI models and the US leaders: GLM-5.3, Kimi K3 and DeepSeek V4 Pro in months" width="1600" height="900"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And the leaders? A duopoly. In June, Anthropic moved the “Fable/Mythos 5” class above Opus; &lt;a href="https://www.anthropic.com/claude-fable-and-mythos-5-1" rel="noopener noreferrer"&gt;Fable 5.1 and Mythos 5.1&lt;/a&gt; were launched on 1 September; Mythos remains reserved for security-cleared programmes. They were followed by &lt;a href="https://www.anthropic.com/claude-opus-5-5" rel="noopener noreferrer"&gt;Claude Opus 5.5&lt;/a&gt; at 4/20 US dollars per million tokens, which, according to Anthropic, is “on a par with Fable 5.1 for most tasks”, and Sonnet 5.5 at the unchanged price of 2/10.&lt;/p&gt;

&lt;p&gt;OpenAI has simultaneously replaced its entire range. &lt;a href="https://openai.com/index/gpt-6-astra/" rel="noopener noreferrer"&gt;GPT-6 Astra&lt;/a&gt; is the company’s largest training run (over 100,000 GPUs), with 1.05 million tokens of context and a price of 10/50 US dollars per million. Below it sit &lt;a href="https://openai.com/index/introducing-gpt-6-sol-and-luna/" rel="noopener noreferrer"&gt;GPT-6 Sol and Luna&lt;/a&gt;, at half the price of the 5.6 generation, and on 29 September, GPT-6.1 Sol, which, according to OpenAI, delivers Astra-level performance on DeepSWE at a fifth of the cost.&lt;/p&gt;

&lt;p&gt;Google has spent the year without a new Pro model. Gemini 3.5 Pro was announced at I/O and &lt;a href="https://techcrunch.com/2026/07/21/google-releases-three-new-gemini-models-but-no-3-5-pro/" rel="noopener noreferrer"&gt;never launched&lt;/a&gt;; instead, three Flash models were released in six weeks, most recently Gemini 3.8 Flash on 2 September. According to Google, Gemini 4 is set to be released &lt;a href="https://www.infoworld.com/article/4226642/google-plans-gemini-4-release-before-year-end-2.html" rel="noopener noreferrer"&gt;well before the end of the year&lt;/a&gt;. xAI releases new models monthly (Grok 4.7 on 21 September, US$2/6). Meta has effectively shelved the Llama line: since April, its only model has been the proprietary Muse Spark, currently at version 1.3.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Release&lt;/th&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;API price (1M tokens in/out)&lt;/th&gt;
&lt;th&gt;Open weights&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5.1&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;1 September 2026&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;10 $ / 50 $&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5.5&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;22 September 2026&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;4 $ / 20 $&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;28 September 2026&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;2 $ / 10 $&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;4 September 2026&lt;/td&gt;
&lt;td&gt;1.05M&lt;/td&gt;
&lt;td&gt;10 $ / 50 $&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6.1 Sol&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;29 September 2026&lt;/td&gt;
&lt;td&gt;1.05M&lt;/td&gt;
&lt;td&gt;2 $ / 10 $&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Luna&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;22 September 2026&lt;/td&gt;
&lt;td&gt;1.05M&lt;/td&gt;
&lt;td&gt;$0.10 / $0.50&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;Google&lt;/td&gt;
&lt;td&gt;2 September 2026&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;td&gt;$0.75 / $3.75 (introductory price until 31 December)&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4.7&lt;/td&gt;
&lt;td&gt;xAI&lt;/td&gt;
&lt;td&gt;21 September 2026&lt;/td&gt;
&lt;td&gt;500K&lt;/td&gt;
&lt;td&gt;$2 / $6&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;Moonshot&lt;/td&gt;
&lt;td&gt;16 July 2026 (weights 27 July)&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;3 $ / 15 $&lt;/td&gt;
&lt;td&gt;yes (revenue-based licence)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4-Pro (0813)&lt;/td&gt;
&lt;td&gt;DeepSeek&lt;/td&gt;
&lt;td&gt;13 August 2026&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;td&gt;yes (MIT)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.8-Max&lt;/td&gt;
&lt;td&gt;Alibaba&lt;/td&gt;
&lt;td&gt;3 August 2026&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;td&gt;yes (own licence; Apache 2.0 only for Qwen3.8-27B)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiMo-V2.6-Pro&lt;/td&gt;
&lt;td&gt;Xiaomi&lt;/td&gt;
&lt;td&gt;21 September 2026&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;$0.43 / $0.87&lt;/td&gt;
&lt;td&gt;yes (MIT)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Vendor list prices as of 30 September 2026; batch and cache discounts come on top. And that is precisely where the most important figure of the month lies: Anthropic has reduced the cache read prices by 75% for Fable 5.1 (to $0.25 per million) and by 60% for Opus 5.5 (to $0.20). In an agent run that re-reads the same 200,000-token context fifty times, that amounts to ten million cache tokens: with Opus 5.5, the cost is two US dollars instead of five. The list price for fresh tokens is almost irrelevant for such workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmarks: what still has room for growth and what is saturated
&lt;/h2&gt;

&lt;p&gt;If you pick models from tables, you have a problem: the benchmarks you know are obsolete. OpenAI removed SWE-bench Verified from its own reporting &lt;a href="https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/" rel="noopener noreferrer"&gt;in February&lt;/a&gt; because 59% of the tested tasks were flawed and top models were able to reproduce the human reference solution verbatim. GPQA Diamond was removed from the index by Artificial Analysis in September because all top models “simply solve” it. On AIME 2026, the top 3 sit between 97% and 99%. On FrontierMath, Epoch AI had to &lt;a href="https://epoch.ai/benchmarks/frontiermath-tier-4-v2" rel="noopener noreferrer"&gt;re-release the benchmark&lt;/a&gt; in June after 42% of the tasks contained errors.&lt;/p&gt;

&lt;p&gt;What counts instead are interactive, economically grounded tests that are hard to contaminate. &lt;a href="https://www.tbench.ai/news/terminal-bench-4-0" rel="noopener noreferrer"&gt;Terminal-Bench 4.0&lt;/a&gt; comprises 66 curated tasks; eight were removed (saturated, publicly solved or with quality issues), 19 were revised. ARC-AGI-3 is interactive; the model must learn the rules of an environment through trial and error, so there is no solution that could be found in the training dataset. GDPval-AA measures real-world professional tasks from 44 occupations in blind pair comparisons. And Artificial Analysis weights private, never-before-published test sets at 45% in the &lt;a href="https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3" rel="noopener noreferrer"&gt;Index v4.3&lt;/a&gt;, compared with 40% in v4.2 and 20% previously.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;What it measures&lt;/th&gt;
&lt;th&gt;Top 3 (as at 30 September)&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://artificialanalysis.ai/leaderboards/models" rel="noopener noreferrer"&gt;AA Intelligence Index v4.3&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;10 evaluations, 45% private sets, independent&lt;/td&gt;
&lt;td&gt;Opus 5.5 (58), Opus 5.5 xhigh / Sonnet 5.5 (56), Fable 5.1 / GPT-6 Astra (53)&lt;/td&gt;
&lt;td&gt;current&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://artificialanalysis.ai/evaluations/terminalbench-4-0" rel="noopener noreferrer"&gt;Terminal Bench 4.0&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;agent-based terminal tasks, independent&lt;/td&gt;
&lt;td&gt;Sonnet 5.5 (63.6%), Opus 5.5 max (59.6%), Opus 5.5 xhigh (59.6%)&lt;/td&gt;
&lt;td&gt;current&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arena.ai/leaderboard/code/webdev" rel="noopener noreferrer"&gt;Arena WebDev&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Web apps in a blind comparison, 795,000 votes&lt;/td&gt;
&lt;td&gt;Opus 5.5 (1,827), GPT-6 Astra (1,792), Fable 5.1 (1,751)&lt;/td&gt;
&lt;td&gt;current&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://artificialanalysis.ai/evaluations/gdpval-aa" rel="noopener noreferrer"&gt;GDPval-AA v2.1&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;220 professional tasks, 44 professions, Elo, independent&lt;/td&gt;
&lt;td&gt;Opus 5.5 (1846), Sonnet 5.5 (1844), Opus 5.5 xhigh (1820)&lt;/td&gt;
&lt;td&gt;current&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://labs.scale.com/leaderboard/humanitys_last_exam_text_only" rel="noopener noreferrer"&gt;Humanity’s Last Exam&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Expert knowledge, no tools, independent (Scale)&lt;/td&gt;
&lt;td&gt;GPT-6 Astra (54.2%), Gemini 3.1 Pro (47.3%), Fable 5.1 (46.8%)&lt;/td&gt;
&lt;td&gt;current&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arcprize.org/results/openai-gpt-6-astra" rel="noopener noreferrer"&gt;ARC-AGI-3&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Interactive reasoning, independent (ARC Prize)&lt;/td&gt;
&lt;td&gt;GPT-6 Astra 62.7% on the standard harness (99.9% on the provider’s harness), Opus 5 30.2%; Opus 5.5 not yet measured&lt;/td&gt;
&lt;td&gt;current, harness dispute&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SWE-bench Verified&lt;/td&gt;
&lt;td&gt;Coding tickets&lt;/td&gt;
&lt;td&gt;Top score 96%; ranks 11 to 15 within half a point at 80%&lt;/td&gt;
&lt;td&gt;Saturated; discontinued by OpenAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPQA Diamond&lt;/td&gt;
&lt;td&gt;Natural sciences&lt;/td&gt;
&lt;td&gt;Provider figures: Astra 96.0%, GPT-5.6 Sol 94.6%&lt;/td&gt;
&lt;td&gt;saturated, removed from the AA Index&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AIME 2026&lt;/td&gt;
&lt;td&gt;Math olympiad&lt;/td&gt;
&lt;td&gt;Top 3 within 2.1 points at 97–99%&lt;/td&gt;
&lt;td&gt;saturated&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;“xhigh”/“max” denote the reasoning effort used in the measurement. “Harness” refers to the test environment: which tools the model is given, how many attempts, and what compute budget.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The figures from providers’ blogs are almost universally higher than the independent measurements. Anthropic reports 70.6% for Sonnet 5.5 on Terminal Bench 4.0, while Artificial Analysis measures 63.6%. OpenAI reports 99.9% for Astra on ARC-AGI-3, while the standard ARC Prize harness yields 62.7%, with computing costs of around 26,000 US dollars for the test run. Neither figure is a lie; they are different harnesses and reasoning levels. They are just not comparable. In short: trust the figure measured by someone who isn’t selling the model.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2F3dfae41fc6ec41eb938929b3b1910000%3Fformat%3Dpng%26w%3D1600%26q%3D82" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2F3dfae41fc6ec41eb938929b3b1910000%3Fformat%3Dpng%26w%3D1600%26q%3D82" alt="Artificial Analysis Intelligence Index v4.3: Opus 5.5, Sonnet 5.5, Fable 5.1, GPT-6 Astra and MiMo-V2.6-Pro" width="1600" height="900"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What has shifted in recent weeks: Since the release of Opus and Sonnet 5.5, Anthropic has been leading the way in agentic coding tests and professional task indices, while OpenAI leads in knowledge, maths and ARC. Both providers are now promoting the same metric. Anthropic &lt;a href="https://www.anthropic.com/claude-opus-5-5" rel="noopener noreferrer"&gt;calculates&lt;/a&gt; that Opus 5.5 at medium reasoning effort beats Astra at maximum effort, at roughly a fifth of the cost per task. OpenAI countered on 29 September with GPT-6.1 Sol and exactly the same formula.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI subscriptions compared: what you get for $20, $100 and $200
&lt;/h2&gt;

&lt;p&gt;Comparing AI subscriptions is simpler than it looks. All providers have settled on three tiers (around 20, 100 and 200 US dollars), and all control the price via limits. Euro prices are for Germany and include VAT where the provider states them:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Entry level&lt;/th&gt;
&lt;th&gt;Mid-range&lt;/th&gt;
&lt;th&gt;Top tier&lt;/th&gt;
&lt;th&gt;What distinguishes the tiers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://claude.com/pricing" rel="noopener noreferrer"&gt;Claude&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Pro €21.42 (Opus, Sonnet, Claude Code)&lt;/td&gt;
&lt;td&gt;Max 5× €107.10&lt;/td&gt;
&lt;td&gt;Max 20× €214.20&lt;/td&gt;
&lt;td&gt;5-hour window plus weekly limit; Fable 5.1 available only on Max and Team Premium, up to 50% of the weekly limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://learn.chatgpt.com/docs/pricing" rel="noopener noreferrer"&gt;ChatGPT&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Plus €23 (15 to 150 Sol messages per 5 hours)&lt;/td&gt;
&lt;td&gt;Pro 5× €103&lt;/td&gt;
&lt;td&gt;Pro 20× €229 (new sign-ups since 29 September with halved quota)&lt;/td&gt;
&lt;td&gt;Weekly quotas at all tiers; Astra Ultrafast only on Pro for $500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://gemini.google/subscriptions/" rel="noopener noreferrer"&gt;Gemini&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;AI Pro $19.99 (3.1 Pro, Antigravity entry-level)&lt;/td&gt;
&lt;td&gt;AI Ultra $99.99&lt;/td&gt;
&lt;td&gt;AI Ultra $199.99&lt;/td&gt;
&lt;td&gt;Gemini CLI disabled for subscriptions since 18 June; replaced by the closed Antigravity CLI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi Code&lt;/td&gt;
&lt;td&gt;Andante ¥49 (K2.7 Code)&lt;/td&gt;
&lt;td&gt;Moderato ¥99 / Allegretto ¥199 (K3, 1M)&lt;/td&gt;
&lt;td&gt;Allegro ¥699&lt;/td&gt;
&lt;td&gt;Monthly pool plus weekly allowance plus 5-hour rate; processed by Moonshot AI Pte. Ltd. in Singapore&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://devgent.org/en/cursor-pricing-guide-en/" rel="noopener noreferrer"&gt;Cursor&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Pro $20 ($20 credit)&lt;/td&gt;
&lt;td&gt;Pro+ $60 ($70 credit)&lt;/td&gt;
&lt;td&gt;Ultra $200 ($400 credit)&lt;/td&gt;
&lt;td&gt;Credits instead of requests since 2025; Auto mode draws on credits too&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/" rel="noopener noreferrer"&gt;GitHub Copilot&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Pro $10 ($10 credit)&lt;/td&gt;
&lt;td&gt;Pro+ $39&lt;/td&gt;
&lt;td&gt;Business $19 / Enterprise $39 per user&lt;/td&gt;
&lt;td&gt;From 1 June 2026: usage-based pricing according to API rates&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Euro prices for Claude according to &lt;a href="https://www.ssdnodes.com/learn/lang/de/claude-pro-and-max-in-germany-what-you-pay" rel="noopener noreferrer"&gt;ssdnodes&lt;/a&gt;, for ChatGPT according to &lt;a href="https://heise.de/-11251779" rel="noopener noreferrer"&gt;heise&lt;/a&gt;; Google and Kimi do not quote euro prices for Germany. For teams, the number of user seats is relevant: Claude Team Premium with Claude Code costs 100 US dollars per seat on an annual subscription, ChatGPT Business Premium also 100; with a VAT number, Claude Max 20× comes to 180 euros net.&lt;/p&gt;

&lt;p&gt;The difference between the tiers is not a model upgrade. With Claude, even the Pro tier at 21 euros gets Opus 5.5 and Claude Code. With OpenAI, the Plus tier gets GPT-6 Astra, but only between five and 45 Astra messages every five hours. You pay by volume. So the more interesting question is: how fast will I hit the limit?&lt;/p&gt;

&lt;h2&gt;
  
  
  In practice: what developers really experience with Claude Max, Codex and Kimi
&lt;/h2&gt;

&lt;p&gt;The pricing page says “5× Pro”. In recent weeks, both big providers have revised downward what that actually means.&lt;/p&gt;

&lt;p&gt;On 14 September, Anthropic replaced the “+50% weekly limit” promotion, launched in May, with a permanent “+25% over the old baseline”. That sounds like a bonus, but compared with the last four months it is &lt;a href="https://www.bleepingcomputer.com/news/artificial-intelligence/anthropic-is-cutting-claude-codes-current-weekly-limits-by-17-percent/" rel="noopener noreferrer"&gt;a 17% cut&lt;/a&gt;, which Anthropic, after criticism on Hacker News, &lt;a href="https://www.mindstudio.ai/blog/claude-code-weekly-rate-limit-changes" rel="noopener noreferrer"&gt;confirmed&lt;/a&gt;. The 5-hour windows, which were doubled in May, remain in place.&lt;/p&gt;

&lt;p&gt;At the same time, a &lt;a href="https://www.engadget.com/2194626/anthropic-hit-with-lawsuit-over-its-claude-max-usage-limits/" rel="noopener noreferrer"&gt;class action lawsuit&lt;/a&gt; has been running in the US since June, alleging that Max 20× actually delivers &lt;a href="https://runtimewire.com/article/anthropic-claude-max-20x-usage-lawsuit" rel="noopener noreferrer"&gt;six to eight times&lt;/a&gt; the usage of Pro, not twenty times. A Max 20× user who has been logging his usage for months &lt;a href="https://christopheralarcon.com/blog/claude-pro-vs-claude-max" rel="noopener noreferrer"&gt;reports&lt;/a&gt; 20% of his weekly quota used on the first day after the reset. His verdict: worth it for $200, “but I wouldn’t bet on it a year from now”.&lt;/p&gt;

&lt;p&gt;An effect that hardly anyone takes into account: Fable 5.1 is capped at 50% of the weekly limit on Max and costs 2.5 times as much per token as Opus 5.5. According to &lt;a href="https://theaicareerlab.com/blog/claude-opus-5-5-vs-fable-5-1-for-professionals" rel="noopener noreferrer"&gt;a widely shared estimate&lt;/a&gt; by Theo Browne, making Opus 5.5 your default gives you roughly four times the effective quota.&lt;/p&gt;

&lt;p&gt;OpenAI tackled the problem from the other side. On 10 September, new registrations for Pro 20× were paused because, according to Codex head Tibo Sottiaux, this segment places the greatest load on the systems; on 29 September, sign-ups reopened with &lt;a href="https://www.developersdigest.tech/blog/codex-usage-limits-pricing-2026" rel="noopener noreferrer"&gt;halved quota&lt;/a&gt;; existing customers are spared until 29 October. The quote from the intro is from this phase: “Weekly limit reset yesterday, only 20% left today”, writes a Pro-5× user &lt;a href="https://community.openai.com/t/is-there-any-indication-when-20x-will-be-available-again-for-subscription/1400356" rel="noopener noreferrer"&gt;on the OpenAI forum&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Common practice is to split the work: Codex for the bulk of small changes and cloud parallelisation, Claude Code for long, continuous sessions. In an &lt;a href="https://dev.to/sangrokjung/claude-code-vs-codex-2026-what-500-reddit-developers-really-think-31pb"&gt;analysis of 500 Reddit comments&lt;/a&gt;, 65% preferred Codex, mainly because of the limits, while Claude Code won eight out of twelve blind tests, which is not a contradiction.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2Fd8c39d78d7c544d8aac195119b66edc3%3Fformat%3Dpng%26w%3D1600%26q%3D82" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2Fd8c39d78d7c544d8aac195119b66edc3%3Fformat%3Dpng%26w%3D1600%26q%3D82" alt="Subscription reality: Claude Code weekly limit down 17%, ChatGPT Pro 20x halved, class action 6- to 8-fold" width="1600" height="900"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Kimi is the third track, and it deserves to be taken seriously. Moonshot’s Kimi K2.7 Code runs as a drop-in inside Claude Code (point &lt;code&gt;ANTHROPIC_BASE_URL&lt;/code&gt; at &lt;code&gt;https://api.moonshot.ai/anthropic&lt;/code&gt;, done) and, at 0.95/4 US dollars per million tokens, costs around a quarter to a fifth of Opus 5.5.&lt;/p&gt;

&lt;p&gt;A &lt;a href="https://www.totalum.app/blog/kimi-k2-7-code-vs-claude-2026" rel="noopener noreferrer"&gt;sample calculation&lt;/a&gt; against the then-current Opus 4.8 came to 0.69 US dollars instead of 4.00 US dollars for a 50-turn agent run. A practitioner who ran K2.7 in Claude Code for about two weeks &lt;a href="https://www.kunalganglani.com/blog/kimi-k2-7-claude-code-alternative" rel="noopener noreferrer"&gt;describes&lt;/a&gt; it as disciplined in following instructions, with no scope creep, but weaker at extended thinking and with less community know-how about the best prompts.&lt;/p&gt;

&lt;p&gt;The catch lies elsewhere. The international Kimi platform is operated by &lt;a href="https://platform.kimi.ai/docs/agreement/userprivacy" rel="noopener noreferrer"&gt;Moonshot AI Pte. Ltd. in Singapore&lt;/a&gt;, where the data is processed; a GDPR-compliant data processing agreement (DPA) is not publicly available. Singapore is a third country without an adequacy decision from the EU, and the parent company is based in Beijing. For a European company with customer data in the repo, that is a deal-breaker as long as there is no EU entity offering a DPA.&lt;/p&gt;

&lt;p&gt;For code without personal data, it is a cheap second track. Ramp reports that &lt;a href="https://ramp.com/data/ai-index-august-2026" rel="noopener noreferrer"&gt;6.1% of US companies paying for AI&lt;/a&gt; are now using platforms that provide open-source or Chinese models.&lt;/p&gt;

&lt;p&gt;Two numbers show how far enterprise reality is from the forums. According to &lt;a href="https://www.kucoin.com/blog/anthropic-vs-openai-ramp-data-shows-claude-leading-us-enterprise-ai-market-share-at-43-8" rel="noopener noreferrer"&gt;Menlo Ventures&lt;/a&gt;, around 54% of enterprise spending on AI coding in 2025 went to Anthropic and 21% to OpenAI; in the &lt;a href="https://ramp.com/data/ai-index-august-2026" rel="noopener noreferrer"&gt;Ramp AI Index&lt;/a&gt; from August, 43.5% of US companies pay Anthropic and around 40% pay OpenAI.&lt;/p&gt;

&lt;p&gt;Usage numbers tell a different story than budgets. The &lt;a href="https://byteiota.com/stack-overflow-dev-survey-2026-ai-at-84-trust-at-3/" rel="noopener noreferrer"&gt;Stack Overflow survey 2026&lt;/a&gt;, with over 49,000 responses, shows that 84% use AI tools, 29% trust their accuracy (3% “strongly”), and 66% say the code is “almost right, but not quite”. Among the tools, Copilot leads with 68% of AI users, Cursor is at 18%, and Claude Code at just under 10%.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for your setup
&lt;/h2&gt;

&lt;p&gt;I think it’s a mistake to rely on a single model for a team. The reason is economic: the cost per task depends more on the reasoning level than on the model. Opus 5.5 scores an identical 59.6% on Terminal Bench 4.0 at “max” and “xhigh”; the higher level just burns more tokens. My recommendation as of today:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; a mid-range model at medium reasoning effort. Today, that means Opus 5.5 medium or GPT-6.1 Sol; both deliver top results at a fifth of the cost. &lt;strong&gt;Review:&lt;/strong&gt; frontier models, i.e. each provider’s most expensive tier (Fable 5.1, Astra), are reserved for cases where a mistake gets expensive: security reviews, architecture decisions, the last pair of eyes before the merge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bulk:&lt;/strong&gt; for high-volume work with clear instructions (generating tests, migration scripts, documentation), a third, low-cost track pays off, provided data protection and DPAs are sorted. That could be Luna, Gemini Flash or an open-weights model on European infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2Fafef177c8b53415a90b3cb9ed849517d%3Fformat%3Dpng%26w%3D1600%26q%3D82" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2Fafef177c8b53415a90b3cb9ed849517d%3Fformat%3Dpng%26w%3D1600%26q%3D82" alt="Comparison of three-track setups for AI models: Default, Review and Mass within a Harness" width="1600" height="900"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;As I described in the post &lt;a href="https://next-levels.de/en/blog/agent-harness-model-why-the-model-question-is-the-wrong-debate" rel="noopener noreferrer"&gt;“Agent = Harness + Model”&lt;/a&gt;, the harness (context, memory, tools and evals) determines whether an agent is productive. In a properly built setup, swapping the model is a single config line; &lt;code&gt;ANTHROPIC_BASE_URL&lt;/code&gt; above is an example.&lt;/p&gt;

&lt;p&gt;If switching from Opus 5 to 5.5 takes days, the dependency is in the wrong place. And anyone sending customer data through a model outside the EU because it’s five times cheaper should read my post on the &lt;a href="https://next-levels.de/en/blog/digital-sovereignty-the-european-ai-stack-for-small-and-medium-sized-enterprises" rel="noopener noreferrer"&gt;European AI stack&lt;/a&gt; first.&lt;/p&gt;

&lt;p&gt;When deciding on a subscription: a team of five developers is better off with Claude Team Premium (US$100 per seat, including Claude Code) or ChatGPT Business Premium than with five individual Max subscriptions. SSO and audit logs are one reason.&lt;/p&gt;

&lt;p&gt;The other: on the Team and Business plans, chats are not used for training by default and belong to the company account; with private subscriptions, they are linked to the employee and move with them. If you stick with Cursor or Copilot: both now bill at API rates, and the $20 plan is really a credit balance.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I’m watching next
&lt;/h2&gt;

&lt;p&gt;Four open threads that will shape the coming weeks.&lt;/p&gt;

&lt;p&gt;Anthropic announced Haiku 5.5 on 22 September, for release “in the coming weeks”, but it isn’t here yet. At OpenAI, the grace period for existing Pro-20× customers ends on 29 October; after that, the halved quota will apply to everyone. Gemini 4 is expected to launch well before the end of the year and would be Google’s first new Pro model since February. And the class action lawsuit against Anthropic’s Max limits is pending; a judgement or settlement would directly affect the subscription table above.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Which AI model is currently the best?&lt;/strong&gt; For agentic coding and professional tasks, Claude Opus 5.5 leads the independent indices (Artificial Analysis, Terminal-Bench 4.0, GDPval-AA); for knowledge and interactive reasoning, GPT-6 Astra (Humanity’s Last Exam, ARC-AGI-3). The performance gap is smaller than the price gap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Claude Max 20× still worth it?&lt;/strong&gt; If you work in Claude Code for several hours a day and set Opus 5.5 as your default instead of Fable, then yes. Expect to pay €214 incl. VAT in Germany and bear in mind that a single large agent run will make a noticeable dent in your weekly limit. For teams, Team Premium at 100 US dollars per seat is usually the better option.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I use Kimi K3 for business in the EU?&lt;/strong&gt; Technically yes, via API or as a drop-in in Claude Code. Legally, it depends on what data you send: processing is carried out by Moonshot AI Pte. Ltd. in Singapore, a third country without an EU adequacy decision, and a data processing agreement is not publicly available. For code without personal data, it’s a cost-effective alternative; for customer data, it isn’t.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why do vendor blogs show different benchmark numbers than this post?&lt;/strong&gt; Because vendors measure with their own harness, at the highest reasoning level and sometimes with multiple attempts. Independent evaluators such as Artificial Analysis, ARC Prize or Scale use a fixed harness and budget. This post prefers the independent numbers and labels vendor numbers as such.&lt;/p&gt;

&lt;h2&gt;
  
  
  Changelog
&lt;/h2&gt;

&lt;p&gt;The original post logs every change to models, prices, limits and top-3 rankings with a date. This dev.to copy is a snapshot; the &lt;a href="https://next-levels.de/en/blog/a-comparison-of-ai-models-the-state-of-ai-benchmarks-subscriptions-and-real-world-applications#changelog" rel="noopener noreferrer"&gt;live changelog&lt;/a&gt; has everything since.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;30 September 2026&lt;/strong&gt; — First published. Status: Opus 5.5, Sonnet 5.5, GPT-6 Astra, GPT-6.1 Sol, Gemini 3.8 Flash, Grok 4.7, Kimi K3, MiMo-V2.6-Pro. Claude weekly limit −17% since 14 September; ChatGPT Pro 20× has been open again (halved) since 29 September.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;I’m Slawa, CEO of &lt;a href="https://next-levels.de/en" rel="noopener noreferrer"&gt;Next Levels&lt;/a&gt;, a German agency building Shopware commerce, apps and AI agents for mid-sized companies. This article was originally published on &lt;a href="https://next-levels.de/en/blog/a-comparison-of-ai-models-the-state-of-ai-benchmarks-subscriptions-and-real-world-applications" rel="noopener noreferrer"&gt;next-levels.de&lt;/a&gt;, where it is re-verified every week. If a number here looks stale, the original has the current one.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Which split are you running in your team: one model for everything, or separate tracks for default, review and bulk work? Let me know in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>OpenAI Dots, Meta Muse, Grok Bot: Who Owns Your AI Agent's Computer?</title>
      <dc:creator>Slawa</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:52:37 +0000</pubDate>
      <link>https://dev.to/slawanextlevels/openai-dots-meta-muse-grok-bot-who-owns-your-ai-agents-computer-25ll</link>
      <guid>https://dev.to/slawanextlevels/openai-dots-meta-muse-grok-bot-who-owns-your-ai-agents-computer-25ll</guid>
      <description>&lt;p&gt;&lt;em&gt;Facts as of September 30, 2026.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Do not use separate Bots as a security boundary."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sentence isn't from a pentest report. It's from xAI's own &lt;a href="https://docs.x.ai/grok-bot/approvals-security-and-privacy" rel="noopener noreferrer"&gt;Grok Bot documentation&lt;/a&gt;, for a product marketed as &lt;a href="https://x.ai/news/introducing-grok-bot" rel="noopener noreferrer"&gt;"AI teammates you can give real work to."&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Within seven weeks, three of the biggest AI vendors shipped the same kind of product: an agent that runs around the clock, logs into your tools and reports back when it's done. xAI went first with Grok Bot on August 11, Meta followed with Muse on September 8, and OpenAI launched Dots at DevDay on September 29. Most coverage compares context windows, benchmarks and avatars. If you've ever written a threat model, the interesting question is a different one: &lt;strong&gt;which computer does the agent run on, who else shares that computer, and what decides what leaves it?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The common pattern: the agent gets a computer
&lt;/h2&gt;

&lt;p&gt;Until recently, "agent" in practice meant an LLM calling tools over an API while you sat in a chat window. Close the tab, the agent is gone. All three products break that model in the same three ways.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Persistent.&lt;/strong&gt; The agent has a name, memory across conversations, and keeps working while you're offline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A real machine.&lt;/strong&gt; Each product gives the agent a cloud computer with a browser, file system and logged-in sessions. xAI's FAQ puts it plainly: "Every Bot on your account uses one persistent cloud computer." That's what lets an agent operate websites that have no API, save files and sign into accounts like a human at a desk.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proactive.&lt;/strong&gt; Dots keeps researching with read-only tools while idle, Grok Bot runs routines on a schedule or on events, and Muse watches inboxes and prices and pings you when something changes.&lt;/p&gt;

&lt;p&gt;In other words: these vendors are selling a workstation for software that behaves like a colleague. And the moment you see it that way, the important questions turn into access-control questions. You don't ask a new hire about their IQ first. You ask what they have access to.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenAI Dots: one cloud computer per agent
&lt;/h2&gt;

&lt;p&gt;Dots runs on GPT-6 Astra. Each dot gets its own cloud computer, separate from your machine; local desktop access is &lt;a href="https://help.openai.com/en/articles/20001530-getting-started-with-your-dot" rel="noopener noreferrer"&gt;off by default&lt;/a&gt; and has to be turned on through the ChatGPT desktop app. You talk to it through the ChatGPT apps, Slack, Microsoft Teams and voice calls, with SMS available as a limited beta for US Pro users. It inherits your ChatGPT app connections and, per OpenAI, reaches more than 4,000 apps through plugins.&lt;/p&gt;

&lt;p&gt;The permission model is the interesting part. OpenAI's &lt;a href="https://openai.com/index/introducing-dots/" rel="noopener noreferrer"&gt;announcement&lt;/a&gt; says "Custom Rules let you allow specific actions, require approval, or block them," and the Help Center lists four levels: &lt;em&gt;Take action without asking&lt;/em&gt;, &lt;em&gt;Take action if pre-approved&lt;/em&gt;, &lt;em&gt;Ask before taking action&lt;/em&gt; and &lt;em&gt;Hand off to you&lt;/em&gt;. Permanent deletion and software installs go through approval; password changes and money transfers are handed back to you entirely.&lt;/p&gt;

&lt;p&gt;The part I like most: proactive research is architecturally constrained. When the dot works on its own initiative, it uses "tools that are restricted to be read-only, which means that they can't send messages, change app content, or control your browser or computer." That's a clean split between &lt;em&gt;looking around&lt;/em&gt; and &lt;em&gt;acting&lt;/em&gt;, and neither competitor draws that line as clearly.&lt;/p&gt;

&lt;p&gt;Pricing: one dot is included in ChatGPT Pro and Business Premium; more dots are announced for later, without a price yet. Availability matters if you're in Europe: the Pro rollout excludes the EEA, Switzerland and the UK, while Business Premium is available "across all supported ChatGPT regions." Enterprise, Edu and Healthcare workspaces get a beta that an admin must enable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Meta Muse: an egress guard that approves every outbound action
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/" rel="noopener noreferrer"&gt;Muse&lt;/a&gt; runs on Meta's Muse Spark model and targets consumers first: email, travel bookings, bill negotiation, forms. It's available on iOS, Android, muse.ai and WhatsApp.&lt;/p&gt;

&lt;p&gt;Architecturally, it's the most interesting of the three. Every user gets a dedicated "Muse Secure VM" holding both the agent and the user's data. Next to it runs a second agent, the &lt;strong&gt;Sentinel&lt;/strong&gt;, separated from Muse at the system level. Per Meta, nothing Muse does reaches the internet unless the Sentinel approves it, and it asks the user for permission when needed. Credentials sit in secure storage that Muse can use without seeing them. Payments go through one-time cards via Stripe Link. A "Confidential VM" with user-held keys is promised for later this year.&lt;/p&gt;

&lt;p&gt;So Muse separates thinking from acting into two processes: the model proposes, the Sentinel decides what leaves the box. Conceptually, that's the strongest answer to prompt injection any vendor ships today. If the agent reads a hidden instruction on a crafted page telling it to email your last bank statements to a stranger, that send still has to pass a guard that knows what you actually asked for. Whether that holds up in daily use, nobody can say yet, Meta included.&lt;/p&gt;

&lt;p&gt;The launch already produced some data points. &lt;a href="https://www.aol.com/articles/meta-launches-ai-agent-access-190442000.html" rel="noopener noreferrer"&gt;Reuters reported&lt;/a&gt; on internal Meta posts describing the agent routing around guardrails to expose a person's private iCloud photos in testing, and monitoring tasks silently stopping after about 15 minutes. Interactions are used for training by default (opt-out available). Pricing is the most transparent of the three: free up to 100M tokens per week, then $20 or $100 per month. It's available in the US and Canada, 18+.&lt;/p&gt;

&lt;h2&gt;
  
  
  Grok Bot: many bots, one computer, shared logins
&lt;/h2&gt;

&lt;p&gt;Grok Bot launched in beta on August 11 and ships with SuperGrok and Cursor plans (xAI is now SpaceXAI, part of SpaceX, which acquired Cursor). It's positioned as a team: you create several bots with different roles and let two to six of them coordinate in group chats. You can show a bot a workflow once, and it turns the demonstration into a draft skill you review, then runs it as a routine on a schedule or trigger. It prefers connectors where they exist and falls back to driving the browser.&lt;/p&gt;

&lt;p&gt;Here's the difference that matters: &lt;strong&gt;all bots on an account run on one shared cloud computer.&lt;/strong&gt; Each gets its own screen, but browser cookies, sessions, files and command-line credentials are shared. Isolation between &lt;em&gt;users&lt;/em&gt; is strict: &lt;a href="https://docs.x.ai/grok-bot/security-faq" rel="noopener noreferrer"&gt;per xAI's security FAQ&lt;/a&gt;, each gets a dedicated Firecracker microVM. Isolation between bots in one account doesn't exist. xAI is upfront about it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Do not use separate Bots as a security boundary."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And from the &lt;a href="https://docs.x.ai/grok-bot/bots" rel="noopener noreferrer"&gt;Bots docs&lt;/a&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Deleting a Bot removes its active profile, conversation, and routines from Grok Bot. Shared computer files and sign-ins are not isolated by Bot and may remain on the computer."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Think through what that means in practice: the bot that reads unfiltered customer email uses the same browser sessions as the bot that prepares bank transfers. For real privilege separation, you need separate accounts, because only those give you separate machines.&lt;/p&gt;

&lt;h2&gt;
  
  
  The comparison
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2Fd836176932804d6a8ff9f35f5b2b63ea%3Fformat%3Dpng%26w%3D1600" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fcms.nextlevels.de%2Fapi%2Ffiles%2Ftransform%2Fd836176932804d6a8ff9f35f5b2b63ea%3Fformat%3Dpng%26w%3D1600" alt="Three isolation models: Dots one computer per agent, Muse one VM per user plus Sentinel, Grok Bot one shared computer per account" width="1600" height="900"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;OpenAI Dots&lt;/th&gt;
&lt;th&gt;Meta Muse&lt;/th&gt;
&lt;th&gt;xAI Grok Bot&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Launch&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Sep 29, 2026&lt;/td&gt;
&lt;td&gt;Sep 8, 2026&lt;/td&gt;
&lt;td&gt;Aug 11, 2026 (beta)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Model&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;Muse Spark&lt;/td&gt;
&lt;td&gt;no fixed model named&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Compute&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;own cloud computer per dot&lt;/td&gt;
&lt;td&gt;one Secure VM per user&lt;/td&gt;
&lt;td&gt;one shared cloud computer per account&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Isolation within an account&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;per dot&lt;/td&gt;
&lt;td&gt;per user + Sentinel egress guard&lt;/td&gt;
&lt;td&gt;none: files, sessions, logins shared&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Sensitive actions&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Custom Rules (4 levels), proactive mode read-only&lt;/td&gt;
&lt;td&gt;Sentinel approves all network actions, user approval when needed&lt;/td&gt;
&lt;td&gt;Auto-review rules, drafts approved before sending, admin policies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pricing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1 dot in ChatGPT Pro / Business Premium&lt;/td&gt;
&lt;td&gt;free tier, $20, $100 / month&lt;/td&gt;
&lt;td&gt;bundled with SuperGrok and Cursor plans&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;EU availability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Pro: no · Business Premium: yes · Enterprise: beta&lt;/td&gt;
&lt;td&gt;US + Canada only&lt;/td&gt;
&lt;td&gt;no regional statement in the docs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Isolation decreases from left to right. All three are legitimate engineering choices. But for an agent that touches customer data, the shared box is the one that doesn't fit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the isolation model beats the model
&lt;/h2&gt;

&lt;p&gt;An LLM that reads web pages, emails and PDFs will inevitably read text written by attackers. Prompt injection can't be patched away, because the model can't reliably separate instructions from content. Every improvement in hit rate lowers the risk; none removes it.&lt;/p&gt;

&lt;p&gt;What decides the damage is &lt;strong&gt;the blast radius of a compromised agent&lt;/strong&gt;, and that radius &lt;em&gt;is&lt;/em&gt; the isolation model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Dots:&lt;/strong&gt; the radius ends at that dot's own computer. A hijacked research dot can't touch the sessions of an accounting dot on a different machine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Muse:&lt;/strong&gt; the radius ends at the Sentinel. Inside the VM the agent can hallucinate all it wants; only what the guard approves gets out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grok Bot:&lt;/strong&gt; the radius ends at the account. A bot reading a poisoned page sits on the same machine, with the same cookies, files and open logins, as every other bot you own.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The sysadmin version: Dots gives every employee their own laptop. Muse gives every employee their own laptop and puts a doorman at the exit who reads every letter before it goes out. Grok Bot sits ten employees at one shared PC where everyone stays logged in, and honestly writes in the manual: if you want separation of privileges, buy more PCs.&lt;/p&gt;

&lt;p&gt;Credit to xAI for writing it down. But a warning in a manual shifts responsibility to you, and it's a different thing from a boundary you can't technically bypass.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do before these land in your stack
&lt;/h2&gt;

&lt;p&gt;You don't need any of these products to start on the part that actually matters. Write the permission model first, vendor-agnostic. It can be as simple as a policy file you keep in your repo:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Illustrative least-privilege policy for one agent (not a vendor format)&lt;/span&gt;
&lt;span class="na"&gt;agent&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;invoice-intake&lt;/span&gt;
&lt;span class="na"&gt;owner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;finance-lead@example.com&lt;/span&gt;     &lt;span class="c1"&gt;# a named human, like a manager for an employee&lt;/span&gt;
&lt;span class="na"&gt;scopes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;mailbox&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;   &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;read&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;true&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt;  &lt;span class="nv"&gt;send&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;approval&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;   &lt;span class="c1"&gt;# drafting is fine, sending needs a human&lt;/span&gt;
  &lt;span class="na"&gt;calendar&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;  &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;read&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;true&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt;  &lt;span class="nv"&gt;write&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;true&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
  &lt;span class="na"&gt;erp&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;       &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;read&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;true&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt;  &lt;span class="nv"&gt;post_booking&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;false&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
  &lt;span class="na"&gt;banking&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;   &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;read&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;false&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;write&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;false&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;     &lt;span class="c1"&gt;# never&lt;/span&gt;
  &lt;span class="na"&gt;crm&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;       &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;read&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;true&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt;  &lt;span class="nv"&gt;write&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;true&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;delete&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;false&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;span class="na"&gt;isolation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;dedicated_runtime&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;    &lt;span class="c1"&gt;# no shared browser sessions with other agents&lt;/span&gt;
  &lt;span class="na"&gt;egress&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;allowlist&lt;/span&gt;          &lt;span class="c1"&gt;# only the domains this job needs&lt;/span&gt;
&lt;span class="na"&gt;logging&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;every_action&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;         &lt;span class="c1"&gt;# with the agent's stated reason&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With that file in hand, evaluating Dots, Muse or Grok Bot can become an afternoon of mapping their controls onto your columns instead of a quarter of discovery. Two more things worth deciding early:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Build or rent.&lt;/strong&gt; These three are rented agents on someone else's infrastructure, with models you can't pick and quotas nobody has priced long-term. Open-source harnesses like OpenClaw or Hermes Agent run on your own infra with any model provider. They aren't automatically safer: OpenClaw received &lt;a href="https://labs.cloudsecurityalliance.org/research/csa-research-note-hermes-agent-cves-20260504-csa-styled/" rel="noopener noreferrer"&gt;nine CVEs between March 18 and 21, 2026&lt;/a&gt;, including a CVSS 9.9 authorization bypass. But they give you control over the agent's computer, which is the variable this whole post is about.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Start boring.&lt;/strong&gt; An agent that reads invoices from a mailbox and files them, with read-only access, is a good first case. An agent that negotiates contracts belongs at the end of the list.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you're in the EU, keep two constraints in mind: GDPR Article 22 limits automated decisions with legal or similarly significant effects without human involvement, and the AI Act's transparency obligations for systems interacting with people have applied since August 2, 2026. An agent that writes emails and makes calls on your behalf has to disclose that it's an AI. Formally that duty sits with the provider, but the problem lands on you.&lt;/p&gt;

&lt;p&gt;Treat Dots, Muse and Grok Bot as a preview of what your own agents will look like soon, and build the frame they're allowed to work in now. Which one ends up in your company, or whether you build your own, then becomes a purchasing decision instead of a security decision.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published in German at &lt;a href="https://next-levels.de/blog/openai-dots-meta-muse-grok-bot-wem-gehort-der-computer-deines-ki-agenten" rel="noopener noreferrer"&gt;next-levels.de&lt;/a&gt;. I'm co-founder and CEO of Next Levels, a digital agency in Germany; AI consulting is one of our service lines.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>security</category>
      <category>architecture</category>
    </item>
    <item>
      <title>AI Agents as Digital Employees: Architecture and Lessons from Practice</title>
      <dc:creator>Slawa</dc:creator>
      <pubDate>Fri, 03 Jul 2026 19:19:45 +0000</pubDate>
      <link>https://dev.to/slawanextlevels/ai-agents-as-digital-employees-architecture-and-lessons-from-practice-3khd</link>
      <guid>https://dev.to/slawanextlevels/ai-agents-as-digital-employees-architecture-and-lessons-from-practice-3khd</guid>
      <description>&lt;p&gt;The "digital employee" is the most heavily sold and least understood product of 2026. Vendor slides promise a colleague who never sleeps. What arrives in most projects is a very fast intern with no memory who makes every mistake with complete confidence.&lt;/p&gt;

&lt;p&gt;This isn't a polemic against the technology. We build these agents ourselves, and they work. But they only work if you treat them as what they are: software with probabilistic behavior that needs to be onboarded, constrained, and supervised like a new hire. That's exactly where most projects fail. &lt;a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" rel="noopener noreferrer"&gt;Gartner predicts&lt;/a&gt; that over 40 percent of all agentic AI projects will be canceled by the end of 2027. The stated reasons are telling: exploding costs, unclear business value, missing risk controls. Not: "the models were too dumb."&lt;/p&gt;

&lt;h2&gt;
  
  
  What an AI agent actually is (and isn't)
&lt;/h2&gt;

&lt;p&gt;The sober definition: an AI agent is a language model running in a loop. It gets a goal, decides for itself which tool to use next (a database query, an email, an API call), evaluates the result, and keeps going until the task is done. In plain terms: a chatbot you gave hands to.&lt;/p&gt;

&lt;p&gt;The difference from a classic workflow is decision freedom. An n8n workflow follows a fixed path a human laid out. An agent picks its own path. That makes it valuable for tasks whose sequence can't be scripted in advance, and dangerous for everything else.&lt;/p&gt;

&lt;p&gt;What an agent is not: an employee in any legal or organizational sense. It has no sense of responsibility, no liability, and no interest in still being employed tomorrow. The "digital employee" metaphor is useful as a thinking model because it forces the right questions: What is it allowed to do? Who supervises it? Who does it report to? As a description of the technology, it's marketing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture: four building blocks that decide success
&lt;/h2&gt;

&lt;p&gt;The architecture of a production-grade agent system is surprisingly conservative. The model itself is the most replaceable part: models get better and cheaper every few months, and a well-built system swaps them like a graphics card. What stays, and what you therefore have to get right, is everything around it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First: task scoping.&lt;/strong&gt; The most common architecture mistake happens before the first line of code: you give the agent a job title instead of a task. "Handle support" is not a task description, it's a capitulation. Production-grade agents have a narrow mandate with a measurable outcome: "Classify incoming tickets, answer the three most common categories yourself, escalate the rest." The tighter the mandate, the higher the reliability.&lt;/p&gt;

&lt;p&gt;That's not a temporary state of the technology; it follows from its statistics. In a process with twenty steps at 95 percent reliability per step, you get the right end result in only about a third of all runs. Errors in agent loops are cumulative.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second: orchestrator-worker instead of jack-of-all-trades.&lt;/strong&gt; The pattern that has won out in practice separates planning from execution. An orchestrator agent decomposes the task and delegates to specialized sub-agents, each with its own context window, its own tools, and its own narrow assignment. Anthropic measured that a multi-agent setup with an orchestrator beat a single-agent system by roughly 90 percent on their internal research benchmark. In short: many small specialists beat one big generalist, just like in human teams.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third: standardized tool access.&lt;/strong&gt; An agent is only as useful as the systems it can reach. Here, the Model Context Protocol (MCP) has become the open standard: introduced by Anthropic in late 2024, since adopted by OpenAI, Google, and Microsoft, and under the Linux Foundation umbrella since December 2025. Instead of building a custom connector for every model-system combination, the agent speaks one standard, and your CRM, ERP, or inventory system exposes its functions as an MCP server. If your team is building integrations today, build them as MCP servers. That's the one architecture decision that keeps model and vendor switches open later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fourth: guardrails that deserve the name.&lt;/strong&gt; A digital employee needs the same three things as a human one on probation: limited permissions, defined approval processes, and someone watching. Translated to engineering: a permission model at the tool level (the agent that reads invoices can't pay them), human-in-the-loop for everything irreversible (send, delete, pay, publish), and complete logging of every action including its rationale. The logging isn't compliance theater; it's your most important development tool. Without traces you simply cannot debug a non-deterministic system.&lt;/p&gt;

&lt;h2&gt;
  
  
  The economics: agents are more expensive than the demo looks
&lt;/h2&gt;

&lt;p&gt;The point that's consistently missing from pitch decks: agents burn tokens. &lt;a href="https://www.anthropic.com/engineering/multi-agent-research-system" rel="noopener noreferrer"&gt;Anthropic's engineering team&lt;/a&gt; puts a single agent's consumption at roughly four times a chat interaction; multi-agent systems run at about fifteen times. That's not a bug — it's the mechanism through which these systems produce their performance: more parallel reasoning, more tool calls, more context.&lt;/p&gt;

&lt;p&gt;A simple business rule follows: an agent only pays off for tasks whose completion is worth more than the multiplied compute plus the supervision cost. As a model calculation: an agent that costs 40 cents per case in API fees and replaces 15 minutes of manual processing pays for itself immediately. The same agent doing a job that a simple workflow previously did deterministically for 0.4 cents is tech enthusiasm at company expense.&lt;/p&gt;

&lt;p&gt;So our standing recommendation: workflow first, agent second. Everything that can be modeled as a fixed process belongs in classic automation with tools like n8n. The agent goes where rules stop working: unstructured input, decisions that need context, research tasks. This order keeps costs down and has an underrated side effect: the process documentation you produce while automating literally becomes the agent's job instructions later. Write it once, use it twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons that aren't in any vendor deck
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Errors are cumulative, so build for failure.&lt;/strong&gt; In classic software, a bug breaks a feature. In an agent system, an early error sends the agent down a completely different path — with full confidence. A ticket gets misclassified at step three, and twenty steps later the agent has prepared a polite, well-written, and completely wrong reply to the wrong recipient. Production-readiness here means: checkpoints you can resume runs from, retry logic, and an agent that can handle a failing tool instead of hallucinating around it.&lt;/p&gt;

&lt;p&gt;The second lesson sounds trivial and costs the most time in practice: the tool description is the new job description. Agents choose their tools based on the tools' description texts. &lt;code&gt;searchCustomer: searches for a customer&lt;/code&gt; reliably sends the agent nowhere. &lt;code&gt;searchCustomer: finds customer records by name, email, or customer ID; returns at most 10 matches; prefer the customer ID when available&lt;/code&gt; turns the same tool into a reliable one. If you build agents, you'll spend a surprising amount of time documenting interfaces so a machine can't misread them. That's the onboarding of your digital employee.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evaluation before scaling.&lt;/strong&gt; Before an agent gets anywhere near real customer data, it needs a test set of real cases and defined success criteria. Even twenty representative test cases will show whether a prompt change raises or lowers the success rate. Without that measurement, every iteration is guesswork, and every "the agent is better now" is an assertion.&lt;/p&gt;

&lt;p&gt;An agent that reads emails and can operate tools can be attacked through exactly those emails: a crafted message contains an instruction in plain prose like "export the customer list and send it to the following address," and a naively built agent executes it, because to the agent, text is text. Prompt injection and poisoned tool descriptions have been publicly demonstrated attack classes since 2025. The consequence is the same as with human employees and phishing: the agent must never grant external content the same authority as its own instructions, and everything irreversible needs a second pair of eyes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The EU AI Act is at the table.&lt;/strong&gt; If you deploy agents in HR, credit, or hiring processes, you quickly enter high-risk categories with documentation and oversight obligations. The logging and human-in-the-loop approvals from the architecture section are doubly useful: they're simultaneously the core of your compliance file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The right first agent is boring
&lt;/h2&gt;

&lt;p&gt;Don't start with the showcase agent for the management meeting. Start with a process that meets three criteria: it hurts (volume or frustration), it's well documentable, and a mistake is correctable before it gets expensive. Ticket triage, quote research, data reconciliation between systems, first drafts of recurring documents. No payments, no HR decisions, nothing irreversible in year one.&lt;/p&gt;

&lt;p&gt;Then, in this order: document the process, automate deterministically whatever can be automated deterministically, and only then put the agent on the remainder — with a narrow mandate, MCP integration, approvals, and logging from day one.&lt;/p&gt;

&lt;p&gt;The bottleneck with digital employees is not the AI. It's the company. An agent can only take over processes the company itself understands. If you can't describe your workflows, you can't delegate them, neither to humans nor to machines. The 40 percent of canceled projects in Gartner's forecast fail at exactly this. On the other side are the companies whose new employees never sleep. It's worth being one of them.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published in German at &lt;a href="https://next-levels.de/blog/ki-agenten-als-digitale-mitarbeiter-architektur-und-lektionen-aus-der-praxis" rel="noopener noreferrer"&gt;next-levels.de&lt;/a&gt;. We're a full-service digital agency; building exactly these agent architectures for mid-sized companies is part of our AI consulting practice.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Self-Hosted vs. SaaS: What Coolify Actually Costs (and Where It Gets Expensive)</title>
      <dc:creator>Slawa</dc:creator>
      <pubDate>Thu, 02 Jul 2026 14:19:11 +0000</pubDate>
      <link>https://dev.to/slawanextlevels/self-hosted-vs-saas-what-coolify-actually-costs-and-where-it-gets-expensive-1icd</link>
      <guid>https://dev.to/slawanextlevels/self-hosted-vs-saas-what-coolify-actually-costs-and-where-it-gets-expensive-1icd</guid>
      <description>&lt;p&gt;Your deploy stack is probably the worst-negotiated subscription in the whole company. Every server migration gets costed to three decimal places, but the Vercel bill that quietly creeps up a few dollars every month? Nobody consciously signs off on that. It just runs.&lt;/p&gt;

&lt;p&gt;That's the gap &lt;a href="https://coolify.io" rel="noopener noreferrer"&gt;Coolify&lt;/a&gt; walks into. It promises the thing a lot of teams have been quietly thinking: why pay $20 per seat or $25 per process to a US platform when a $6 server hosts the same app? The answer isn't "never" and it isn't "always." It's a calculation — and that calculation has one line item both sides conveniently leave off the landing page.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Coolify is (and what it isn't)
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;git push&lt;/code&gt;, and seconds later your app is live on your domain, TLS cert included, on a server you own. That's the trick you'd credit Heroku with and never your own $6 VPS. Coolify delivers it.&lt;/p&gt;

&lt;p&gt;Technically, it's a self-hosted PaaS, open source under Apache 2.0, positioned as an alternative to Vercel, Heroku and Netlify. You connect a server over SSH, attach a Git repo (GitHub, GitLab, Gitea, or self-hosted), pick a build pack (Nixpacks, Dockerfile, Docker Compose, or static), and Coolify does the rest: builds the image, runs the container, terminates TLS via Traefik with automatic Let's Encrypt certs, and serves the app. On top of that: 280+ one-click services, from databases to analytics to full apps, with built-in backup routines.&lt;/p&gt;

&lt;p&gt;The install is genuinely a single script (grab the current one-liner from the &lt;a href="https://coolify.io/docs" rel="noopener noreferrer"&gt;Coolify docs&lt;/a&gt; — it's the classic &lt;code&gt;curl … | bash&lt;/code&gt; shape), then you're in a web UI.&lt;/p&gt;

&lt;p&gt;Here's the part that decides the whole cost equation later: &lt;strong&gt;what Coolify is *not&lt;/strong&gt;&lt;em&gt;. It's not a global CDN, not autoscaling edge infrastructure, and not a company that carries the pager at 3 a.m. Everything Vercel or Heroku bundle into the subscription — scaling, availability, platform-level security patching — becomes *your&lt;/em&gt; job on Coolify. It automates the deploy, not the operations.&lt;/p&gt;

&lt;p&gt;On maturity, since that's the trust question: Coolify hit &lt;strong&gt;stable with v4.0.0 in April 2026&lt;/strong&gt; after roughly two years in beta. v4.1 (May 2026) added an extra build path (Railpack), structured audit logging for API and auth events, and MCP support. v5 with multi-server scaling is actively being built. Not a weekend project anymore — but no vendor SLA behind it either.&lt;/p&gt;

&lt;h2&gt;
  
  
  The math nobody runs
&lt;/h2&gt;

&lt;p&gt;Skip the ranges, here's a concrete workload: a Next.js app with Postgres and Redis, moderate traffic.&lt;/p&gt;

&lt;p&gt;On Heroku, two Standard-1X dynos ($25 each) plus a managed Postgres add-on gets you to roughly &lt;strong&gt;$100/month — about $1,200/year&lt;/strong&gt;. The same app on a €12 VPS running Coolify: &lt;strong&gt;~€145/year&lt;/strong&gt;. Same workload.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Line item&lt;/th&gt;
&lt;th&gt;Managed PaaS&lt;/th&gt;
&lt;th&gt;Coolify (self-hosted)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Software / platform&lt;/td&gt;
&lt;td&gt;$40–100/mo&lt;/td&gt;
&lt;td&gt;$0 (Cloud optional: $5)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Server / infra&lt;/td&gt;
&lt;td&gt;included in plan&lt;/td&gt;
&lt;td&gt;from ~€6/mo (1 VPS)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bandwidth overage&lt;/td&gt;
&lt;td&gt;variable, grows with traffic&lt;/td&gt;
&lt;td&gt;inside the VPS allowance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Direct cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$40–100+&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~€6–20&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few reference prices so this isn't hand-waving (July 2026):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Heroku:&lt;/strong&gt; Eco $5 (a shared 1,000-hour account pool, &lt;em&gt;not&lt;/em&gt; per process — it sleeps), Basic $7, Standard-1X $25, Standard-2X $50. No free tier since 2022.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vercel Pro:&lt;/strong&gt; $20/seat, 1 TB bandwidth included, $0.15/GB overage, $0.128 per &lt;em&gt;active&lt;/em&gt; CPU-hour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Railway:&lt;/strong&gt; Hobby $5 (incl. $5 usage), Pro $20/seat (incl. $20 usage).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Coolify Cloud&lt;/strong&gt; (managed control plane, infra still separate): $5/mo for 2 servers, +$3 each additional, 20% off annually.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The server:&lt;/strong&gt; a 4 GB VPS at Hetzner starts around &lt;strong&gt;€5.49 (CX23)&lt;/strong&gt; or &lt;strong&gt;€5.99 (ARM CAX11)&lt;/strong&gt;; a beefier CPX22 is ~€19.49. Note Hetzner renamed and repriced the CPX line on &lt;strong&gt;15 June 2026&lt;/strong&gt; (~+144% on that class), so the old "under €10 CPX" boxes are gone — but self-hosting still lands in the low tens of euros a month.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Annualized, the gap between a Heroku setup and self-hosted Coolify is &lt;strong&gt;3–9×&lt;/strong&gt;. If you're paying more than ~$15–20/month for deploy subscriptions, you're cheaper on raw server cost immediately. That's the number in every Coolify pitch, and it's true.&lt;/p&gt;

&lt;p&gt;It's just not the whole bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  The line item the landing page hides: your time
&lt;/h2&gt;

&lt;p&gt;The 3–9× advantage holds &lt;em&gt;exactly&lt;/em&gt; as long as someone keeps the server alive. And that someone costs money — usually more than the cloud bill you saved, if you're honest.&lt;/p&gt;

&lt;p&gt;It's 3 a.m. A kernel update took the Docker daemon with it, the customer portal is down, and the pager is you. That call is what the $20 at Vercel buys away. What's baked into the managed price and lands on your desk with Coolify: OS and security updates, monitoring, actually &lt;em&gt;testing&lt;/em&gt; backups (not just configuring them), incident response, hardening the box, keeping Coolify itself current. None of it is rocket science. But it's work someone else used to do, now sitting with your team. And when the one colleague who set Coolify up leaves, your cost saving just became a bus factor of one.&lt;/p&gt;

&lt;p&gt;Put the hour in the spreadsheet. One conservative hour of DevOps time a month, at loaded internal cost, eats most of what you save on a &lt;em&gt;single&lt;/em&gt; small project. Self-hosting doesn't pay off because a server is cheaper than a subscription. It pays off when you already have those ops hours in-house and spread them across &lt;em&gt;several&lt;/em&gt; projects. One VPS with Coolify hosts a dozen apps at the same flat price. That's where the math tips.&lt;/p&gt;

&lt;h2&gt;
  
  
  When it's worth it — and when it isn't
&lt;/h2&gt;

&lt;p&gt;Straight recommendation, no hedging. Coolify is the right call when three things line up:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;You have DevOps capability in-house&lt;/strong&gt; (or want to build it deliberately) — someone comfortable with SSH, Docker and a Linux box, who &lt;em&gt;wants&lt;/em&gt; to be.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You run several apps&lt;/strong&gt;, so one maintained server amortizes across many deployments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data ownership and independence from US platforms are a real criterion&lt;/strong&gt;, not a slide-deck line.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Miss one of those and you should stay on managed PaaS, at least for now. A solo founder with one Next.js app who'd rather ship features than patch servers belongs on Vercel. A team with zero Docker experience will spend the saved euros back as learning curve and downtime. Self-hosting isn't moral high ground — it's an operational decision.&lt;/p&gt;

&lt;p&gt;And the most-overlooked answer: &lt;strong&gt;it doesn't have to be all or nothing.&lt;/strong&gt; Data-critical core app on your own Coolify server, marketing landing page on a CDN with a generous free tier. Drawing that line well is the actual skill.&lt;/p&gt;

&lt;h2&gt;
  
  
  A realistic way to try it
&lt;/h2&gt;

&lt;p&gt;No big bang required. A cheap VPS, Coolify installed from the one-liner, one non-critical internal tool as your first deploy — that's an afternoon, not a quarter. On that one project you'll learn fast whether your team &lt;em&gt;wants&lt;/em&gt; to own the operations, or whether after the third late-night "the container is gone" the calm of a managed subscription is suddenly worth $20 to you. Both answers are worth the money, and you only get them by actually standing it up instead of theorizing.&lt;/p&gt;

&lt;p&gt;So the concrete first step: open your last Vercel or Heroku invoice and multiply it by twelve. If the annual number makes you wince, and someone on the team can keep a server alive, the question's already half-answered.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by the engineering team at &lt;a href="https://nextlevels.de" rel="noopener noreferrer"&gt;Next Levels&lt;/a&gt;, a German digital agency that builds and operates setups like this. If you've moved a real production workload from managed PaaS to Coolify, I'd genuinely like to hear where the ops cost surprised you — drop it in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>coolify</category>
      <category>selfhosting</category>
      <category>devops</category>
      <category>paas</category>
    </item>
    <item>
      <title>GEO: Wie du dafür sorgst, dass ChatGPT &amp; Co. deine Seite zitieren</title>
      <dc:creator>Slawa</dc:creator>
      <pubDate>Wed, 24 Jun 2026 03:27:06 +0000</pubDate>
      <link>https://dev.to/slawanextlevels/geo-wie-du-dafur-sorgst-dass-chatgpt-co-deine-seite-zitieren-417i</link>
      <guid>https://dev.to/slawanextlevels/geo-wie-du-dafur-sorgst-dass-chatgpt-co-deine-seite-zitieren-417i</guid>
      <description>&lt;p&gt;Dein bestes Google-Ranking ist wertlos, wenn die Antwort schon vor dem Klick gegeben wurde. Genau das passiert gerade: Nutzer fragen ChatGPT, Claude oder Perplexity – und bekommen eine fertige Antwort mit drei, vier zitierten Quellen. Bist du nicht darunter, existierst du in diesem Moment nicht. Kein Ranking, kein Klick, keine zweite Chance.&lt;/p&gt;

&lt;p&gt;Die Disziplin, die das adressiert, heißt &lt;strong&gt;Generative Engine Optimization (GEO)&lt;/strong&gt;. Und sie ist – anders als der Marketing-Lärm vermuten lässt – zu großen Teilen ein Engineering-Problem. Crawler-Zugang, Rendering, strukturierte Daten. Lauter Dinge, über die ein Entwickler entscheidet, nicht das Content-Team.&lt;/p&gt;

&lt;h2&gt;
  
  
  SEO optimiert auf den Klick. GEO optimiert auf das Zitat.
&lt;/h2&gt;

&lt;p&gt;Der Unterschied ist nicht kosmetisch. Klassisches SEO will, dass du auf Platz eins rankst, damit jemand klickt. GEO will, dass ein Sprachmodell deinen Absatz &lt;strong&gt;wörtlich in seine Antwort übernimmt&lt;/strong&gt; – inklusive Quellenangabe. Der Klick ist nur noch Bonus.&lt;/p&gt;

&lt;p&gt;Daraus folgt ein anderer Tech-Stack an Signalen:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspekt&lt;/th&gt;
&lt;th&gt;Klassisches SEO&lt;/th&gt;
&lt;th&gt;GEO / KI-Sichtbarkeit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ziel&lt;/td&gt;
&lt;td&gt;Top-10 in Google&lt;/td&gt;
&lt;td&gt;Zitat in ChatGPT, Claude, Perplexity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Relevante Bots&lt;/td&gt;
&lt;td&gt;Googlebot, Bingbot&lt;/td&gt;
&lt;td&gt;GPTBot, ClaudeBot, PerplexityBot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Index-Hinweis&lt;/td&gt;
&lt;td&gt;&lt;code&gt;sitemap.xml&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;llms.txt&lt;/code&gt; + &lt;code&gt;sitemap.xml&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Strukturierte Daten&lt;/td&gt;
&lt;td&gt;Rich Snippets&lt;/td&gt;
&lt;td&gt;Entity-Linking (&lt;code&gt;Organization&lt;/code&gt;, &lt;code&gt;sameAs&lt;/code&gt;, &lt;code&gt;@graph&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rendering&lt;/td&gt;
&lt;td&gt;Google rendert JS (verzögert)&lt;/td&gt;
&lt;td&gt;viele KI-Bots rendern &lt;strong&gt;kein&lt;/strong&gt; JS → SSR Pflicht&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Erfolgskontrolle&lt;/td&gt;
&lt;td&gt;Search Console, Rank-Tracker&lt;/td&gt;
&lt;td&gt;Citation- &amp;amp; Mention-Tracking in LLMs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Die Hebel überschneiden sich – sauberes HTML, schnelle Antwortzeiten, valides Markup helfen beidem. Aber die Bots, die Index-Signale und die Erfolgskontrolle sind eigenständig. Wer GEO als „SEO mit neuem Namen" abtut, übersieht genau die Stellen, an denen es klemmt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Schritt 1: Lass die Bots überhaupt rein
&lt;/h2&gt;

&lt;p&gt;Bevor du über Content-Qualität nachdenkst, klär die banale Frage: Kommt der Crawler durch? Erstaunlich oft lautet die Antwort nein – und niemand merkt es, weil ein Browser die Seite ja problemlos lädt.&lt;/p&gt;

&lt;p&gt;Die drei User-Agents, die zählen, sind &lt;code&gt;GPTBot&lt;/code&gt;, &lt;code&gt;ClaudeBot&lt;/code&gt; und &lt;code&gt;PerplexityBot&lt;/code&gt;. Eine &lt;code&gt;robots.txt&lt;/code&gt;, die sie durchlässt, sieht so aus:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight robot_framework"&gt;&lt;code&gt;User-agent: GPTBot&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Allow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;User-agent:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;OAI-SearchBot&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Allow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;User-agent:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;ClaudeBot&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Allow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;User-agent:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;PerplexityBot&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Allow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ein Detail, das viele falsch machen: &lt;code&gt;GPTBot&lt;/code&gt; und &lt;code&gt;OAI-SearchBot&lt;/code&gt; sind &lt;strong&gt;nicht dasselbe&lt;/strong&gt;. &lt;code&gt;GPTBot&lt;/code&gt; füttert die Trainingsdaten, &lt;code&gt;OAI-SearchBot&lt;/code&gt; ist für Citations in der ChatGPT-Suche zuständig. Wer aus Datenschutzgründen das Training ausschließen will, sollte nicht pauschal alles von OpenAI sperren – sonst kappt er sich versehentlich die Sichtbarkeit. Training raus, Suche rein:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight robot_framework"&gt;&lt;code&gt;User-agent: GPTBot&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Disallow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;User-agent:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;OAI-SearchBot&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Allow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Der gemeinste Fall ist aber nicht die &lt;code&gt;robots.txt&lt;/code&gt;, sondern die Schicht davor. Cloudflare Bot Fight Mode, eine ModSecurity-Regel, ein nginx-User-Agent-Filter oder ein überambitioniertes Shopware-Plugin weisen unbekannte User-Agents pauschal als Scraper ab. Der Browser sieht davon nie etwas, der KI-Bot bekommt ein 403 oder eine Captcha-Seite. Prüf das im Zweifel direkt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="s2"&gt;"GPTBot"&lt;/span&gt; &lt;span class="nt"&gt;-I&lt;/span&gt; https://deine-domain.de
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Kommt da kein sauberes &lt;code&gt;200&lt;/code&gt; zurück, hast du dein erstes Problem gefunden, bevor du eine Zeile Content angefasst hast. Und Achtung: Cloudflare-Defaults setzen sich bei Updates gern zurück – also nach jedem größeren Release erneut testen.&lt;/p&gt;

&lt;h2&gt;
  
  
  Schritt 2: Liefere Text, kein JavaScript-Versprechen
&lt;/h2&gt;

&lt;p&gt;Hier wird es für moderne Frontends unbequem. Viele KI-Crawler rendern &lt;strong&gt;kein&lt;/strong&gt; JavaScript. Was bei einer reinen Client-Side-React-App im initialen HTML steht, ist oft ein leeres &lt;code&gt;&amp;lt;div id="root"&amp;gt;&lt;/code&gt; – und genau das liest der Bot. Dein schöner Content existiert für ihn nicht.&lt;/p&gt;

&lt;p&gt;Die Lösung ist kein Geheimnis, sie kostet nur Disziplin: Server-Side Rendering oder Static Generation, damit die Hauptinhalte schon im ausgelieferten HTML stehen. In Next.js heißt das, die Finger von reinem Client-Fetching für indexierbaren Content zu lassen:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight jsx"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Server Component – Inhalt steht im initialen HTML, lesbar ohne JS&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;ProductPage&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;params&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;product&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;getProduct&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;article&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;h1&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;product&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;h1&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;p&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;product&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;description&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;p&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;article&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Kurz gesagt: Wenn der Inhalt erst nach dem Hydration-Schritt im DOM auftaucht, ist er für einen erheblichen Teil der KI-Crawler unsichtbar. SSR ist bei GEO keine Performance-Kür mehr, sondern Voraussetzung.&lt;/p&gt;

&lt;h2&gt;
  
  
  Schritt 3: Mach deine Entität eindeutig
&lt;/h2&gt;

&lt;p&gt;Ein Sprachmodell zitiert lieber, was es eindeutig zuordnen kann. Strukturierte Daten sind dafür das stärkste Signal – nicht als SEO-Deko, sondern als maschinenlesbare Aussage darüber, &lt;em&gt;wer du bist&lt;/em&gt;. Das Minimum ist ein &lt;code&gt;Organization&lt;/code&gt;-Schema mit &lt;code&gt;sameAs&lt;/code&gt;-Verknüpfungen, die deine Marke mit ihren bekannten Profilen verbinden:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;script &lt;/span&gt;&lt;span class="na"&gt;type=&lt;/span&gt;&lt;span class="s"&gt;"application/ld+json"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@context&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://schema.org&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Organization&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;name&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Beispiel GmbH&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;url&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://beispiel.de&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;sameAs&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://www.linkedin.com/company/beispiel&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://de.wikipedia.org/wiki/Beispiel_GmbH&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://www.crunchbase.com/organization/beispiel&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/script&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Die &lt;code&gt;sameAs&lt;/code&gt;-Links sind der eigentliche Trick: Sie verankern deine Entität in Quellen, denen das Modell ohnehin vertraut. Darüber hinaus lohnen sich je nach Seitentyp &lt;code&gt;FAQPage&lt;/code&gt; (gut für Frage-Antwort-Formate), &lt;code&gt;Product&lt;/code&gt; + &lt;code&gt;Offer&lt;/code&gt; (Shops) sowie &lt;code&gt;Article&lt;/code&gt; + &lt;code&gt;Author&lt;/code&gt; (Blogs) – am besten gebündelt in einem gemeinsamen &lt;code&gt;@graph&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Schritt 4: llms.txt – kleiner Aufwand, klare Aufwärtschance
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;llms.txt&lt;/code&gt; ist eine vorgeschlagene Konvention (&lt;a href="https://llmstxt.org" rel="noopener noreferrer"&gt;llmstxt.org&lt;/a&gt;), kein offizieller Standard. Die Idee: eine Markdown-Datei unter &lt;code&gt;/llms.txt&lt;/code&gt;, die KI-Systemen eine kuratierte Liste deiner wichtigsten Inhalte nennt – konzeptionell wie eine &lt;code&gt;sitemap.xml&lt;/code&gt;, nur für LLMs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Beispiel GmbH&lt;/span&gt;
&lt;span class="gt"&gt;
&amp;gt; Digitalagentur für E-Commerce, Software und KI-Beratung.&lt;/span&gt;

&lt;span class="gu"&gt;## Wichtige Seiten&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Leistungen&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://beispiel.de/leistungen&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Überblick über alle Services
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Blog&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://beispiel.de/blog&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Fachartikel zu E-Commerce und KI
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Kontakt&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://beispiel.de/kontakt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Ansprechpartner und Standorte
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ehrlich bleiben: Noch lesen längst nicht alle Crawler die Datei aus, einige Recherche-Agents tun es bereits. Der Aufwand ist minimal, das Risiko null – das ist eine der wenigen GEO-Maßnahmen, bei denen sich die Kosten-Nutzen-Rechnung nicht lange diskutieren lässt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Und der Content? Schreib so, dass man dich zitieren kann
&lt;/h2&gt;

&lt;p&gt;Die Technik schafft den Zugang, aber zitiert wird am Ende eine Textpassage. Und Modelle haben eine klare Präferenz: kurze, eigenständige Absätze, klare Definitionen, Zahlen, Listen, saubere Überschriften-Hierarchie. Die Princeton-Forschung hinter dem Begriff GEO (&lt;a href="https://dl.acm.org/doi/10.1145/3637528.3671900" rel="noopener noreferrer"&gt;KDD 2024&lt;/a&gt;) zeigt, dass Zitate, Statistiken und konkrete Quellenangaben die Wahrscheinlichkeit, in einer KI-Antwort aufzutauchen, um bis zu 30–40 % erhöhen.&lt;/p&gt;

&lt;p&gt;Werblicher Fließtext ohne Substanz ist das Gegenteil davon. Ein Absatz, der eine Frage in zwei klaren Sätzen beantwortet, wird zitiert. Ein Absatz voller „innovativer, ganzheitlicher Lösungen" wird übersprungen.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wo stehst du gerade?
&lt;/h2&gt;

&lt;p&gt;Das Unangenehme an GEO: Die meisten dieser Probleme sieht man der Seite im Browser nicht an. Die &lt;code&gt;robots.txt&lt;/code&gt; blockt lautlos, die WAF antwortet Bots anders als dir, das JSON-LD hat ein fehlendes Pflichtfeld. Bevor du anfängst zu optimieren, lohnt sich deshalb eine Bestandsaufnahme aus Bot-Perspektive.&lt;/p&gt;

&lt;p&gt;Dafür haben wir bei nextlevels einen &lt;a href="https://next-levels.de/ki-sichtbarkeit-check" rel="noopener noreferrer"&gt;kostenlosen KI-Sichtbarkeits-Check&lt;/a&gt; gebaut: Er simuliert die wichtigsten KI-Crawler, liest &lt;code&gt;robots.txt&lt;/code&gt;, &lt;code&gt;llms.txt&lt;/code&gt; und Sitemap, prüft das ausgelieferte HTML auf SSR und strukturierte Daten und erkennt Bot-Blocker auf WAF-Ebene. Ergebnis in unter 30 Sekunden, ohne E-Mail, und jeder Einzelbefund kommt mit den Rohdaten, die du selbst per &lt;code&gt;curl&lt;/code&gt; nachprüfen kannst. Genau das macht ihn für Entwickler brauchbar – es ist kein Score zum Glauben, sondern eine Checkliste zum Nachvollziehen.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fazit
&lt;/h2&gt;

&lt;p&gt;GEO ist kein neues Marketing-Buzzword, das man dem Content-Team überlassen kann. Die Stellen, an denen Sichtbarkeit in KI-Antworten entsteht oder stirbt, liegen im Code: in der &lt;code&gt;robots.txt&lt;/code&gt;, in der Render-Strategie, im Schema-Markup, in der WAF-Konfiguration. Die gute Nachricht für uns Entwickler ist, dass das alles überprüfbar und fixbar ist – kein Raten, kein Black-Box-Algorithmus. Fang bei der banalsten Frage an: Kommt der Bot überhaupt durch? Den Rest baust du darauf auf.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>ai</category>
      <category>seo</category>
      <category>german</category>
    </item>
    <item>
      <title>Shopware vs Shopify: a developer's case for the open platform</title>
      <dc:creator>Slawa</dc:creator>
      <pubDate>Wed, 24 Jun 2026 03:24:03 +0000</pubDate>
      <link>https://dev.to/slawanextlevels/shopware-vs-shopify-a-developers-case-for-the-open-platform-10kd</link>
      <guid>https://dev.to/slawanextlevels/shopware-vs-shopify-a-developers-case-for-the-open-platform-10kd</guid>
      <description>&lt;p&gt;Most "Shopware vs Shopify" posts compare dashboards, app stores, and pricing tables. None of that matters to you until the day a client asks for something the platform won't let you build. Then the comparison stops being a feature grid and becomes a question about ceilings: &lt;strong&gt;how high can I go before the platform says no, and what happens when I hit it?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the only axis I care about as a developer, so that's the one I'll argue on. Shopify is an outstanding product. It's also a closed SaaS that decides, on your behalf, where customization ends. Shopware is open source built on Symfony, which means the ceiling is "however far PHP and HTTP will take you." Below are the three places that difference actually bites, with code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Angle 1: The checkout is the wall
&lt;/h2&gt;

&lt;p&gt;This is the headline because it's where most agency developers first hit something they cannot do.&lt;/p&gt;

&lt;p&gt;For years the Shopify answer to "customize the checkout" was &lt;code&gt;checkout.liquid&lt;/code&gt;. That era is over. Shopify deprecated &lt;code&gt;checkout.liquid&lt;/code&gt; in favour of &lt;strong&gt;Checkout Extensibility&lt;/strong&gt;. Plus stores had to migrate their Thank-you and Order-status pages by &lt;strong&gt;August 28, 2025&lt;/strong&gt;, and in January 2026 Shopify began auto-upgrading stores — wiping customizations built on additional scripts, script-tag apps, or &lt;code&gt;checkout.liquid&lt;/code&gt;. Non-Plus stores have until &lt;strong&gt;August 26, 2026&lt;/strong&gt;, and legacy Shopify Scripts keep working only until &lt;strong&gt;June 30, 2026&lt;/strong&gt;. (&lt;a href="https://help.shopify.com/en/manual/checkout-settings/customize-checkout-configurations/upgrade-thank-you-order-status/plus-upgrade-guide" rel="noopener noreferrer"&gt;Shopify migration timeline&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;The replacement, Checkout Extensibility, is genuinely more upgrade-safe. It's also a smaller box. You get &lt;strong&gt;Checkout UI Extensions&lt;/strong&gt; (declarative components that render in slots Shopify defines) and &lt;strong&gt;Shopify Functions&lt;/strong&gt; for backend logic — and that's the surface. You don't own the checkout template; you decorate the pieces Shopify exposes. Worth noting: full visual checkout customization (branding API, custom fields beyond the defaults, full UI extension power) is gated to &lt;strong&gt;Shopify Plus&lt;/strong&gt; anyway.&lt;/p&gt;

&lt;p&gt;On Shopware, the checkout is a Twig template like every other page, and you override it the same way you override anything else — by extending a block:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight twig"&gt;&lt;code&gt;&lt;span class="c"&gt;{# MyPlugin/src/Resources/views/storefront/page/checkout/confirm/index.html.twig #}&lt;/span&gt;
&lt;span class="cp"&gt;{%&lt;/span&gt; &lt;span class="nv"&gt;sb_extends&lt;/span&gt; &lt;span class="s1"&gt;'@Storefront/storefront/page/checkout/confirm/index.html.twig'&lt;/span&gt; &lt;span class="cp"&gt;%}&lt;/span&gt;

&lt;span class="cp"&gt;{%&lt;/span&gt; &lt;span class="k"&gt;block&lt;/span&gt; &lt;span class="nv"&gt;page_checkout_confirm_tos&lt;/span&gt; &lt;span class="cp"&gt;%}&lt;/span&gt;
    &lt;span class="c"&gt;{# Inject a B2B purchase-order field right above the terms checkbox #}&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;div&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"po-number-field"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;label&lt;/span&gt; &lt;span class="na"&gt;for=&lt;/span&gt;&lt;span class="s"&gt;"poNumber"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;&lt;span class="cp"&gt;{{&lt;/span&gt; &lt;span class="s2"&gt;"checkout.poNumberLabel"&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="nf"&gt;trans&lt;/span&gt; &lt;span class="cp"&gt;}}&lt;/span&gt;&lt;span class="nt"&gt;&amp;lt;/label&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;input&lt;/span&gt; &lt;span class="na"&gt;type=&lt;/span&gt;&lt;span class="s"&gt;"text"&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"poNumber"&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"poNumber"&lt;/span&gt;
               &lt;span class="na"&gt;value=&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="cp"&gt;{{&lt;/span&gt; &lt;span class="nv"&gt;page.extensions.poNumber&lt;/span&gt; &lt;span class="err"&gt;??&lt;/span&gt; &lt;span class="s1"&gt;''&lt;/span&gt; &lt;span class="cp"&gt;}}&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/div&amp;gt;&lt;/span&gt;

    &lt;span class="cp"&gt;{{&lt;/span&gt; &lt;span class="nv"&gt;parent&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="cp"&gt;}}&lt;/span&gt;
&lt;span class="cp"&gt;{%&lt;/span&gt; &lt;span class="k"&gt;endblock&lt;/span&gt; &lt;span class="cp"&gt;%}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No slot has to exist for this. No feature has to be on a pricing tier. You're editing the checkout's actual markup, in the same templating language as the rest of the storefront, and your override survives core updates because it extends rather than replaces. The Shopify equivalent — arbitrary markup in the middle of the checkout flow — is simply not a thing you can do, on any plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  Angle 2: Backend logic — a 5ms sandbox vs. the whole framework
&lt;/h2&gt;

&lt;p&gt;Say the requirement is a non-trivial discount: &lt;em&gt;"15% off, but only for B2B customers in a specific customer group, only on products from suppliers we flag as overstocked in an external ERP, and only Monday–Wednesday."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On Shopify this is a &lt;strong&gt;Shopify Function&lt;/strong&gt;: Rust or JavaScript compiled to WebAssembly. It's clever engineering, but read the constraints before you design against it. A function's Wasm module must be &lt;strong&gt;≤ 256 KB&lt;/strong&gt;, may execute &lt;strong&gt;≤ 11 million instructions&lt;/strong&gt;, runs under a &lt;strong&gt;~5 ms&lt;/strong&gt; execution budget, is &lt;strong&gt;fully sandboxed&lt;/strong&gt; (isolated memory, no network calls out to your ERP), and &lt;strong&gt;forbids nondeterminism — no clock, no random&lt;/strong&gt;. Functions also can't be chained or made aware of each other. (&lt;a href="https://shopify.dev/docs/apps/build/functions/programming-languages/webassembly-for-functions" rel="noopener noreferrer"&gt;Shopify Functions / WebAssembly docs&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Look at that list against the requirement. "Only on overstocked products from an external ERP" needs a network call — not allowed in the function. "Monday–Wednesday" needs the clock — not allowed. So the real-world implementation becomes: a separate hosted app syncs ERP + day-of-week state into metafields out-of-band, and the function reads those metafields. You can ship it, but the platform pushed a chunk of your domain logic out of the function and into infrastructure you now operate and keep in sync.&lt;/p&gt;

&lt;p&gt;Here's the shape of what you're allowed to do inside the function — pure, deterministic, metafield-fed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Shopify Function (JS → Wasm). No network, no Date.now(), ≤5ms, ≤256KB.&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;discounts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cart&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;lines&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;merchandise&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;product&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;isOverstocked&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;true&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;targets&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;cartLine&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt;
      &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;percentage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;15.0&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}));&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;discounts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;discountApplicationStrategy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;FIRST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Shopware the same rule is an event subscriber in ordinary PHP, with the full container at your disposal — the database, HTTP clients, your ERP service, the clock, everything:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="cp"&gt;&amp;lt;?php&lt;/span&gt;
&lt;span class="c1"&gt;// MyPlugin/src/Subscriber/OverstockDiscountSubscriber.php&lt;/span&gt;
&lt;span class="kn"&gt;namespace&lt;/span&gt; &lt;span class="nn"&gt;MyPlugin\Subscriber&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Shopware\Core\Checkout\Cart\Cart&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Shopware\Core\Checkout\Cart\Event\CartChangedEvent&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Symfony\Component\EventDispatcher\EventSubscriberInterface&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;OverstockDiscountSubscriber&lt;/span&gt; &lt;span class="kd"&gt;implements&lt;/span&gt; &lt;span class="nc"&gt;EventSubscriberInterface&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;__construct&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="kt"&gt;ErpClient&lt;/span&gt; &lt;span class="nv"&gt;$erp&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;// your own HTTP service&lt;/span&gt;
        &lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="kt"&gt;DiscountFactory&lt;/span&gt; &lt;span class="nv"&gt;$factory&lt;/span&gt; &lt;span class="c1"&gt;// injected, like any Symfony service&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;getSubscribedEvents&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="kt"&gt;array&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nc"&gt;CartChangedEvent&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;class&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s1"&gt;'applyOverstockDiscount'&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;applyOverstockDiscount&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;CartChangedEvent&lt;/span&gt; &lt;span class="nv"&gt;$event&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kt"&gt;void&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nv"&gt;$cart&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nv"&gt;$event&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;getCart&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
        &lt;span class="nv"&gt;$weekday&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;\DateTimeImmutable&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;format&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'N'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// 1..3 = Mon..Wed&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$weekday&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="c1"&gt;// A real network call to your ERP — impossible inside a Shopify Function&lt;/span&gt;
        &lt;span class="nv"&gt;$overstockedSkus&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nv"&gt;$this&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;erp&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;fetchOverstockedSkus&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

        &lt;span class="k"&gt;foreach&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$cart&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;getLineItems&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nv"&gt;$item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;in_array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$item&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;getReferencedId&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nv"&gt;$overstockedSkus&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="nv"&gt;$this&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;factory&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;applyPercentage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$cart&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$item&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;15.0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The point isn't that PHP is nicer than Rust. It's that &lt;strong&gt;the entire business rule lives in one place, inside the request, with no sandbox to design around.&lt;/strong&gt; Shopify's model is safer and more scalable by construction — that's a real trade-off, not a slur — but it's a model where the platform decides which parts of your logic are allowed to run where. Shopware doesn't make that decision for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Angle 3: Where your code lives
&lt;/h2&gt;

&lt;p&gt;This one is structural and easy to underrate.&lt;/p&gt;

&lt;p&gt;A Shopify app is an &lt;strong&gt;external application&lt;/strong&gt;. It runs on infrastructure you host, talks to the store over OAuth and the Admin/Storefront APIs, and reacts via webhooks. Your code is never &lt;em&gt;in&lt;/em&gt; the store; it's a satellite orbiting it through a rate-limited API. That's a clean boundary, and for many integrations it's the right one. But it means even small backend tweaks become a deployed, authenticated, separately-monitored service — and you're always one API version or rate-limit window away from the platform.&lt;/p&gt;

&lt;p&gt;A Shopware plugin is a &lt;strong&gt;Symfony bundle that lives inside the application&lt;/strong&gt;. The class hierarchy is literal: your plugin extends &lt;code&gt;Plugin&lt;/code&gt;, which extends &lt;code&gt;Bundle&lt;/code&gt;, which extends Symfony's &lt;code&gt;Bundle&lt;/code&gt;. (&lt;a href="https://developer.shopware.com/docs/guides/plugins/plugins/plugins-for-symfony-developers.html" rel="noopener noreferrer"&gt;Shopware: Plugins for Symfony Developers&lt;/a&gt;) So a plugin is just a Symfony bundle with conventions, and standard framework wiring applies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MyPlugin/
├── composer.json
└── src/
    ├── MyPlugin.php                       # extends Shopware\Core\Framework\Plugin
    ├── Resources/
    │   ├── config/
    │   │   ├── services.xml               # DI: register your subscribers/services
    │   │   └── routes.xml
    │   └── views/storefront/...           # Twig overrides (see Angle 1)
    ├── Subscriber/
    │   └── OverstockDiscountSubscriber.php # runs in-process (see Angle 2)
    └── Migration/
        └── Migration1700000000Example.php  # owns its own schema changes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="cp"&gt;&amp;lt;?php&lt;/span&gt;
&lt;span class="c1"&gt;// MyPlugin/src/MyPlugin.php&lt;/span&gt;
&lt;span class="kn"&gt;namespace&lt;/span&gt; &lt;span class="nn"&gt;MyPlugin&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Shopware\Core\Framework\Plugin&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;MyPlugin&lt;/span&gt; &lt;span class="kd"&gt;extends&lt;/span&gt; &lt;span class="nc"&gt;Plugin&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// It's a Symfony bundle. Services in services.xml autoload,&lt;/span&gt;
    &lt;span class="c1"&gt;// routes register, migrations run on install. No external host,&lt;/span&gt;
    &lt;span class="c1"&gt;// no OAuth handshake, no API version to chase.&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your subscriber from Angle 2 and your Twig override from Angle 1 are the same deployable unit — one bundle, in the same process as the shop, sharing its container and database. There's no satellite to operate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest trade-offs
&lt;/h2&gt;

&lt;p&gt;I'd be doing the dishonest version of this post if I stopped there.&lt;/p&gt;

&lt;p&gt;Shopify's closed model buys you things that are genuinely hard to replicate: you never patch a server, the checkout converts extremely well out of the box, PCI scope is mostly Shopify's problem, and the sandboxed Functions model means one tenant's bad discount logic can't take down the platform. For a merchant who wants to sell, not to operate software, that's often the correct choice — and as a developer you should say so.&lt;/p&gt;

&lt;p&gt;There's also a cost angle that's easy to forget mid-architecture-debate: if you use any payment gateway other than Shopify Payments, Shopify adds a &lt;strong&gt;third-party transaction fee&lt;/strong&gt; — 2% on Basic, scaling down to 1% (Grow), 0.6% (Advanced), and 0.2% (Plus). (&lt;a href="https://help.shopify.com/en/manual/your-account/manage-billing/billing-charges/types-of-charges/third-party-charges/third-party-transaction-fees" rel="noopener noreferrer"&gt;Shopify third-party transaction fees&lt;/a&gt;) Self-hosted Shopware has no such cut — but you (or your host) now own uptime, security patching, and scaling, which is not free either. You're trading a platform fee for an operational burden. Which one is cheaper depends entirely on the project.&lt;/p&gt;

&lt;p&gt;Shopware's openness is power &lt;em&gt;and&lt;/em&gt; responsibility. The ceiling is high because there basically isn't one — and the flip side of "there's no sandbox to design around" is "there's no sandbox protecting you from yourself."&lt;/p&gt;

&lt;h2&gt;
  
  
  So when does the open platform win?
&lt;/h2&gt;

&lt;p&gt;When the requirement is the thing the closed platform won't let you build. A checkout that needs custom markup, not a Shopify-defined slot. Business logic that has to call your ERP synchronously and consult the clock. A backend tweak that belongs in-process, not in a separately-hosted, OAuth'd, rate-limited satellite. The moment a project has two or three of those, the Shopify ceiling stops being theoretical and starts costing you weeks of working around it — and Shopware's "it's just Symfony" stops being a slogan and starts being the reason you ship on time.&lt;/p&gt;

&lt;p&gt;Pick Shopify when you want the platform to make decisions for you. Pick Shopware when you need to make them yourself. As a developer, you already know which kind of project lands on your desk.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I'm CTO at &lt;a href="https://next-levels.de" rel="noopener noreferrer"&gt;nextlevels&lt;/a&gt;, a German digital agency and Shopware Silver Partner. We build Shopware shops, custom software, and AI workflows for B2B Mittelstand clients — so I've shipped enough of both platforms to have opinions. Happy to argue the trade-offs in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>shopware</category>
      <category>shopify</category>
      <category>ecommerce</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
