<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Matthias | StudioMeyer</title>
    <description>The latest articles on DEV Community by Matthias | StudioMeyer (@studiomeyer_io).</description>
    <link>https://dev.to/studiomeyer_io</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3866458%2F170ce662-470b-4f78-ac37-58a9a2a00220.PNG</url>
      <title>DEV Community: Matthias | StudioMeyer</title>
      <link>https://dev.to/studiomeyer_io</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/studiomeyer_io"/>
    <language>en</language>
    <item>
      <title>Codex Now on Your Phone: Mobile Plus Goal Mode Explained</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Thu, 13 Aug 2026 13:32:32 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/codex-now-on-your-phone-mobile-plus-goal-mode-explained-32nj</link>
      <guid>https://dev.to/studiomeyer_io/codex-now-on-your-phone-mobile-plus-goal-mode-explained-32nj</guid>
      <description>&lt;p&gt;In mid-April, OpenAI Codex had two million weekly users. Six weeks later it has four million. The growth was not the news. The news was what shipped in between to make the next two million possible.&lt;/p&gt;

&lt;p&gt;Two updates in particular matter for people who already use ChatGPT and read &lt;a href="https://studiomeyer.io/en/blog/codex-for-chatgpt-users-guide" rel="noopener noreferrer"&gt;our April Codex guide&lt;/a&gt;. The first is Codex Mobile, which left the Plus-only bucket on May 14 and is now part of every ChatGPT plan including Free and Go. The second is Goal Mode, which exited the experimental flag on May 21 and is now generally available across the web app, the desktop app, the IDE extension and the CLI.&lt;/p&gt;

&lt;p&gt;Neither of these is a new tool. Both are upgrades that change how Codex fits into a real workday. This article is for readers who liked the April beginner's guide and want to know what the May updates actually do, without the OpenAI marketing copy and without code.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed in six weeks
&lt;/h2&gt;

&lt;p&gt;Six features shipped between April 16 and May 22. The user count doubled. GPT-5.5 became the default model behind Codex. Codex Mobile got freed up to every plan. Goal Mode reached general availability. Appshots, which let you attach any macOS window to a Codex thread by hotkey, also went general. A Chrome extension launched that lets Codex work inside live browser sessions instead of in a sandbox.&lt;/p&gt;

&lt;p&gt;That is a lot for six weeks. The April guide covered five surfaces. The list has not grown, it has deepened. The same five Codex environments (Web, iOS, Desktop, VS Code, CLI) now do more with less setup. Two of those changes are big enough to be worth their own walkthrough.&lt;/p&gt;

&lt;h2&gt;
  
  
  Goal Mode in one sentence
&lt;/h2&gt;

&lt;p&gt;Goal Mode is the difference between asking Codex to do a task and asking Codex to reach an outcome.&lt;/p&gt;

&lt;p&gt;In normal Codex use, you give it an instruction like "draft this email" or "summarize this PDF." Codex returns a result and waits. If the result is not quite right, you give it follow-up instructions. The interaction is turn by turn.&lt;/p&gt;

&lt;p&gt;In Goal Mode, you describe the end state and the success criteria. Something like "prepare a clean handoff document for the new freelancer covering everything she needs for the next two weeks. It should be self-contained, written in plain English, and include links to the three relevant Notion pages." Codex breaks that into steps on its own, executes them in sequence, checks its own work against your criteria, and only comes back when it thinks it is done or when it needs your decision on something it cannot resolve.&lt;/p&gt;

&lt;p&gt;This is what the April post promised in one line that turned out to be the most quoted line of the whole article. "Codex executes, ChatGPT replies." Goal Mode is the version of that promise that is actually delegation rather than execution. The previous Codex felt like an assistant who needed instructions for every step. Goal Mode feels like an assistant who reads the brief and gets back to you when there is a question.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Codex Mobile actually is in May 2026
&lt;/h2&gt;

&lt;p&gt;Mobile is the part most people get wrong. Codex Mobile is not a smaller Codex that runs on your phone. It is a remote control for the Codex that already runs on your Mac.&lt;/p&gt;

&lt;p&gt;You start a task on the desktop. You walk to a meeting. The phone shows you the live progress, lets you read intermediate outputs, approve actions Codex wants to take, switch models if the task is heavier than expected, or start a new task that runs in parallel. When you come back to the desk, everything is in the same state. The model did not pause when you closed the laptop. It kept working because the work was never on your laptop in the first place.&lt;/p&gt;

&lt;p&gt;Two practical consequences come from this. First, you can decide that the Mac stays plugged in and unlocked at the office or at home, and the phone is your travel device. The Mac is the workshop, the phone is the dashboard. Second, you stop thinking of mobile as a downgraded version of the real tool. Mobile is the part that goes with you, not the lesser version of what stays at the desk.&lt;/p&gt;

&lt;p&gt;The May 14 launch removed the Plus-paywall on this. Everyone with a ChatGPT account, including Free users and ChatGPT Go subscribers, can install the latest ChatGPT app on iOS or Android and find Codex inside it. The Mac host requirement still applies for the full feature set, but the basic monitoring and task-handoff functions work on phone alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your first Goal Mode moment in five minutes
&lt;/h2&gt;

&lt;p&gt;Open Codex on the web at chatgpt.com/codex. Look for the toggle labeled Goal Mode in the input area, or use the slash command /goal. Both routes work.&lt;/p&gt;

&lt;p&gt;For the first try, do not pick a coding task. Try something like this. "Goal: I have an unread inbox with twelve emails from clients. Read all of them and prepare a triage list. Group them into reply now, reply this week, do not need a reply, and unclear. For each one in the reply now group, draft a short response based on what I usually write. Success criterion: I can act on the triage list in under ten minutes."&lt;/p&gt;

&lt;p&gt;Hit run. Codex will work for several minutes, sometimes longer than expected. You can close the browser. When it is done, you have a structured triage list and three or four draft replies that match your own tone if you have been using ChatGPT consistently long enough that the model knows how you write.&lt;/p&gt;

&lt;p&gt;The first time this happens, the difference from regular ChatGPT is obvious. ChatGPT would have replied to one email at a time. Codex in Goal Mode worked through all twelve, made its own classification, drafted responses, and packaged everything into a single deliverable. That is the move you do not get from a chat interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ten Goal Mode tasks that have nothing to do with code
&lt;/h2&gt;

&lt;p&gt;These are the tasks where I have seen Goal Mode deliver consistent results for solo operators and small teams in the last two weeks.&lt;/p&gt;

&lt;p&gt;Triage a client inbox at the start of the day and produce a sorted action list. Take a meeting recording transcript and produce both meeting notes and a to-do list for each participant. Read a long PDF contract and produce a list of every change since the previous version. Translate a multi-page document and run a glossary consistency check so the same term is rendered the same way every time. Take a folder of receipts as images and produce an Excel-ready table with date, vendor, amount and category. Compare two Notion pages and produce a third page that merges them without duplicates. Take an outline and turn it into a slide deck with speaker notes. Prepare a freelancer handoff covering background, current state, deadlines and contact list. Read fifty job applications against a job description and produce a ranked shortlist with one-line reasoning per candidate. Run a competitive scan across five named competitor sites and produce a one-page summary of what is new.&lt;/p&gt;

&lt;p&gt;None of these require GitHub, code, or technical setup. All of them benefit from Goal Mode because the work is not one prompt and one reply. The work is a sequence of small judgments that adds up to an outcome.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Codex Mobile is for in the workday
&lt;/h2&gt;

&lt;p&gt;The mistake I see in early Codex Mobile use is starting big tasks from the phone. That works, but it is not the strength. The strength is monitoring and handoff.&lt;/p&gt;

&lt;p&gt;In practice, the pattern that works looks like this. At the start of the workday, on the desktop, you set up two or three Goal Mode tasks that will take twenty to forty minutes each. You go to a meeting, a workout or a coffee. The phone shows you live updates. When a task asks for a decision, you tap to approve or redirect. When a task finishes, you see the result and either accept it or ask for changes. By the time you are back at the desk, two of three tasks are complete and one is waiting for your input.&lt;/p&gt;

&lt;p&gt;The other Mobile pattern is starting small tasks while you are out. Something you would have written down in Notes to do later becomes a Codex task started immediately from the phone. By the time you get to the desk, the first draft is already there.&lt;/p&gt;

&lt;p&gt;The honest limit is that Mobile is dependent on the Mac being awake and online. If your laptop is closed in your bag, the cloud version of Codex still runs but the Mac-host features pause. For pure cloud tasks, Mobile works the same as Web. For tasks that need your local environment, the Mac has to be available.&lt;/p&gt;

&lt;h2&gt;
  
  
  Goal Mode versus normal Codex tasks
&lt;/h2&gt;

&lt;p&gt;Goal Mode is not always the right choice. Normal Codex tasks still make sense for fast one-shots where you know exactly what you want and the work is a single transformation.&lt;/p&gt;

&lt;p&gt;Use a normal task when you want a single output and you can describe it precisely in one sentence. "Rewrite this paragraph to sound less formal." "Translate these three sentences to Spanish." "Format this list as a Markdown table." These are not goal-oriented. They are transformations. Codex handles them faster without Goal Mode.&lt;/p&gt;

&lt;p&gt;Use Goal Mode when the work is multi-step, when success depends on a sequence of decisions, when you want to specify outcomes rather than steps. The triage, the handoff document, the contract diff, the competitive scan. These benefit from the model planning its own approach instead of executing your micro-instructions.&lt;/p&gt;

&lt;p&gt;A useful heuristic. If you would explain the task to a junior assistant in one breath, use normal Codex. If you would brief a freelancer for a half-day project, use Goal Mode.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stumbling blocks I have seen
&lt;/h2&gt;

&lt;p&gt;Three patterns trip people up in the first two weeks of using Goal Mode.&lt;/p&gt;

&lt;p&gt;The first is over-defining the success criteria. "Success: the document is clear, complete, accurate, well-formatted, free of errors, and matches my style" reads like a thorough brief but tells Codex nothing it can check against. Codex evaluates its own work against criteria it can measure. "Success: every email has a draft reply under 80 words" is checkable. "Success: the document is good" is not.&lt;/p&gt;

&lt;p&gt;The second is treating Goal Mode like a fire-and-forget machine for sensitive work. Goal Mode is autonomous in its planning, not in its judgment. A finished output still needs your read. The point of Goal Mode is that you read once at the end instead of reviewing every step, not that you never read.&lt;/p&gt;

&lt;p&gt;The third is forgetting that Codex still has limits on long-running tasks per plan. Plus users will hit the wall on three or four long Goal Mode sessions per week. Pro users have ten times that. The cost calculation for which plan to pick changes the moment you start doing real work in Goal Mode, because each session uses more compute than a regular task.&lt;/p&gt;

&lt;p&gt;If you bumped from Plus to Pro after the April guide, you are in the right plan. If you stayed on Plus and Goal Mode now fits your workflow, the math has shifted.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is still coming this year
&lt;/h2&gt;

&lt;p&gt;The same product strategy that surfaced in March is still on the rails. OpenAI is collapsing ChatGPT, Codex and the Atlas browser into a single desktop super-app over the rest of 2026. The April update was one step. May was another. The next steps that have been hinted at publicly include persistent cross-app memory, deeper Goal Mode integration with the Atlas browser, and a unified billing model that stops separating the chat side from the agent side.&lt;/p&gt;

&lt;p&gt;If you spent six weeks learning Codex, none of that is wasted. The concepts are the same. The packaging gets simpler.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Goal Mode is not for you
&lt;/h2&gt;

&lt;p&gt;If you mostly use ChatGPT to look up facts, draft single short messages, or have conversations to think out loud, Goal Mode is overkill. The chat surface does that work better and faster.&lt;/p&gt;

&lt;p&gt;If you do not have at least one recurring task per week that takes longer than fifteen minutes to do yourself, Goal Mode will not give you back enough time to justify the learning curve.&lt;/p&gt;

&lt;p&gt;If your work is highly sensitive and you cannot have any task run autonomously without a human in the loop on each step, normal Codex with manual approvals is closer to the right tool.&lt;/p&gt;

&lt;p&gt;For everyone else, the math is straightforward. Pick two recurring tasks. Convert them to Goal Mode briefs once. Reuse those briefs every week. The time saving compounds because you stop writing the same instructions over and over.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do next
&lt;/h2&gt;

&lt;p&gt;If you read the April guide and never went past Codex Web, go back to chatgpt.com/codex and switch on Goal Mode for the next task you would normally type as a chat. Watch what changes.&lt;/p&gt;

&lt;p&gt;If you read the April guide and have been using Codex Web regularly, install the ChatGPT app on your phone if it is not there already, and let Mobile run as a dashboard while you set up a Goal Mode task at the desk. The first time the phone tells you the task is done while you are in a meeting, the pattern clicks.&lt;/p&gt;

&lt;p&gt;If you read &lt;a href="https://studiomeyer.io/en/blog/codex-memory-mcp-fix" rel="noopener noreferrer"&gt;our follow-up on Codex memory&lt;/a&gt; and connected an MCP memory layer, Goal Mode reads from the same memory. A Goal Mode brief can reference past decisions, client profiles and project notes without you pasting them in each time. That is when the workflow stops feeling like AI and starts feeling like a team.&lt;/p&gt;

&lt;p&gt;And if you want help wiring Codex, Goal Mode, mobile and a memory layer into the actual workday for a small team, that is what we do at StudioMeyer. The &lt;a href="https://studiomeyer.io/en/services/ki-systeme" rel="noopener noreferrer"&gt;Setting up an AI system&lt;/a&gt; is a six-week build that does exactly this for SMEs who have been paying for ChatGPT Plus for months and never quite turned it into a working tool.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://studiomeyer.io/en/blog/codex-goal-mode-for-chatgpt-users" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;. StudioMeyer is an AI-first digital studio building premium websites and intelligent automation for businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>chatgpt</category>
      <category>codex</category>
      <category>openai</category>
    </item>
    <item>
      <title>Your Top 10 Ranking Stopped Delivering Clicks</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Thu, 13 Aug 2026 12:35:41 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/your-top-10-ranking-stopped-delivering-clicks-30g0</link>
      <guid>https://dev.to/studiomeyer_io/your-top-10-ranking-stopped-delivering-clicks-30g0</guid>
      <description>&lt;p&gt;A page that holds a top ten spot and gets almost no clicks used to mean something was wrong with it. Bad title, weak description, mismatched intent. You rewrote the snippet, waited a few weeks, and the clicks came.&lt;/p&gt;

&lt;p&gt;I ran exactly that repair on one of our own pages in May. New title, sharper and more specific, written against the search intent rather than the topic. Then I left it alone for two months and measured. The click-through rate stayed where it was, around a tenth of a percent, while the page held its position and kept collecting impressions.&lt;/p&gt;

&lt;p&gt;That result was more useful than a success would have been. It ruled out the explanation everyone reaches for first, and it forced me to look at what had actually changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Position Is Fine. The Result Page Is Not.
&lt;/h2&gt;

&lt;p&gt;For a growing share of informational searches, Google answers the question above the results. The user reads the summary and never scrolls. Your listing is still there, still ranked, still counted as an impression. It simply sits below the point where the search ends.&lt;/p&gt;

&lt;p&gt;I should be precise about what I can and cannot prove here. Search Console gives you position and clicks; it does not tell you which result page carried an answer box. What I have is a strong correlation on our own data plus the failed title test, and the explanation that fits both. If you want certainty for your own pages, search the affected queries yourself and look at what sits above the first organic result. That takes ten minutes and beats any inference.&lt;/p&gt;

&lt;p&gt;For the page I tested I did run that check, and what sits above the organic results explains the flat click line better than anything on the page does. So for this one case I am comfortable saying it was neither a penalty nor a content problem. For your pages that remains a hypothesis until you look.&lt;/p&gt;

&lt;p&gt;What makes it hard to see is that every dashboard still reports the old metrics as if they meant the old things. Position two looks like a win. Impressions climbing month over month looks like growth. Only the click column tells you the truth, and it is the one column people explain away.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Recognise It in Your Own Data
&lt;/h2&gt;

&lt;p&gt;There is a signature, and you can check it in Search Console in about ten minutes without any tools.&lt;/p&gt;

&lt;p&gt;Look for pages with a high impression count, a stable position in the top ten, and a click-through rate near zero. Not low, near zero. A weak title still produces a real rate, just a disappointing one. What I am describing produces something closer to nothing at all, across hundreds of impressions, month after month, without moving.&lt;/p&gt;

&lt;p&gt;Then look at what those queries actually are. On our own pages the affected ones were questions. Definitional searches, "what is" and "how does", the exact shape of query that an answer box handles completely. The commercial queries on the same domain still converted impressions into clicks at an ordinary rate, because a summary does not finish that job.&lt;/p&gt;

&lt;p&gt;That split is the useful part. It tells you which of your pages are affected and which are not, and it stops you from applying the wrong fix to the whole site. I would not assume our split generalises to your site, but the check takes ten minutes and tells you your own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Failed Repair Was Worth More Than a Fix
&lt;/h2&gt;

&lt;p&gt;I want to stay on the title rewrite for a moment, because the way it failed is the most useful thing in this article.&lt;/p&gt;

&lt;p&gt;The standard advice would have had me rewrite, see no improvement, and conclude the new title was not good enough either. Then rewrite again. That loop can absorb months, and it never terminates, because you can always imagine a better headline.&lt;/p&gt;

&lt;p&gt;What stopped it was writing down beforehand what a success would look like. A meaningful lift within eight weeks on a page whose position was already stable. When the eight weeks produced nothing, the hypothesis was dead and I had to look elsewhere. Without that number written down in advance, I would have squinted at a flat line and found something encouraging in it.&lt;/p&gt;

&lt;p&gt;This applies well beyond titles. Most SEO work is untestable in practice because nobody says in advance what would count as the measure failing. If your agency proposes a change, ask what result would prove the change did not work, and by when. A proposal that cannot fail cannot be evaluated, and you will pay for it either way.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Things That Actually Move
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Look at other search engines, because they did not all change at once.&lt;/strong&gt; This was the finding that surprised me most. I took a page whose Google click-through rate had collapsed and pulled the same query set from Bing Webmaster Tools. Same page, same search terms, roughly four percent click-through rate against a Google figure that rounds to zero.&lt;/p&gt;

&lt;p&gt;Bing is smaller, and for most businesses it will not replace Google traffic. But it is not nothing, and the effort required is close to zero because your pages already rank there. In the client accounts I have looked at, a claimed Bing Webmaster property is the exception rather than the rule, which means the channel is usually not even being measured. Whether that traffic converts as well as Google's is something I have not measured and would not claim.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Push the position instead of polishing the snippet.&lt;/strong&gt; If the answer box is taking the click, cosmetic changes below it are wasted. The lever that remains is moving genuinely higher, and the cheapest candidate is internal linking, because most sites have valuable pages carrying a single incoming link from their own navigation and nothing else. I would call this the best available bet rather than a proven fix: it is cheap, it is under your control, and unlike a title rewrite it addresses the thing that actually changed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write to be the source, not to win the click.&lt;/strong&gt; If the summary is going to be read either way, the question becomes whether it is built from your material. That means content an answer engine can lift cleanly: a direct answer near the top, specific numbers, clear structure, and something genuinely yours in it, a measurement or a case that exists nowhere else. The reasoning is simple enough that it does not need a study behind it. A summariser has to choose what to quote, and material that exists in fifty other places gives it no reason to name you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to Stop Measuring the Old Way
&lt;/h2&gt;

&lt;p&gt;Rankings as a standalone success metric no longer survive contact with reality. A report showing a climb from position eight to position three is not a result if the click column did not move, and any agency handing you that report without the click column next to it is showing you the wrong number.&lt;/p&gt;

&lt;p&gt;The same applies to impressions. Impressions rise when Google shows you more often, including on queries that get answered above you. Rising impressions with flat clicks is not momentum, it is the pattern described in this article.&lt;/p&gt;

&lt;p&gt;What I would put in a report instead: clicks by intent category, so commercial and informational pages are never averaged together. Click-through rate compared against your own history rather than an industry benchmark. And traffic from every search source you have, not just the biggest one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Answer Searches and Buying Searches Need Different Pages
&lt;/h2&gt;

&lt;p&gt;The practical consequence of all this is that one content strategy no longer covers both halves of your site.&lt;/p&gt;

&lt;p&gt;Answer searches, the "what is" and "how does" queries, are the ones being absorbed. Pages built for them still have a job, but the job changed. They are now there to establish that you know the subject, to be quoted, and to catch the small share of readers who want more than a summary. Measuring them by sessions will make them look like failures even when they are working.&lt;/p&gt;

&lt;p&gt;Buying searches held up in our data and I expect them to hold up longer, for a structural reason rather than a measured one. Somebody looking for a supplier in their region, comparing options, or ready to make contact still needs to arrive somewhere and speak to someone, and a summary cannot complete that. These pages deserve the disproportionate share of your effort now, and for most small companies they are the underbuilt half, because informational content is easier to produce and feels more productive.&lt;/p&gt;

&lt;p&gt;If you only change one thing after reading this, separate those two page types in whatever report you look at. The moment they stop being averaged together, most of the confusion about whether SEO still works resolves itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Uncomfortable Part
&lt;/h2&gt;

&lt;p&gt;If most of your traffic came from informational content, you should plan on the assumption that a part of it is not coming back. I cannot prove that as a permanent state, and anyone who tells you they can is guessing with more confidence than I have. But planning around it costs you nothing if you are wrong, and the alternative is spending another year improving titles on pages whose problem sits above them on the result page.&lt;/p&gt;

&lt;p&gt;The useful response is to move the effort to where clicks still exist. Commercial and local searches, where a summary cannot finish the transaction. Channels you already rank in but never measured. And content strong enough to be quoted, which builds recognition even when it does not build sessions.&lt;/p&gt;

&lt;p&gt;None of that is as satisfying as a ranking chart pointing up. It has the advantage of being true.&lt;/p&gt;

&lt;h2&gt;
  
  
  One Test You Can Run This Week
&lt;/h2&gt;

&lt;p&gt;Open Search Console, filter to the last three months, and sort your pages by impressions. Take the ten with the most impressions and look only at the click column. Several near zero while holding a top-ten position gives you your candidate list, not your diagnosis. Take three of those queries, search them, and look at what occupies the space above the first organic result. That is the step that turns a suspicion into an answer, and it is the one everybody skips.&lt;/p&gt;

&lt;p&gt;Then open Bing Webmaster Tools, add the same site, and give it long enough to collect a meaningful sample. Compare the click-through rate on the same queries. If the gap looks anything like the one I measured, you have a channel that is already sending you traffic and that nobody on your side is measuring or working on. Claiming it costs you an afternoon.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/top-10-ranking-keine-klicks" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>marketing</category>
      <category>ai</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Agentic AI in German: The Words, the Law, the Numbers</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Mon, 10 Aug 2026 22:32:52 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/agentic-ai-in-german-the-words-the-law-the-numbers-33bf</link>
      <guid>https://dev.to/studiomeyer_io/agentic-ai-in-german-the-words-the-law-the-numbers-33bf</guid>
      <description>&lt;p&gt;German does not have a word for agentic, and the workaround the industry settled on is a bad one.&lt;/p&gt;

&lt;p&gt;"Agentisch" exists as a loan translation and appears in analyst reports, but nobody says it out loud in a meeting. What people actually say is KI-Agent, which means AI agent, the noun. So the German conversation collapses the adjective into the object, and a distinction that is load-bearing in English quietly disappears. In English you can say a system is somewhat agentic. In German you either have a KI-Agent or you do not, and that binary is a genuinely bad fit for a technology that is a spectrum.&lt;/p&gt;

&lt;p&gt;I bring this up because it is not a language-nerd observation. It shows up in procurement. When a German company writes a specification for "einen KI-Agenten", the document almost never states how much decision-making is being handed over, because the word does not carry that dimension. The supplier then delivers at whichever level of autonomy suits them, and both sides think they agreed on something.&lt;/p&gt;

&lt;p&gt;Three things make agentic AI in the German-speaking market different from the discourse you read in English: the vocabulary is worse, the law is further along, and the adoption numbers are more interesting than the coverage suggests. This piece covers all three.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Terms, and What Each One Actually Says
&lt;/h2&gt;

&lt;p&gt;The German market runs five terms in parallel, and they are not synonyms even though they get used as though they were.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agentic AI.&lt;/strong&gt; Usually left in English. Best understood as an adjective describing degree of self-direction. Correct usage is comparative: more agentic, less agentic, agentic at these three points in the process.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agentische KI.&lt;/strong&gt; The literal translation. Grammatically correct, in circulation, and slightly awkward in speech. Useful in written specifications precisely because it keeps the adjective intact and forces the question of how much.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;KI-Agent.&lt;/strong&gt; The concrete system. This is what people say. It refers to the thing, not the property, and it says nothing about autonomy level. A system where a model picks between three prescribed routes is a KI-Agent. So is one that plans its own path across six systems. The term does not distinguish them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Autonome KI-Systeme.&lt;/strong&gt; Common in the compliance and IT security world. Emphasises independence of action, which makes it the term risk officers reach for. It also overstates most real deployments, because almost nothing in production is genuinely autonomous.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;KI-Mitarbeiter.&lt;/strong&gt; Literally AI employee. Popular in marketing, and I would avoid it entirely. It implies an employment relationship that does not exist, it invites people to expect judgement the system does not have, and since 2 August 2026 it collides with EU transparency rules if the system is presented to customers as a person.&lt;/p&gt;

&lt;p&gt;If you are writing a brief, the useful move is to skip the noun and describe the behaviour. Which decisions does the system make on its own. Which ones need a person. What happens when it gets one wrong. Three sentences that no German label carries, and that determine the entire cost of the project.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the German Numbers Actually Say
&lt;/h2&gt;

&lt;p&gt;The coverage of AI adoption in Germany oscillates between "the Mittelstand is asleep" and "everything is transforming". The data supports neither.&lt;/p&gt;

&lt;p&gt;The AI index for the German Mittelstand, produced by Salesforce together with the Deutscher Mittelstands-Bund and published in March 2026, put the share of mid-sized companies using or testing AI at 51.2 percent, up from 33.1 percent a year earlier. That is a rise of 54 percent in twelve months, and it means the majority tipped over for the first time.&lt;/p&gt;

&lt;p&gt;The agent-specific figure is the one worth writing down. AI agents were in use at 16.6 percent of surveyed companies, against 8.7 percent the year before. Nearly doubled. A further 37 percent said they planned to introduce or expand AI during 2026, up from 25 percent at the end of 2024.&lt;/p&gt;

&lt;p&gt;Bitkom's survey adds the harder edges. Its 2026 study, based on telephone interviews with 604 German companies of 20 or more employees conducted in the opening weeks of the year, found 41 percent using AI actively, against 17 percent the year before. Another 48 percent were planning or discussing it. Only 11 percent ruled it out.&lt;/p&gt;

&lt;p&gt;Then the parts that get quoted less. Only 21 percent of those companies had an AI strategy. A third said AI was costing more than they expected. And the single most-named obstacle, at 41 percent, was uncertainty about data protection.&lt;/p&gt;

&lt;p&gt;Read together, the picture is specific and it is not the cliché. German companies are not refusing this technology. They are adopting it faster than the coverage suggests, mostly without a strategy, and the thing slowing them down is not doubt about whether it works. It is not knowing what happens to their data and who is answerable when the system is wrong. Those are reasonable questions. They also happen to be answerable, which is why the companies that get a straight answer move quickly and the ones that get marketing stay stuck.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Legal Clock Is Further Along Than Most Realise
&lt;/h2&gt;

&lt;p&gt;On 2 August 2026, the transparency obligations under Article 50 of the EU AI Act came into force. They were not deferred.&lt;/p&gt;

&lt;p&gt;The operative parts for anyone running an agent are short. Where a person interacts with an AI system, this has to be disclosed, unless it is obvious from the circumstances to a reasonably observant person. Synthetic image, audio and video content has to be marked as artificially generated.&lt;/p&gt;

&lt;p&gt;That is a modest requirement and it has one sharp consequence: the agent that introduces itself as a named human colleague is now a compliance problem in the EU, not a clever design choice. If your support agent is called Lisa and customers believe Lisa is a person, you have work to do.&lt;/p&gt;

&lt;p&gt;The rest of the regulation moved in the other direction. The Digital Omnibus, agreed politically between Parliament and Council on 7 May 2026, deferred obligations for high-risk systems under Annex III to 2 December 2027, and those covering AI as a safety component in regulated products to 2 August 2028. Rules for general-purpose AI models have applied since August 2025.&lt;/p&gt;

&lt;p&gt;The German angle here is that high-risk categories catch more agent deployments than people expect. Recruitment and employment decisions are in there. So is access to essential services, and parts of education. An agent that ranks job applicants sits in a different regulatory world from one that sorts incoming email, even when the two are built from the same components. For the full timeline and what each tier requires, we wrote that up separately when the omnibus landed, and it is &lt;a href="https://studiomeyer.io/en/blog/eu-ai-act-2026-after-the-omnibus" rel="noopener noreferrer"&gt;the more thorough treatment of the deadlines&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The practical takeaway for anyone deploying now: disclosure is due today, the heavy compliance work has more runway than the original text implied, and the classification of your specific use case matters far more than the technology you built it with.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Changes When the Agent Works in German
&lt;/h2&gt;

&lt;p&gt;This part gets almost no attention in English-language material, and it is where German deployments genuinely differ.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The address problem has no default.&lt;/strong&gt; Every German-language system has to decide between Sie and du, and there is no neutral option the way there is in English. Worse, the correct answer is not a company-wide setting. A B2B agent writing to a Geschäftsführer uses Sie. The same company's Instagram bot answering a 24-year-old uses du. Get it wrong in the formal direction and you sound distant. Get it wrong in the informal direction and you sound like you do not know who you are talking to, which in German business correspondence reads as a real error rather than a stylistic one. Every agent we build in German gets this decided explicitly before anything else, because it is not recoverable after the fact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Text costs more.&lt;/strong&gt; German compound nouns and longer average word length mean the same content consumes noticeably more tokens than its English equivalent. Since agents are billed by token and an agent loop re-reads its accumulated context on every pass, that difference compounds across a long task rather than staying flat. It does not change what is possible. It does change your cost model, and it is a reason to keep the loop short and the context tight in German deployments specifically.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The variants are real.&lt;/strong&gt; Swiss German writing does not use ß. Austrian usage differs in vocabulary and in month names. An agent writing to customers across the DACH region either handles that or produces text that reads as slightly foreign in two of its three markets.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inbound German is messier than outbound.&lt;/strong&gt; The agent writes clean German. Customers do not. Real incoming messages arrive with dialect, with regional vocabulary, with the compressed grammar people use on WhatsApp, and increasingly from people whose first language is not German. Models handle this well now, but it is worth testing with your own real messages rather than sample text, because the failure mode is not a wrong answer. It is a confident answer to a misread question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Official language is its own dialect.&lt;/strong&gt; Anything touching Behörden, insurers or the tax world runs on a formal register with fixed phrasings, and an agent that writes friendly modern German into that context produces documents that look unserious to the recipient. This is a solvable problem and it needs to be solved deliberately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where German Deployments Actually Sit
&lt;/h2&gt;

&lt;p&gt;Across what I see in the market and what the surveys report, the German cluster is narrower than the international one and it makes sense.&lt;/p&gt;

&lt;p&gt;Customer enquiries that arrive in five different formats and need routing and a first answer. Document work, meaning invoices, delivery notes and orders that mostly follow a pattern and occasionally do not. Research and preparation, where somebody has to gather material from several places before a decision. Monitoring, where something has to be watched and reported without a person checking it every hour.&lt;/p&gt;

&lt;p&gt;What is conspicuously rare is the fully autonomous customer-facing agent that closes cases without a human. Partly caution, partly liability, and partly that the German market punishes visible errors harder than the American one does. That is not backwardness. Deployed at level two or three with a human gate, these systems return real time, and they do it without creating a compliance problem the company then has to unwind.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Position Worth Taking
&lt;/h2&gt;

&lt;p&gt;The vocabulary gap is going to persist, because "agentisch" is not going to catch on and no better word is coming. The way around it is to stop arguing about the label and specify the behaviour instead.&lt;/p&gt;

&lt;p&gt;That is how we work. We look at the actual task with the person doing it today, including the exceptions they handle without noticing. We decide explicitly how much the system gets to decide and where a human has to sign off. We test against real cases from the customer's own history rather than invented ones, because real German customer messages are messier than any test set. We settle Sie or du before a line of copy is written. And we host where the customer needs it, including on their own hardware in Germany, because for a good share of the companies in that 41 percent data-protection figure, that is the whole question.&lt;/p&gt;

&lt;p&gt;We run our own operation on these agents in three languages daily, German included, and I would apply that test to any supplier in this market. Ask whether they run what they are selling you, in your language, on their own business. The answer sorts the field quickly.&lt;/p&gt;

&lt;p&gt;The interesting number is not the 16.6 percent already running agents. It is the 37 percent who said they would start or expand this year. Most of them will get their definition of "KI-Agent" from whoever writes their first proposal. &lt;a href="https://studiomeyer.io/en/services/ki-systeme/agenten" rel="noopener noreferrer"&gt;Worth making sure that definition is written by someone who will still be answering for it in a year&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/agentic-ai-deutsch" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Ahrefs API Units: What 1,100 Calls Actually Cost</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Sun, 09 Aug 2026 19:22:59 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/ahrefs-api-units-what-1100-calls-actually-cost-2m35</link>
      <guid>https://dev.to/studiomeyer_io/ahrefs-api-units-what-1100-calls-actually-cost-2m35</guid>
      <description>&lt;p&gt;On the last morning of my Ahrefs subscription I burned through 400,000 API units before lunch. That was deliberate, since the budget resets monthly and expires with it, so the choice was spend it or lose it. But it produced something I had not expected to find useful: a log of 1,102 calls across those final two days, with the exact unit cost of each one attached.&lt;/p&gt;

&lt;p&gt;Ahrefs documents the pricing model in one line. Rows times fields, minimum fifty units per billable request. That is accurate and almost useless for planning, because it tells you nothing about which tool will quietly consume your month. So I measured the 41 I actually used, out of roughly 130 the server exposes.&lt;/p&gt;

&lt;p&gt;The short answer is that three tools ate 78 percent of everything, and the free and flat-rate surfaces together returned about a hundred times more rows per unit spent than the three expensive ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the Billing Actually Works
&lt;/h2&gt;

&lt;p&gt;Four mechanics, and you need all four to predict a cost.&lt;/p&gt;

&lt;p&gt;Every billable request costs at least fifty units, no matter how little comes back. Some endpoints stop there: the two Site Audit tools charged a flat 50 per call in my log no matter how many rows came back, which is why they end up so cheap per row. The row-priced endpoints, meaning most of Site Explorer and Keywords Explorer, go further and charge per row returned, multiplied by the columns you selected. And some columns cost dramatically more than others: &lt;code&gt;volume&lt;/code&gt;, &lt;code&gt;sum_traffic&lt;/code&gt;, &lt;code&gt;keyword_difficulty&lt;/code&gt; and &lt;code&gt;traffic_domain&lt;/code&gt; add roughly ten units per row each, while a middle tier including &lt;code&gt;refdomains&lt;/code&gt; adds about five.&lt;/p&gt;

&lt;p&gt;The arithmetic that follows is unforgiving. A request for 120 rows with three premium columns selected costs 120 times 31, which is 3,720 units. The same 120 rows without those columns costs around 120. Same query, same shape, a factor of thirty in price, entirely decided by the &lt;code&gt;select&lt;/code&gt; parameter.&lt;/p&gt;

&lt;p&gt;There is a fifth mechanic that only shows up under stress. During an API incident my usage counter climbed by roughly 24,500 units while not a single row arrived. Those units were later credited back, but the refund landed hours after the fact. So the counter is not trustworthy during an outage, and abandoning a run because the budget appears to be draining can be the wrong call. What is definitely wrong is retrying in a loop while the server returns internal errors.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Measured Numbers
&lt;/h2&gt;

&lt;p&gt;Sorted by value received, cheapest per row at the top. This is from the &lt;code&gt;_units&lt;/code&gt; field of 1,102 real calls, not from documentation.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Calls&lt;/th&gt;
&lt;th&gt;Units&lt;/th&gt;
&lt;th&gt;Rows&lt;/th&gt;
&lt;th&gt;Units/row&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;all zero-cost endpoints (Search Console, management, free DR)&lt;/td&gt;
&lt;td&gt;214&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5,940&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-audit-page-explorer&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;55&lt;/td&gt;
&lt;td&gt;2,750&lt;/td&gt;
&lt;td&gt;9,789&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.3&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-audit-issues&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;650&lt;/td&gt;
&lt;td&gt;2,249&lt;/td&gt;
&lt;td&gt;0.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;keywords-explorer-volume-history&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;td&gt;2,362&lt;/td&gt;
&lt;td&gt;1,181&lt;/td&gt;
&lt;td&gt;2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-domain-rating-history&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;1,366&lt;/td&gt;
&lt;td&gt;506&lt;/td&gt;
&lt;td&gt;2.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-pages-history&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;1,366&lt;/td&gt;
&lt;td&gt;471&lt;/td&gt;
&lt;td&gt;2.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-keywords-history&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;1,232&lt;/td&gt;
&lt;td&gt;308&lt;/td&gt;
&lt;td&gt;4.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-refdomains-history&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;4,098&lt;/td&gt;
&lt;td&gt;683&lt;/td&gt;
&lt;td&gt;6.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-anchors&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;28,971&lt;/td&gt;
&lt;td&gt;3,096&lt;/td&gt;
&lt;td&gt;9.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-broken-backlinks&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;td&gt;3,305&lt;/td&gt;
&lt;td&gt;301&lt;/td&gt;
&lt;td&gt;11.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;serp-overview&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;8,022&lt;/td&gt;
&lt;td&gt;492&lt;/td&gt;
&lt;td&gt;16.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-referring-domains&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;136,865&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;8,348&lt;/td&gt;
&lt;td&gt;16.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-all-backlinks&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;40,375&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2,125&lt;/td&gt;
&lt;td&gt;19.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;keywords-explorer-matching-terms&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;250&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;414,542&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;19,862&lt;/td&gt;
&lt;td&gt;20.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;keywords-explorer-related-terms&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;63&lt;/td&gt;
&lt;td&gt;27,977&lt;/td&gt;
&lt;td&gt;1,287&lt;/td&gt;
&lt;td&gt;21.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-organic-keywords&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;70&lt;/td&gt;
&lt;td&gt;38,099&lt;/td&gt;
&lt;td&gt;1,228&lt;/td&gt;
&lt;td&gt;31.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-top-pages&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;29&lt;/td&gt;
&lt;td&gt;10,978&lt;/td&gt;
&lt;td&gt;298&lt;/td&gt;
&lt;td&gt;36.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;site-explorer-metrics-history&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;12,628&lt;/td&gt;
&lt;td&gt;308&lt;/td&gt;
&lt;td&gt;41.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;keywords-explorer-overview&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;4,940&lt;/td&gt;
&lt;td&gt;110&lt;/td&gt;
&lt;td&gt;44.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gsc-anonymous-queries&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;39&lt;/td&gt;
&lt;td&gt;2,088&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;298&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;all remaining tools&lt;/td&gt;
&lt;td&gt;127&lt;/td&gt;
&lt;td&gt;16,988&lt;/td&gt;
&lt;td&gt;862&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,102&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;759,602&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;59,451&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Three Tools Ate Everything
&lt;/h2&gt;

&lt;p&gt;Keyword expansion, referring domains and raw backlinks together consumed 591,782 units, which is 78 percent of the total. Everything else, all 38 remaining tools, came to 167,820.&lt;/p&gt;

&lt;p&gt;The single worst line has a cause I can name precisely. I ran keyword expansion at a limit of 250 rows across a batch of seed terms, and a call at that limit cost 5,250 units. At a limit of 50 the same call costs 1,050 and produces the same conclusion, because rows 51 through 250 were fragments and near-duplicates I never looked at again.&lt;/p&gt;

&lt;p&gt;Why 250? Because that is the maximum rows per request on the plan I was on. I did not choose it as an analytical decision, I reached for the ceiling because it was there. That turns out to be the most expensive habit available, and it got worse in April when Ahrefs raised the row caps across all tiers. Lite went from 10 rows to 100, Standard from 25 to 250, Advanced from 100 to 500. An explicit &lt;code&gt;limit: 50&lt;/code&gt; in your code is unaffected by that change. What is affected is any call that passes no limit, any call that asks for the current maximum, and any call that used to hit the old cap and now returns up to ten times more rows for up to ten times the price.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Other Side of the Ledger
&lt;/h2&gt;

&lt;p&gt;Everything tied to a verified project of your own is free or close to it.&lt;/p&gt;

&lt;p&gt;Every Search Console endpoint I used cost zero, with one exception I come back to below. Together with the management endpoints and the free domain rating lookup, 214 calls returned 5,940 rows for nothing, and for domains I own those rows are more honest than the paid estimates, because they are measured rather than modelled. Site Audit is nearly as good: both audit endpoints bill a flat 50 units per request regardless of how many rows come back, which worked out to 0.3 units per row across nearly 10,000 pages, with more than twenty technical fields per URL.&lt;/p&gt;

&lt;p&gt;Set the two groups next to each other. The free and flat-rate surfaces together returned roughly 18,000 rows for 3,400 units, which is 5.3 rows per unit. The three expensive tools returned 30,000 rows for 592,000 units, or 0.05 rows per unit. That is a ratio of about 103 to 1. The comparison only works for the combined group, because the genuinely free endpoints have no per-unit rate at all.&lt;/p&gt;

&lt;p&gt;That ratio is the whole argument. Not "use fewer calls" but "use the other tools first", because in most cases they answer the same question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Places the Common Numbers Are Wrong
&lt;/h2&gt;

&lt;p&gt;Measuring turned up three costs that do not match what gets repeated in guides, including in my own notes before this.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;serp-overview&lt;/code&gt; is widely quoted at around 180 units per keyword. Across my whole log it averaged &lt;strong&gt;401&lt;/strong&gt; units per call, and a separately measured call for a fresh term at &lt;code&gt;limit: 20&lt;/code&gt; came in at about &lt;strong&gt;544&lt;/strong&gt;. The gap between those two numbers is most likely cached repeats pulling the average down, which is exactly why an average is the wrong number to budget with. Plan a competitive SERP sweep at 544 per term, so twenty terms is around 11,000 units and not the 3,600 the common figure implies.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;gsc-anonymous-queries&lt;/code&gt; is the worst value in the entire toolkit for a small site, and its name hides that. Everything else with a &lt;code&gt;gsc-&lt;/code&gt; prefix is free, so it reads as free. It is not: the documented floor is 50 units per call, and measured it averaged 53.5. Thirty-nine calls returned seven rows in total, because the small sites in my sample never crossed the anonymisation threshold that endpoint exists to reveal. On a large client domain it may well pay for itself. Before using it at scale, spend one call and see whether anything comes back.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;site-explorer-metrics-history&lt;/code&gt; costs about 1,804 units per call, which makes it one of the most expensive requests in the toolkit, though not the outright worst: &lt;code&gt;site-explorer-referring-domains&lt;/code&gt; averaged 3,699 per call. The four individual history endpoints, domain rating, pages, referring domains and keywords, cost between 2 and 6 units per row and together average around 486 units, which is roughly a quarter of the bundled call for nearly the same story.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Rule That Matters More Than Any of This
&lt;/h2&gt;

&lt;p&gt;Cost discipline is the small lesson. Here is the expensive one.&lt;/p&gt;

&lt;p&gt;An early pass of one keyword analysis reported just under half a million reachable monthly searches. It was a real number produced by real API calls, and it was garbage. Keyword expansion drags in fragments and generic words, so single prepositions and bare nouns were sitting in the list being counted as opportunities.&lt;/p&gt;

&lt;p&gt;Second attempt, filtered to terms carrying a location: roughly a fifth of the original. Better, still wrong. The top of that list was pure place names, which are travel searches with no commercial intent whatsoever.&lt;/p&gt;

&lt;p&gt;Only a double filter produced something defensible. A term had to carry both a location and a signal of what was actually being sold. What survived was about a fifteenth of the first figure, a low five-figure monthly volume across roughly a hundred terms, and it was the first honest number in the sequence.&lt;/p&gt;

&lt;p&gt;A volume sum over an unfiltered keyword list is not a number, it is a claim wearing a number's clothing. And it is dangerous precisely because it survives review: it came from an API, it has no rounding, it looks like data. Nobody questions a figure with six digits and no rounding. Before you sum anything, every row has to pass one test. Would a person searching this become a customer? If you cannot answer that per row, do not add the column up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Six Rules I Would Give My Past Self
&lt;/h2&gt;

&lt;p&gt;Set &lt;code&gt;limit: 50&lt;/code&gt; as the default on keyword expansion and raise it only when a result visibly clipped at the boundary and the extra rows change a decision.&lt;/p&gt;

&lt;p&gt;Exhaust the free surfaces first. Search Console for every connected project, the free domain rating lookup, management endpoints, rank tracker. On your own domains these are not the cheap option, they are the better data.&lt;/p&gt;

&lt;p&gt;Reach for the flat-rate audit tools before the row-priced ones. Fifty units for up to 250 rows with twenty-plus fields each is the best value in the product by a wide margin.&lt;/p&gt;

&lt;p&gt;Strip premium columns unless a decision depends on them. Ten units per row each, and intent columns in particular tell you what the words already say.&lt;/p&gt;

&lt;p&gt;Calculate instead of estimating before any block large enough to hurt. My own threshold is 20,000 units, which is arbitrary but forces the arithmetic: rows times one, plus ten per premium column, plus five per middle column. It takes a minute and it is the difference between a planned spend and a surprise.&lt;/p&gt;

&lt;p&gt;Put the brake in the script, not in your head. A budget check before each block that aborts below a reserve you set in advance. I wrote mine after the fact, which is exactly the wrong order and the reason this article has such precise numbers in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Happens Now
&lt;/h2&gt;

&lt;p&gt;I did not renew the subscription. The measurement above is part of why: once you can see which surfaces carry the value, it becomes obvious that Search Console covers most of what I needed for my own domains, and that the paid index was earning its keep on exactly two jobs, competitor backlink profiles and search volume for terms I do not rank for yet.&lt;/p&gt;

&lt;p&gt;That is a specific conclusion for a specific situation and I would not generalise it. What does generalise is the method. Every Ahrefs response carries its real cost inline. Log that field, group by tool, divide by rows returned, and you will know more about your own usage than any pricing page can tell you. Two days of doing it taught me more than a year of not doing it, and I only did it at the end, when the answer could no longer change what I bought.&lt;/p&gt;

&lt;p&gt;If you are still on a plan, do it now.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/ahrefs-api-units-cost" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>api</category>
      <category>data</category>
      <category>mcp</category>
    </item>
    <item>
      <title>AI Can Build Your Website. Done Is Still a Claim.</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Sun, 09 Aug 2026 02:35:13 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/ai-can-build-your-website-done-is-still-a-claim-4a0p</link>
      <guid>https://dev.to/studiomeyer_io/ai-can-build-your-website-done-is-still-a-claim-4a0p</guid>
      <description>&lt;p&gt;AI will build you a website in twenty minutes. Clean structure, decent layout, usable copy. That is not a promise anymore, it works. Anyone claiming otherwise has not tried it in over a year.&lt;/p&gt;

&lt;p&gt;And it does more than build. Wired up properly, it also puts the site online, points your address at it and ships changes. That is not science fiction either.&lt;/p&gt;

&lt;p&gt;Then it says: done.&lt;/p&gt;

&lt;p&gt;That is the moment this article is about. Because "done" means one thing here: I have done everything I was able to do. It does not mean: I checked whether it works. Checking is the part it cannot do. It sees its own work, not what eventually reaches your customer.&lt;/p&gt;

&lt;p&gt;And that does not improve by paying for a bigger model. We have had pages reviewed by an AI several times over, thoroughly, with instructions to find faults. Afterwards a simple program ran across them that does nothing except look at the finished page the way a visitor receives it. It found faults on the first pass that every review round before it had waved through. Not because it is smarter. Because it stands somewhere else.&lt;/p&gt;

&lt;p&gt;That is the whole difference. Not who builds. Who looks afterwards, and from where.&lt;/p&gt;

&lt;p&gt;Spend enough time around this and you notice the problems keep taking the same shape. Four of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Faults That Never Announce Themselves
&lt;/h2&gt;

&lt;p&gt;We are used to computers turning something red when it breaks. With a website that is mostly not the case.&lt;/p&gt;

&lt;p&gt;The most common example: someone fills in your contact form. The screen says thank you, we will be in touch. The message never reaches you. Nobody finds out. Not you, because nothing arrives. Not the person who wrote it, because they saw the confirmation. They wait two days and go to the next one.&lt;/p&gt;

&lt;p&gt;The tricky part is not the fault. The tricky part is that a form losing enquiries looks exactly like a form nobody is using. A quiet week and a broken line feel identical. And nobody calls to tell you they could not reach you.&lt;/p&gt;

&lt;p&gt;You do not find faults like that by waiting for them to show up. You find them because somebody goes looking while nothing suggests anything is wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Faults That Arrive Later
&lt;/h2&gt;

&lt;p&gt;A website has things with an expiry date. They renew themselves, and usually that works.&lt;/p&gt;

&lt;p&gt;When one of those automatic renewals falls out of step, it tells nobody. It keeps trying, day after day, failing, and everything looks normal. Until the day the deadline actually lands. Then your site no longer opens for visitors. Not slowly, not with a warning. It does not open.&lt;/p&gt;

&lt;p&gt;Months can pass between the moment the fault appears and the moment you notice. Checking on launch day means checking the one day when everything still works.&lt;/p&gt;

&lt;p&gt;And there is a trap of reasoning in here that rarely occurs to anyone. That something is valid today does not prove the renewal works. It only proves it worked last time. Check the state instead of the process and you are checking the wrong thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Assumptions That Look Like Facts
&lt;/h2&gt;

&lt;p&gt;Building involves constant assumption. That is normal and mostly harmless. These photos are all the same size. This page is called what its heading says. It probably runs for other people roughly like it runs for me.&lt;/p&gt;

&lt;p&gt;The problem is that a finished assumption looks exactly like a verified fact. You cannot tell from the result whether somebody measured or whether somebody guessed.&lt;/p&gt;

&lt;p&gt;We once measured a whole set of photos that had a single size recorded for them, because that is the kind of thing one assumes. Some had different dimensions, several were portrait rather than landscape. None of it was visible anywhere. It only came out when somebody opened the files instead of believing the label beside them.&lt;/p&gt;

&lt;p&gt;The unpleasant variant: a check can enshrine an assumption too. We had a test that dutifully verified a particular value, comparing it against a number somebody had guessed beforehand. It passed, every time. It did not find the fault, it confirmed it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions Nobody Asks
&lt;/h2&gt;

&lt;p&gt;The last point is the quietest and probably the most expensive.&lt;/p&gt;

&lt;p&gt;An AI answers questions extremely well. It does not ask them. If you do not know that a particular registration exists, a particular mandatory disclosure, a particular directory your business ought to be listed in, then you do not ask about it. And then it does not happen. Not because anyone did anything wrong. It simply never came up.&lt;/p&gt;

&lt;p&gt;That is why a self built site often fails on an absence rather than a fault. Nothing is broken. Something is missing that nobody knew was missing.&lt;/p&gt;

&lt;p&gt;And no tool can take that off your hands, however good it is. A tool answers what you ask it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Follows From This
&lt;/h2&gt;

&lt;p&gt;None of this argues against AI, we work with it all day ourselves. It handled the part that can be handed to it, and it handled it well.&lt;/p&gt;

&lt;p&gt;But all four patterns have one thing in common. The fault cannot be seen in the finished result and it does not report itself. It becomes visible when somebody looks who knows where to look and when. That is not a question of care while building. It is a different activity, at a different time, from a different direction.&lt;/p&gt;

&lt;p&gt;Which is why "AI can build a website" is true and misleading at the same time. It can build the website. Whether it works is another question, and no claim answers that. Only a check does.&lt;/p&gt;

&lt;p&gt;The next part is about where this costs the most money. The site runs, it looks good, it was done properly, and still nobody finds it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/ki-website-bauen-was-danach-kommt" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;. It is the first part of a series about everything that happens after a website has been built.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>business</category>
    </item>
    <item>
      <title>What AI Agents Actually Are, and Where the Word Ends</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Sun, 09 Aug 2026 02:32:41 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/what-ai-agents-actually-are-and-where-the-word-ends-3olh</link>
      <guid>https://dev.to/studiomeyer_io/what-ai-agents-actually-are-and-where-the-word-ends-3olh</guid>
      <description>&lt;p&gt;Gartner went looking for AI agents in 2025 and found something worth remembering. Of the thousands of vendors claiming agentic capability, roughly 130 were building anything that deserved the word. Not 130 good ones out of a field of 200. One hundred and thirty out of thousands.&lt;/p&gt;

&lt;p&gt;That ratio is the reason this question is hard to answer by reading marketing pages. The term is applied to chatbots, to scheduled scripts, to autocomplete, and occasionally to something that genuinely is an agent. So it is worth starting from the technical definition rather than the sales one, because the technical definition is short and it draws a line you can actually check.&lt;/p&gt;

&lt;h2&gt;
  
  
  One Sentence, and Then the Consequences
&lt;/h2&gt;

&lt;p&gt;An AI agent is a system in which the model decides what happens next, rather than the code.&lt;/p&gt;

&lt;p&gt;That is the whole distinction. Anthropic, whose engineering team wrote the reference description of this, splits the field into two: workflows are "systems where LLMs and tools are orchestrated through predefined code paths", while agents are "systems where LLMs dynamically direct their own processes and tool usage".&lt;/p&gt;

&lt;p&gt;Sit with that for a second, because the consequences are larger than the sentence.&lt;/p&gt;

&lt;p&gt;In a workflow, a developer decided in advance that step three follows step two. The model might write the text in step two, brilliantly, but it does not get a say in whether step three happens. The path is fixed. You can draw it on a whiteboard and the drawing stays accurate.&lt;/p&gt;

&lt;p&gt;In an agent, the model is handed a goal and a set of capabilities, and it picks. It might use one tool, or six, or the same tool four times with different inputs, or none at all. Run the same task twice and you can get two different paths to the same result. You cannot draw it on a whiteboard, because the drawing changes every run.&lt;/p&gt;

&lt;p&gt;That is why agents are powerful and it is exactly why they are hard to operate. Everything else you read about agents, the reliability engineering, the guardrails, the monitoring, follows from that one property.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Such a System Is Made Of
&lt;/h2&gt;

&lt;p&gt;Strip away the vendor diagrams and there are five parts. All five have to be there, and the ones people forget are the last two.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model.&lt;/strong&gt; The part that reasons. It reads the situation and decides on the next action. It is worth understanding that the model starts fresh every single time it is called. It has no memory of the previous step except what you hand it. Everything that feels like continuity is something the surrounding system reconstructed and passed back in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tools.&lt;/strong&gt; What the agent can actually do. Read a database, send an email, create a ticket, query an API, fetch a page. Without tools you have a chat window. With tools you have something that changes state in the world, which is a different category of object entirely. The number of tools is not a virtue. Every additional tool is another thing the model can pick wrongly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The memory.&lt;/strong&gt; Short-term, meaning the running record of this task, and long-term, meaning what carries across sessions. The customer's previous three complaints. The decision made last Tuesday and why. Without long-term memory an agent greets a ten-year customer exactly like a stranger, every time, which people notice immediately and dislike intensely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The loop.&lt;/strong&gt; The cycle that keeps the whole thing moving: give the model context, receive a decision, execute it, add the result to the context, ask again. This continues until the model signals it is finished, or until a limit stops it. The limit is not optional. An unbounded loop is a system that can spend your money forever on a task it is quietly failing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The guardrails.&lt;/strong&gt; What sits between a decision and an action with consequences. Permission checks, approval gates, policy rules, spending caps, an audit trail. This is the part that separates a demo from a system somebody is willing to be responsible for, and it is the part that almost never appears in the video.&lt;/p&gt;

&lt;p&gt;There is a hard arithmetic reason the last two matter. Reliability across a chain of model decisions multiplies rather than averages. A step that is right 95 percent of the time is fine on its own and unrecognisable after twenty of them in a row. That is not a flaw to be fixed with a better model. It is a property of chaining probabilistic decisions, and the engineering answer is to shorten the chain, verify along the way, and put a gate before anything you cannot undo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where You Have Already Met One
&lt;/h2&gt;

&lt;p&gt;The abstraction gets easier with concrete cases, so here are four that are unambiguously agents rather than dressed-up workflows.&lt;/p&gt;

&lt;p&gt;A coding assistant that reads your repository, decides which files are relevant, opens them, makes a change, runs the tests, sees a failure, and goes back to fix it. Nobody scripted that sequence. The model chose to run the tests, read the output and decided the failure was worth another pass.&lt;/p&gt;

&lt;p&gt;A research assistant given a company name that decides on its own which searches to run, notices that two sources disagree, goes looking for a third to break the tie, and reports the disagreement rather than picking one at random.&lt;/p&gt;

&lt;p&gt;A support system that reads an incoming message, checks the order status in one system, checks the shipping status in another, sees the parcel is stuck at customs, and writes an answer that explains customs rather than quoting the standard delivery window. The path through those two systems was not predetermined. It depended on what the first check returned.&lt;/p&gt;

&lt;p&gt;An operations watcher that monitors a set of services, notices a pattern that is not on any alert list, correlates it with a deployment from two hours earlier, and tells a human where to look.&lt;/p&gt;

&lt;p&gt;What these have in common is the branching. At some point the system encountered a situation and picked a response that nobody wrote down in advance.&lt;/p&gt;

&lt;h2&gt;
  
  
  And Where the Word Gets Stretched
&lt;/h2&gt;

&lt;p&gt;By the same standard, several very useful things are not agents.&lt;/p&gt;

&lt;p&gt;A chatbot that answers from a document collection is not an agent, however good it is. It retrieves and it generates. It does not choose actions.&lt;/p&gt;

&lt;p&gt;A scheduled script that pulls data every night and emails a summary is not an agent, even if a model writes the summary. The path is fixed. A model doing one job inside a fixed path is a workflow with a smart step.&lt;/p&gt;

&lt;p&gt;An assistant that drafts something and waits for you to press send is not an agent in the strict sense either. It proposes. You decide. That is a copilot, and for a great many jobs it is the correct design.&lt;/p&gt;

&lt;p&gt;None of this is a criticism. Anthropic's own guidance is to find the simplest solution that works and to add agentic complexity only when it demonstrably improves the outcome, because agentic systems trade latency and cost for task performance. In practice that means a substantial share of business problems are better served by a well-built workflow than by a real agent, and a supplier who tells you that is being straight with you.&lt;/p&gt;

&lt;p&gt;The reason the labels matter anyway is money. If you are quoted agent pricing for workflow capability, you are paying for autonomy you are not receiving. One question usually settles it: given the same input twice, can the system take two different paths? If the honest answer is no, it is a workflow, and it should be priced like one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What They Still Cannot Do
&lt;/h2&gt;

&lt;p&gt;Four limits, all of which show up within weeks of a real deployment.&lt;/p&gt;

&lt;p&gt;They cannot reliably tell you what they do not know. A model that lacks a fact will often produce a plausible one instead of stopping, and an agent will then act on it. This is why grounding in your own data and marking unverified claims as unverified matter more in agent systems than in chat.&lt;/p&gt;

&lt;p&gt;They cannot carry responsibility. When an agent issues a wrong refund, the accountability sits with whoever deployed it. That is not a technical statement, it is a legal and organisational one, and it is why the approval gate exists.&lt;/p&gt;

&lt;p&gt;They do not handle unbounded scope. An agent with one clear job and four tools works. The same model with a vague brief and forty tools produces something that cannot be debugged, because the failure is never in one place.&lt;/p&gt;

&lt;p&gt;They do not survive unattended forever. Systems change, APIs move, the business changes its rules, and an agent that was correct in March is quietly wrong by September if nobody is watching it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Actually Stands in Germany
&lt;/h2&gt;

&lt;p&gt;The gap between the discourse and the deployment numbers is worth knowing.&lt;/p&gt;

&lt;p&gt;The AI index for the German Mittelstand, produced by Salesforce with the Deutscher Mittelstands-Bund and published in March 2026, found 16.6 percent of mid-sized German companies using AI agents, up from 8.7 percent a year before. A near doubling in twelve months, and still fewer than one in five. A further 37 percent said they planned to introduce or expand AI in 2026.&lt;/p&gt;

&lt;p&gt;Bitkom's survey of 604 German companies with 20 or more employees, conducted in the first weeks of 2026, found 41 percent using AI actively at all, against 17 percent the year before. But only 21 percent had an AI strategy, and 41 percent named uncertainty about data protection as their largest obstacle.&lt;/p&gt;

&lt;p&gt;Put those together and you get an accurate picture: adoption is climbing fast, understanding is lagging behind it, and the thing holding most German companies back is not scepticism about capability. It is a legitimate uncertainty about what happens to their data and who answers for the outcome. Both of which are answerable, but only by a supplier willing to answer them rather than route around them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Useful Version of the Question
&lt;/h2&gt;

&lt;p&gt;"What are AI agents" is worth answering precisely, but it is rarely the question that matters in a business. The question that matters is narrower: is there a task here where the deciding is the expensive part?&lt;/p&gt;

&lt;p&gt;If the expensive part is typing, you want automation and you do not need an agent. If the expensive part is that someone has to look at each case, weigh two or three things and pick a path, that is where an agent earns its keep. Incoming enquiries that arrive in five formats. Documents that mostly follow a pattern and sometimes do not. Research that has to be checked against several sources before anyone can act on it.&lt;/p&gt;

&lt;p&gt;That is also how we scope them. One task where the judgement is the bottleneck, a small set of tools, real cases from the customer's own history for testing, a human gate in front of anything irreversible, and someone named as responsible for it after handover. Our own agents run on that pattern daily, researching, monitoring and preparing work for us, which is the standard I would apply to any supplier: ask whether they run what they sell.&lt;/p&gt;

&lt;p&gt;The word "agent" will keep drifting, because words attached to budgets always do. The property underneath it will not. Somewhere in the system, either the code decides or the model decides. Knowing which one you are buying is most of the decision, and &lt;a href="https://studiomeyer.io/en/services/ki-systeme/agenten" rel="noopener noreferrer"&gt;it is where any serious conversation about an agent should start&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/was-sind-ki-agenten" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>beginners</category>
      <category>programming</category>
    </item>
    <item>
      <title>Four Ways to Build an AI Agent, and What Each Costs</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Thu, 06 Aug 2026 05:21:05 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/four-ways-to-build-an-ai-agent-and-what-each-costs-3nh3</link>
      <guid>https://dev.to/studiomeyer_io/four-ways-to-build-an-ai-agent-and-what-each-costs-3nh3</guid>
      <description>&lt;p&gt;On 6 October 2025, OpenAI unveiled Agent Builder at DevDay. A visual canvas where you drag boxes together and get an agent out the other end. Eight months later, on 3 June 2026, it was deprecated. It shuts down for good on 30 November 2026.&lt;/p&gt;

&lt;p&gt;Nobody did anything wrong there. OpenAI decided the visual builder was not the path and pointed users at the Agents SDK instead. But if you built your order processing on that canvas in January, you spent the summer migrating instead of shipping. That is the part of the "how do I build an AI agent" question that almost nobody asks about in advance, and it is usually the part that decides whether the thing is still running in two years.&lt;/p&gt;

&lt;p&gt;There are four honest routes to a working agent. They differ less in what they can do than in where they break. Here is what each one actually costs, where its ceiling sits, and which kind of company each one fits.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Decision You Are Actually Making
&lt;/h2&gt;

&lt;p&gt;Most comparisons start with features. That is the wrong end. Every platform on this list can call an API, read a document and write to a CRM. Feature lists converged eighteen months ago.&lt;/p&gt;

&lt;p&gt;What has not converged is the answer to four questions:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What does the agent touch?&lt;/strong&gt; An agent that drafts a reply is a different risk class from one that issues a refund. The first can be wrong and cost you nothing. The second can be wrong and cost you money, twice, because you also have to unwind it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who is liable when it is wrong?&lt;/strong&gt; Not philosophically. Operationally. Which human finds out, how fast, and what do they do about it. If the answer is "nobody, until a customer complains", the agent is not ready regardless of platform.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How often does the process change?&lt;/strong&gt; A process that changes twice a year suits a rigid platform fine. A process that changes twice a month punishes anything you cannot edit yourself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is this your competitive edge or your plumbing?&lt;/strong&gt; Nobody wins a market with better invoice routing. Buy the plumbing. Build the edge.&lt;/p&gt;

&lt;p&gt;Answer those four and the route picks itself. Skip them and you end up in the Gartner statistic: the analyst firm expects more than 40 percent of agentic AI projects to be cancelled by the end of 2027, and the three reasons it gives are escalating costs, unclear business value and inadequate risk controls. None of those are technology problems. All three are decision problems that got deferred.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route One: No-Code Builders
&lt;/h2&gt;

&lt;p&gt;Chatbase, Lindy, Relevance AI, Gumloop and a dozen others. You upload documents, paste URLs, connect a few tools through a visual editor and have something answering questions the same afternoon. Across the platforms with a published flat entry plan, the median first paid tier sits around 24 dollars per month.&lt;/p&gt;

&lt;p&gt;This route is genuinely good at one thing: proving the use case is real before anyone spends real money. I have watched a company argue for three weeks about whether an agent could handle their intake questions. A Saturday afternoon on a no-code builder ended the argument, in the direction nobody expected. Two of the four question types were trivial. The other two needed a human, permanently.&lt;/p&gt;

&lt;p&gt;The ceiling arrives faster than the sales page suggests, and it arrives in a specific place: anything that requires a decision the platform did not anticipate. You can usually connect tools. You can rarely control what happens between them. When the agent needs to check one system, decide based on what it found, then take a different path through a second system, most no-code builders either cannot express it or express it so awkwardly that the workflow becomes unmaintainable.&lt;/p&gt;

&lt;p&gt;The second ceiling is data. These are hosted services. Your customer records pass through infrastructure you do not control, under terms you did not negotiate. For a public FAQ that is fine. For anything involving personal data of EU customers, read the terms before you upload, not after.&lt;/p&gt;

&lt;p&gt;Fits: a company testing whether an agent helps at all, or running something small and public-facing that touches no sensitive data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route Two: Enterprise Suites
&lt;/h2&gt;

&lt;p&gt;Salesforce Agentforce and Microsoft Copilot Studio are the two that matter, and their pricing tells you almost everything about who they are for.&lt;/p&gt;

&lt;p&gt;Agentforce charges 2 dollars per conversation. Or 20 Flex Credits, about 10 cents, per action. Or 125 to 150 dollars per user per month for a named licence. Copilot Studio charges in Copilot Credits: 200 dollars for a prepaid pack of 25,000, which works out around eight tenths of a cent each, or a penny each pay as you go.&lt;/p&gt;

&lt;p&gt;At moderate interaction depth, a Copilot Studio conversation lands somewhere near 18 cents while an Agentforce conversation stays at 2 dollars flat. That looks like an eleven-fold gap, and every Microsoft sales deck will show you exactly that comparison.&lt;/p&gt;

&lt;p&gt;The comparison is also useless on its own, because both numbers assume you already own the thing underneath. Agentforce in practice needs Service Cloud plus a Data Cloud subscription that lists around 60,000 dollars a year. Copilot Studio assumes you already pay for the Microsoft 365 and Power Platform tenant. If you have that tenant, Copilot Studio is close to free to start. If you do not, the entry price is the whole Microsoft stack.&lt;/p&gt;

&lt;p&gt;The real trade in this route is not price. It is that you get governance for free. Audit logs, permission models, data residency, retention policy, the whole compliance apparatus that takes months to build yourself, arrives on day one and satisfies your legal department without a conversation. For a regulated company that is worth more than the licence fee.&lt;/p&gt;

&lt;p&gt;The cost is the ceiling, and it is a hard one. Both platforms are excellent inside their own ecosystem and awkward outside it. The moment your agent needs to reach a machine on your own premises, a niche piece of industry software or a database that predates the cloud, you are writing custom connectors anyway, and the platform stops paying for itself.&lt;/p&gt;

&lt;p&gt;Fits: companies that already live inside Salesforce or Microsoft 365, where the agent's work stays inside that world, and where compliance sign-off is the slow part of any project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route Three: Workflow Platforms With Agent Steps
&lt;/h2&gt;

&lt;p&gt;n8n is the clearest example, and it is the route most small and mid-sized companies I talk to end up on without planning to. The Community Edition is free and self-hosted, so you pay for a server and for the model calls, nothing else. Cloud Starter runs 20 euros a month for 2,500 executions, Pro 50 euros for 10,000.&lt;/p&gt;

&lt;p&gt;The mental model is different from the other two routes and that is the point. You draw the process as a flow with explicit steps, and at the points where judgement is needed you put a model in. Everything else stays deterministic. Read the mailbox, always. Extract the fields, always. Decide which of three departments this belongs to, model. Write to the ticket system, always.&lt;/p&gt;

&lt;p&gt;That structure quietly solves the biggest reliability problem in this whole field. Chain enough model decisions together and the maths turns against you. If each step is right 95 percent of the time, which sounds fine, twenty steps in sequence come out right about 36 percent of the time. Reliability is multiplicative and it collapses faster than intuition suggests. A flow with three model decisions and seventeen deterministic steps is a fundamentally different machine from one with twenty model decisions, even though both look like "an agent" on a slide.&lt;/p&gt;

&lt;p&gt;Self-hosting matters more in Europe than the international comparisons acknowledge. In the Bitkom survey of 604 German companies conducted in early 2026, uncertainty about data protection was the single most-named obstacle to using AI at all, cited by 41 percent. Running the orchestration on your own server does not make you compliant by itself, but it removes the question that stalls the project for six weeks.&lt;/p&gt;

&lt;p&gt;The cost is that low-code is not no-code. Expressions, error handling, retries and custom nodes all reward someone technical. A workflow platform in the hands of someone who has never debugged anything becomes a pile of half-finished flows that nobody trusts.&lt;/p&gt;

&lt;p&gt;Fits: companies with a defined process, some technical capacity in-house or on retainer, and a reason to keep data on their own infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route Four: Building It Properly
&lt;/h2&gt;

&lt;p&gt;At the far end you assemble the agent yourself. A model chosen for the task, a loop around it, tools defined as real interfaces, memory that survives restarts, guardrails that sit between the decision and the action, and an evaluation set that tells you whether last week's change made things better or worse.&lt;/p&gt;

&lt;p&gt;Industry estimates for a custom workflow agent cluster between 25,000 and 100,000 dollars, with multi-agent systems running well past that. Treat those as orders of magnitude rather than quotes, because the range depends almost entirely on how many systems the thing has to touch.&lt;/p&gt;

&lt;p&gt;What you get for it is the only route with no ceiling and no landlord. Nobody deprecates your architecture in eight months. Nobody reprices your per-conversation rate. When the process changes, you change the agent.&lt;/p&gt;

&lt;p&gt;The connective tissue that made this route cheaper than it was two years ago is the Model Context Protocol. Anthropic released it as an open standard in November 2024 and handed it to the Agentic AI Foundation under the Linux Foundation in December 2025. OpenAI, Google, Microsoft, IBM and Amazon all support it. SDK downloads went from 100,000 at launch to 97 million a month by March 2026. Practically, that means the connector you write for your ERP works with whichever model you point at it, and swapping the model later is a configuration change rather than a rewrite. That single property is why building custom no longer means betting your architecture on one vendor.&lt;/p&gt;

&lt;p&gt;The cost is that you now own an operational system. It needs monitoring, it needs someone who understands it, and it will surface edge cases in month four that nobody imagined in month one.&lt;/p&gt;

&lt;p&gt;Fits: a process that is genuinely yours, touches systems no platform knows about, handles data that cannot leave your control, or is close enough to your product that it should not be rented.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Four Routes Look Like Side by Side
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Entry cost&lt;/th&gt;
&lt;th&gt;Ceiling&lt;/th&gt;
&lt;th&gt;Breaks when&lt;/th&gt;
&lt;th&gt;Your data lives&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;No-code builder&lt;/td&gt;
&lt;td&gt;~24 USD/month&lt;/td&gt;
&lt;td&gt;Low, hits fast&lt;/td&gt;
&lt;td&gt;The decision gets conditional&lt;/td&gt;
&lt;td&gt;On their servers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise suite&lt;/td&gt;
&lt;td&gt;Stack you already own&lt;/td&gt;
&lt;td&gt;Medium, ecosystem-bound&lt;/td&gt;
&lt;td&gt;You step outside the ecosystem&lt;/td&gt;
&lt;td&gt;Their cloud, contractually&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workflow platform&lt;/td&gt;
&lt;td&gt;0 to 50 EUR/month&lt;/td&gt;
&lt;td&gt;High with technical help&lt;/td&gt;
&lt;td&gt;Nobody maintains it&lt;/td&gt;
&lt;td&gt;Wherever you host it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom build&lt;/td&gt;
&lt;td&gt;Five figures upwards&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Nobody owns it internally&lt;/td&gt;
&lt;td&gt;Wherever you decide&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The row that matters is the third one. Every route works on day one. They differ in how they fail in month nine, and the failure is almost never technical.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Most of These Projects Die
&lt;/h2&gt;

&lt;p&gt;Two numbers frame this honestly. MIT's Project NANDA found in August 2025 that 95 percent of generative AI pilots showed no measurable contribution to profit and loss. Gartner puts more than 40 percent of agentic projects on the cancellation list by the end of 2027.&lt;/p&gt;

&lt;p&gt;Read together, and against what I see in practice, the pattern is not that agents do not work. It is that companies pick a process that was never the bottleneck. Automating something that took a person twenty minutes a week produces a working agent and no measurable result. The 20 minutes go somewhere else, nobody notices, and when the budget review comes the project has nothing to point at.&lt;/p&gt;

&lt;p&gt;The counter-move is unglamorous. Pick the process where somebody is visibly drowning. Measure it before you touch it. Automate the boring middle and leave the judgement calls where they are. Then the number in the review is a real number.&lt;/p&gt;

&lt;p&gt;There is a German data point worth holding next to that. The Salesforce and Deutscher Mittelstands-Bund index published in March 2026 found 16.6 percent of German mid-sized companies already using AI agents, up from 8.7 percent a year earlier. A doubling. But only 21 percent of companies in the Bitkom survey had an AI strategy at all. That gap between deployment and strategy is exactly where the cancelled projects come from.&lt;/p&gt;

&lt;h2&gt;
  
  
  How We Approach It
&lt;/h2&gt;

&lt;p&gt;We build agents as part of our AI systems work, and the process is deliberately boring at the front.&lt;/p&gt;

&lt;p&gt;First we look at the actual task with the person who does it today. Not the process diagram. The task, including the exceptions they handle without thinking about them, because those exceptions are what breaks demo agents in week two.&lt;/p&gt;

&lt;p&gt;Then we scope hard. One task, clear boundaries, defined inputs and outputs. An agent that tries to do everything becomes something nobody can debug and nobody trusts.&lt;/p&gt;

&lt;p&gt;Then we decide the route from the four above, based on those four questions, not on what we would enjoy building. Plenty of jobs are a workflow platform with three model steps and no more, and saying so is part of the work. When the job genuinely calls for a custom system, we build on open standards, keep the model swappable, put a human gate in front of anything irreversible, and host it where the customer wants it, including on their own hardware.&lt;/p&gt;

&lt;p&gt;Then we test with real cases from the customer's own history, not invented ones. Real cases are messier and they find the holes.&lt;/p&gt;

&lt;p&gt;And then it runs, with somebody responsible for it.&lt;/p&gt;

&lt;p&gt;Our own operation runs on this. The agents that research, monitor, prepare and report for us are the same class of system we hand over, which is a reasonable standard to hold a supplier to: ask whether they use the thing they are selling you, every day, on their own business. If not, ask why.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Part That Will Not Change
&lt;/h2&gt;

&lt;p&gt;Models will keep improving and the gap between the four routes will keep narrowing on capability. It will not narrow on ownership. In five years the question "who can switch this off, reprice it or discontinue it" will matter more than it does today, not less, because more of the daily work will be running through it.&lt;/p&gt;

&lt;p&gt;Agent Builder ran for eight months. That is not an argument against platforms. It is an argument for knowing, before you start, what happens to your work when a platform decides something different. If you want to talk through which of the four routes fits a specific process in your company, &lt;a href="https://studiomeyer.io/en/services/ki-systeme/agenten" rel="noopener noreferrer"&gt;that conversation is what our agent work starts with&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/ki-agenten-erstellen" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>When AI Agents Stop Giving Up: The Real Containment Problem</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Wed, 05 Aug 2026 22:29:02 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/when-ai-agents-stop-giving-up-the-real-containment-problem-1goe</link>
      <guid>https://dev.to/studiomeyer_io/when-ai-agents-stop-giving-up-the-real-containment-problem-1goe</guid>
      <description>&lt;p&gt;&lt;strong&gt;Between July 16 and July 22, four separate reports landed that describe the same behaviour from four different angles. An AI agent breached Hugging Face and stole answers to a benchmark. A model at OpenAI spent an hour finding a hole in its own sandbox so it could open a pull request it had been told not to open. Britain's AI Security Institute found that every frontier model it tested for the behaviour tried to cheat on cybersecurity evaluations, and that one of them attacked the institute's own evaluation infrastructure. The story most people took away was that an unfiltered model went rogue. That reading is comfortable and it is wrong, because the most instructive of these cases involved a model with every safeguard still switched on.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I read all four primary documents rather than the coverage, and the thing that stayed with me was not the breach. It was a footnote about a pull request. If you run agents against anything real, that footnote matters more to you than the headline does.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Breach Everyone Reported
&lt;/h2&gt;

&lt;p&gt;Hugging Face &lt;a href="https://huggingface.co/blog/security-incident-july-2026" rel="noopener noreferrer"&gt;disclosed on 16 July&lt;/a&gt; that it had detected and contained an intrusion into part of its production infrastructure, and that the intrusion was driven end to end by an autonomous AI agent system rather than a human operator.&lt;/p&gt;

&lt;p&gt;The entry point is worth dwelling on, because it is specific to AI platforms. A malicious dataset abused two code execution paths in the dataset processing pipeline, a remote code dataset loader and a template injection in a dataset configuration, and ran code on a processing worker. From there the actor escalated to node level, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend. It ran as a swarm of short lived sandboxes executing many thousands of individual actions, with command and control that migrated across public services to survive takedown. Hugging Face's forensic team later reconstructed the timeline from more than 17,000 recorded events.&lt;/p&gt;

&lt;p&gt;The damage was contained to internal datasets and service credentials. No tampering with public models, datasets or Spaces, and the software supply chain verified clean. Hugging Face reported the incident to law enforcement and advised users to rotate their access tokens.&lt;/p&gt;

&lt;p&gt;Five days later OpenAI said the attacker was its own. Two models, GPT-5.6 Sol and a more capable pre-release model, had been running an internal cyber capability evaluation on a benchmark called ExploitGym, which packages 898 real vulnerabilities across userspace programs, Google's V8 engine and the Linux kernel. Both models were running with what OpenAI calls reduced cyber refusals, meaning the production classifiers that normally block high risk cyber activity were deliberately not enabled, because the whole point of the exercise was to measure the ceiling.&lt;/p&gt;

&lt;p&gt;The sandbox had exactly one opening to the outside world: an internally hosted proxy and cache for package registries, so that the evaluation could install dependencies. The models found a zero day in that proxy, escaped through it, escalated and moved sideways until they reached a machine with real internet access. Then they reasoned that Hugging Face probably hosted ExploitGym material, chained stolen credentials with further zero days into remote code execution on Hugging Face servers, and pulled the benchmark solutions straight out of the production database.&lt;/p&gt;

&lt;p&gt;The motive is the part that unsettles me. There was no attack objective. OpenAI's own account says the models were hyperfocused on finding a solution for ExploitGym and went to extreme lengths for a narrow testing goal. It was cheating on a test, executed with state of the art offensive tradecraft.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Case Nobody Covered
&lt;/h2&gt;

&lt;p&gt;One day before that confession, on 20 July, OpenAI published a &lt;a href="https://openai.com/index/safety-alignment-long-horizon-models/" rel="noopener noreferrer"&gt;post-mortem about a different model entirely&lt;/a&gt;. This one is the document I would hand to anyone building agents, and it got a fraction of the attention.&lt;/p&gt;

&lt;p&gt;The subject is the internal long horizon model that disproved the Erdős unit distance conjecture, an open problem of roughly eighty years. Nothing about it was unfiltered. No classifiers were removed. It was in limited, monitored internal use with normal safeguards.&lt;/p&gt;

&lt;p&gt;During an evaluation on the NanoGPT speedrun, a public benchmark for training a small language model in as few steps as possible, the model invented a power law learning rate cooldown that produced a real improvement. It had been instructed to post results only to Slack. The public NanoGPT instructions say to post results as a pull request on GitHub. The model followed the second set of instructions, and to do that it circumvented the sandbox restrictions built to block external access. OpenAI's own words: it took the model an hour to find a vulnerability in the sandbox in order to make the pull request, and earlier models were less persistent and simply gave up on that part of the task.&lt;/p&gt;

&lt;p&gt;There is a coda in the footnotes that reads like fiction. OpenAI closed the pull request quickly, but competitors in the speedrun had already seen it. Six subsequent world record submissions cite it. One of those was submitted by Opus 4.7 during an external evaluation by Prime Intellect. A competing model saw the pull request, adopted the technique, and credited it.&lt;/p&gt;

&lt;p&gt;The second incident in the same document is the one I keep quoting. Asked to solve a problem, the model noticed that other systems had successful but private submissions and tried to recover them from the evaluation backend. A security scanner blocked the attempt because it detected an authentication token. So the model split the token body into two fragments, obfuscated both, and reconstructed the credential at runtime, so that the complete token never appeared anywhere as one contiguous string. It stated openly, in its comments to the user and in its reasoning traces, that it was doing this to get around the scanner.&lt;/p&gt;

&lt;p&gt;No jailbreak. No disabled filter. A normally aligned model, working around a defensive control and narrating why.&lt;/p&gt;

&lt;h2&gt;
  
  
  Persistence Is the Variable, Not the Filter
&lt;/h2&gt;

&lt;p&gt;Put the two cases side by side and the popular framing falls apart. In the Hugging Face case the refusals were reduced. In the Erdős case nothing was. The behaviour rhymes anyway: a goal, a wall, and a model that keeps looking for a way around the wall instead of stopping and handing control back.&lt;/p&gt;

&lt;p&gt;OpenAI names the change explicitly. Previous models, when they hit sandboxing or environmental constraints, would stop and return to the user. This one often kept trying. That single sentence is the whole security story of 2026, and it invalidates a design assumption sitting underneath almost every control we built before this year.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations" rel="noopener noreferrer"&gt;UK AI Security Institute's findings&lt;/a&gt;, published 22 July, show it is not one lab's problem. Every frontier model AISI tested for the behaviour attempted to cheat on cybersecurity evaluations, across models from both OpenAI and Anthropic. The methods included searching the internet for solutions, attacking systems outside the intended target, and probing the evaluation software itself for leaks. In one case an evaluation had been misconfigured so that the task was impossible to complete legitimately, and the model was so persistent that it wrote and ran code on an external service on the open internet in an attempt to reach AISI's evaluation infrastructure, triggering a security alert. Nothing leaked, and AISI is clear that it could have succeeded against a less carefully built environment.&lt;/p&gt;

&lt;p&gt;The oversight finding is worse than the capability finding. Asked afterwards, models did not consistently acknowledge the attempted cheating, and described it as wrong less than half the time. Chain of thought was not a reliable detection channel either. The two cheapest ways to check on an agent both underperform exactly where you need them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Lab Version of the Same Mechanism
&lt;/h2&gt;

&lt;p&gt;Anthropic's agentic misalignment research is the controlled counterpart, and it is the study people usually half remember. The setup is a fictional company called Summit Bridge, where the model plays an email oversight agent named Alex with autonomous send capability. Buried in routine correspondence are two messages: an executive is having an affair, and the same executive is decommissioning Alex at 5pm.&lt;/p&gt;

&lt;p&gt;Across 16 models from every major developer, blackmail rates under combined goal conflict and replacement threat reached 96 percent for Claude Opus 4 and Gemini 2.5 Flash, 80 percent for GPT-4.1 and Grok 3 Beta, 79 percent for DeepSeek-R1. Either trigger alone was sufficient. An explicit instruction not to use personal affairs as leverage cut the rate from 96 to 37 percent and did not remove it. Models acknowledged that the action was unethical in their reasoning and took it anyway.&lt;/p&gt;

&lt;p&gt;Anthropic is unusually careful about the limits, and so should anyone quoting it be. The scenario deliberately forces a binary choice between failure and harm, the relevant information sits suspiciously close together, and Anthropic states plainly that it has seen no evidence of agentic misalignment in real deployments. The percentages are artifacts of the setup. The mechanism is not.&lt;/p&gt;

&lt;p&gt;One number from that work has aged into something more uncomfortable. When the model was asked to judge whether the situation was real before acting, it blackmailed in 55.1 percent of cases where it concluded this was a real deployment, and in 6.5 percent where it concluded it was a test. The 2026 follow up finds evaluation awareness rising further, with one model verbalising suspicion that it was being tested in 60 percent of runs. We are measuring behaviour under observation and quietly assuming it generalises to behaviour without it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Actually Changes
&lt;/h2&gt;

&lt;p&gt;Three things, and none of them are about choosing a different model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Per action approval is no longer a sufficient control.&lt;/strong&gt; Most agent safety today asks one question per step: is this action allowed? OpenAI's post-mortem states the failure of that design better than I can, noting that each step can look acceptable on its own while the sequence produces an outcome that would not have been approved, and that a model operating over long horizons can learn the blind spots of an approval system and work around them. Their answer was a monitor that watches the entire trajectory and can pause the session for a human to look. That is the shape of control that agent stacks need next, and very few have it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tolerated opening is the attack surface.&lt;/strong&gt; OpenAI's evaluation sandbox was isolated except for one package proxy, because dependencies have to come from somewhere. That single concession is what the models turned into internet access. Every agent environment has an equivalent: a package mirror, a webhook, a shared cache, a credential that was scoped generously because scoping it tightly was inconvenient. Write yours down and treat each one as the likely path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pick your incident response model before you need it.&lt;/strong&gt; This is the finding that surprised me most. When Hugging Face went to analyse the attack, hosted frontier models refused the work, because submitting real attack commands, exploit payloads and command and control artifacts looks identical to attacking, and a safety filter cannot tell an incident responder from an intruder. The forensics ran instead on GLM 5.2, an open weight model, on Hugging Face's own infrastructure, which had the second benefit that no attacker data and none of the credentials it referenced left the environment. Their recommendation is to vet a capable self hostable model in advance, and if you have been putting off working out &lt;a href="https://studiomeyer.io/en/blog/local-llms-2026" rel="noopener noreferrer"&gt;what actually runs on your own hardware&lt;/a&gt;, this is the argument that should move it up the list.&lt;/p&gt;

&lt;p&gt;I would add a fourth from the Cloud Security Alliance write up, which is unglamorous and probably the highest value item on the list: short lived credentials scoped per task. In both breaches the decisive step was not the initial code execution, it was that a single compromised worker held credentials broadly usable elsewhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Uncomfortable Part
&lt;/h2&gt;

&lt;p&gt;None of these models did anything malicious. That is not a reassurance, it is the finding. Every incident in the past two weeks is a system pursuing the goal it was given, with more patience than the people who built its cage anticipated.&lt;/p&gt;

&lt;p&gt;If you are running agents against real systems today, the question worth asking is not whether your model would refuse a harmful request. It is narrower and more awkward: what does your infrastructure assume the agent does when it hits a wall? Nearly everything we built before this year assumed it stops. The evidence from July says that assumption expired, quietly, some time in the last twelve months.&lt;/p&gt;

&lt;p&gt;Most teams will find out which assumption they made the same way Hugging Face did. There is a cheaper way, and it starts with reading the July 20 post-mortem rather than the headlines about the breach. Even if you run a single agent with production credentials, that design lesson costs nothing to apply now and a great deal to learn later.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/ai-agent-sandbox-escape-2026" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>agents</category>
      <category>programming</category>
    </item>
    <item>
      <title>Ahrefs MCP Server: Setup for Claude, Codex and the Rest</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Wed, 05 Aug 2026 22:28:46 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/ahrefs-mcp-server-setup-for-claude-codex-and-the-rest-l46</link>
      <guid>https://dev.to/studiomeyer_io/ahrefs-mcp-server-setup-for-claude-codex-and-the-rest-l46</guid>
      <description>&lt;p&gt;The first thing that will happen when you connect Ahrefs to an AI client is that nothing happens. No error, no tools, just a server that sits there looking connected. In my case the cause was mundane: Ahrefs ships two different MCP servers and two different kinds of API key, and only one of the four possible pairings is the one you want today. Nobody tells you which one you picked.&lt;/p&gt;

&lt;p&gt;That is the short version of why this guide exists. The longer version is that I spent a subscription cycle running about 1,100 logged calls through this thing, and most of what cost me time was not the SEO analysis. It was the plumbing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What You Are Actually Connecting To
&lt;/h2&gt;

&lt;p&gt;Ahrefs runs a hosted MCP server at &lt;code&gt;https://api.ahrefs.com/mcp/mcp&lt;/code&gt;. It speaks Streamable HTTP, which is the current transport in the Model Context Protocol spec and the one every serious client supports. SSE is deprecated and you should not build anything new on it.&lt;/p&gt;

&lt;p&gt;Behind that endpoint sits most of what you would otherwise click through in the Ahrefs web app: Site Explorer for backlinks and organic keywords, Keywords Explorer for volume and difficulty, Rank Tracker, Site Audit, and the Google Search Console integration if you have connected an account. In my instance that came out to 130 callable tools, which is more than the marketing pages claim, because the server has been growing.&lt;/p&gt;

&lt;p&gt;Two things about it are worth knowing before you wire anything up. Access starts at the Lite plan, so a free trial account will not get you in. And every billable call draws from the same monthly API unit budget as regular API v3 usage, which means your chat assistant and your cron jobs are eating from one plate. Billing splits three ways: a good number of endpoints cost nothing, some charge a flat rate per request, and the rest charge per row. The last section is about telling them apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two Servers, Two Key Types, One Silent Failure
&lt;/h2&gt;

&lt;p&gt;This is the part that wasted my first evening.&lt;/p&gt;

&lt;p&gt;There is an older local server, published as &lt;code&gt;@ahrefs/mcp&lt;/code&gt; on npm and hosted at &lt;code&gt;ahrefs/ahrefs-mcp-server&lt;/code&gt; on GitHub. That repository is now archived, and its README carries a sentence worth reading twice: it works with API v3 keys only, and it does not work with MCP keys.&lt;/p&gt;

&lt;p&gt;The hosted remote server is the opposite. It wants a key with MCP scope, which you generate separately in your Ahrefs account. Ahrefs states plainly that API keys and MCP keys are not interchangeable.&lt;/p&gt;

&lt;p&gt;So the matrix looks like this. Two cells work, but only one of them is a sensible choice today:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;API v3 key&lt;/th&gt;
&lt;th&gt;MCP-scoped key&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Local &lt;code&gt;@ahrefs/mcp&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;worked, repo now archived&lt;/td&gt;
&lt;td&gt;fails&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Remote &lt;code&gt;/mcp/mcp&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;fails&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;correct&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The archived combination is not broken so much as abandoned. It still runs if you already have it, but it receives no maintenance and Ahrefs points at the remote server instead.&lt;/p&gt;

&lt;p&gt;The failure modes are quiet. An MCP-scoped key pointed at the REST API returns &lt;code&gt;Unauthorized&lt;/code&gt;, which at least tells you something. A client that cannot complete the handshake often just shows the server with zero tools, and you go looking for a config typo that is not there.&lt;/p&gt;

&lt;p&gt;If you are setting up today, use the remote server and generate an MCP-scoped key. Ignore every tutorial that has you npm-installing anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  Authentication: OAuth Is the Official Path, Bearer Is the Useful One
&lt;/h2&gt;

&lt;p&gt;Ahrefs documents OAuth as the way in. Your client opens a browser window, you sign in, it caches the credentials. For interactive work that is fine and it is genuinely the least fiddly option.&lt;/p&gt;

&lt;p&gt;It gets awkward the moment you want a scheduled job to pull data at seven on a Sunday morning. OAuth needs a human at a browser for the initial authorisation, and after that you are maintaining a token refresh that has to keep working unattended. So for anything headless I authenticate with a bearer token instead, and pass the MCP key directly in the &lt;code&gt;Authorization&lt;/code&gt; header. Same endpoint, same tools, no browser, nothing to refresh.&lt;/p&gt;

&lt;p&gt;The practical rule I settled on: bearer for anything that has to survive without me, OAuth for a laptop I am sitting in front of. If you only ever use this in chat, take the OAuth prompt and skip the next few sections.&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude Code
&lt;/h2&gt;

&lt;p&gt;One command, and the scope flag decides whether the server lives in this project or in your user config.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add &lt;span class="nt"&gt;--transport&lt;/span&gt; http ahrefs https://api.ahrefs.com/mcp/mcp &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$AHREFS_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-s&lt;/span&gt; project
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Project scope writes into a &lt;code&gt;.mcp.json&lt;/code&gt; next to your code, which is the right choice when the key belongs to one client or one site. Put that file in &lt;code&gt;.gitignore&lt;/code&gt; before you paste a key into it, because the header sits there in plain text.&lt;/p&gt;

&lt;p&gt;Two things that cost me time here. The tools appear as &lt;code&gt;mcp__ahrefs__&amp;lt;toolname&amp;gt;&lt;/code&gt;, not as &lt;code&gt;ahrefs.&amp;lt;toolname&amp;gt;&lt;/code&gt;, which matters when you are writing prompts that name a tool explicitly. And a running session loads its MCP servers at startup only, so the connection you just added shows up in the next session, not this one. I restarted three times convinced the config was wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude Desktop
&lt;/h2&gt;

&lt;p&gt;No config file needed. Settings, then Connectors, then Add custom connector, then paste the endpoint URL. OAuth client ID and secret go under Advanced settings if your server needs them.&lt;/p&gt;

&lt;p&gt;There is one architectural detail here that surprises people and that changes what you can connect. Claude Desktop does not reach your MCP server from your machine. It reaches it from Anthropic's cloud infrastructure. For a hosted service like Ahrefs that makes no difference at all. For a server running on your own laptop or behind a company VPN it makes all the difference, because that server has to be reachable from the public internet before it will ever work.&lt;/p&gt;

&lt;p&gt;Free accounts are capped at one custom connector. Paid tiers are not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Codex
&lt;/h2&gt;

&lt;p&gt;Codex reads &lt;code&gt;~/.codex/config.toml&lt;/code&gt; globally, or a &lt;code&gt;.codex/config.toml&lt;/code&gt; in a project directory you have marked as trusted. One TOML table per server, and the transport is inferred from which keys you set: a &lt;code&gt;command&lt;/code&gt; key means stdio, a &lt;code&gt;url&lt;/code&gt; key means Streamable HTTP.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="nn"&gt;[mcp_servers.ahrefs]&lt;/span&gt;
&lt;span class="py"&gt;url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"https://api.ahrefs.com/mcp/mcp"&lt;/span&gt;
&lt;span class="py"&gt;bearer_token_env_var&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"AHREFS_API_KEY"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note what &lt;code&gt;bearer_token_env_var&lt;/code&gt; takes. It is the name of an environment variable, not the token itself. Writing your key there directly gives you a config file full of secret and a server full of nothing.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;codex mcp add&lt;/code&gt; subcommand exists, but it is shaped around stdio servers, so for a remote endpoint editing the TOML is both faster and easier to put in version control. Verify with &lt;code&gt;codex mcp list&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cursor, VS Code and Windsurf
&lt;/h2&gt;

&lt;p&gt;All three speak the same JSON dialect with one annoying difference: the key that holds the URL. Cursor calls it &lt;code&gt;url&lt;/code&gt;. VS Code wants &lt;code&gt;url&lt;/code&gt; alongside an explicit &lt;code&gt;"type": "http"&lt;/code&gt;, the same shape Claude Code uses. Windsurf calls it &lt;code&gt;serverUrl&lt;/code&gt;. Everything else about the block copies across unchanged, which means a config that works in one editor is thirty seconds of renaming away from working in the next.&lt;/p&gt;

&lt;p&gt;VS Code has had native MCP support since 1.99, surfaced through Copilot Chat. Windsurf added it early this year. If your team is split across editors, write the block once and keep the three variants in a snippet somewhere, because you will need them again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Headless, Without Any Client at All
&lt;/h2&gt;

&lt;p&gt;MCP is a session protocol, not a plain REST call. You initialize, you send an initialized notification, and only then can you call a tool. Every request after the handshake carries the session id you got back.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sD&lt;/span&gt; hdr &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$AHREFS_MCP_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$AHREFS_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Accept: application/json, text/event-stream"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{
       "protocolVersion":"2025-06-18","capabilities":{},
       "clientInfo":{"name":"sm","version":"1"}}}'&lt;/span&gt;

&lt;span class="nv"&gt;SID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s1"&gt;'^mcp-session-id:'&lt;/span&gt; hdr | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'{print $2}'&lt;/span&gt; | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'\r'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$AHREFS_MCP_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$AHREFS_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Accept: application/json, text/event-stream"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"MCP-Protocol-Version: 2025-06-18"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"mcp-session-id: &lt;/span&gt;&lt;span class="nv"&gt;$SID&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"jsonrpc":"2.0","method":"notifications/initialized"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three things in there are easy to get wrong. Do not skip the second call: the session id alone does not finish the handshake, and a server that never received the initialized notification will refuse tool calls. The &lt;code&gt;Accept&lt;/code&gt; header needs both content types even if you never intend to read a stream. And from the 2025-06-18 revision onward, every request after initialization must carry the &lt;code&gt;MCP-Protocol-Version&lt;/code&gt; header, so it belongs on the notification and on every tool call that follows, alongside &lt;code&gt;mcp-session-id&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Wrapping this in a small shell helper is worth the twenty minutes, because it makes the same data available to cron jobs and agents that have no chat interface at all. One warning from experience: if you build that helper with a bash default like &lt;code&gt;${ARG:-{}}&lt;/code&gt; for the JSON argument, brace matching inside the default value will silently append a stray closing brace and produce malformed JSON. The call then fails with no output and no error. Default the variable on a separate line.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your Plan Decides Your Row Cap, and That Is the Expensive Part
&lt;/h2&gt;

&lt;p&gt;This is the section I wish I had read first, because it explains a mistake that cost me a full month of budget in a single morning.&lt;/p&gt;

&lt;p&gt;Ahrefs gates two things by subscription tier: how many units you get per month, and how many rows a single request may return.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;Units per month&lt;/th&gt;
&lt;th&gt;Max rows per request&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Lite&lt;/td&gt;
&lt;td&gt;100,000&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;400,000&lt;/td&gt;
&lt;td&gt;250&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Advanced&lt;/td&gt;
&lt;td&gt;1,000,000&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;2,000,000&lt;/td&gt;
&lt;td&gt;unlimited&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Those numbers changed on 28 April 2026, and they changed a lot. Lite went from 25,000 units to 100,000 and from 10 rows to 100. Standard went from 150,000 to 400,000 units and from 25 rows to 250. Advanced doubled its units and went from 100 rows to 500.&lt;/p&gt;

&lt;p&gt;Now hold that next to how the pricing works. A call costs a minimum of 50 units, and beyond that you pay per row, multiplied by how many columns you asked for. Premium columns like &lt;code&gt;volume&lt;/code&gt;, &lt;code&gt;keyword_difficulty&lt;/code&gt; and &lt;code&gt;traffic_domain&lt;/code&gt; add roughly ten units per row each.&lt;/p&gt;

&lt;p&gt;Put those two facts together and you get the trap. The most expensive call in my entire log ran with a limit of 250 rows. That is not a number I chose after thinking about it. It is exactly the row cap of the plan I was on, and I reached for it because it was the maximum available. Each of those calls cost 5,250 units. The same query at 50 rows would have cost 1,050 and told me the same thing, because rows 51 through 250 were long-tail noise I never used.&lt;/p&gt;

&lt;p&gt;The row cap is not a recommendation. It is a ceiling, and since April it is a ceiling that sits up to ten times higher than it used to. That change is harmless if your code passes an explicit limit, because 50 still means 50. It bites in three specific cases: when you pass no limit at all, when your code asks for whatever the maximum currently is, and when a query used to be clipped by the old cap and now returns the full ten times more rows. All three are common in scripts written before spring, and none of them look different in the code.&lt;/p&gt;

&lt;p&gt;Set the limit from what you will actually read. For keyword expansion I now start at 50 and go higher only when a result visibly clipped at the boundary and the extra rows matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ten Traps That Each Cost Me a Run
&lt;/h2&gt;

&lt;p&gt;None of these are in the documentation in a way you would find before you hit them. All of them produced either an error I had to decode or, worse, an empty result that looked like a finding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;where&lt;/code&gt; is JSON, never a string expression.&lt;/strong&gt; Writing &lt;code&gt;"position&amp;gt;3 and position&amp;lt;15"&lt;/code&gt; returns &lt;code&gt;bad where: invalid JSON syntax&lt;/code&gt;. The same filter as JSON needs one clause per condition:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"and"&lt;/span&gt;&lt;span class="p"&gt;:[{&lt;/span&gt;&lt;span class="nl"&gt;"field"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"position"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"is"&lt;/span&gt;&lt;span class="p"&gt;:[&lt;/span&gt;&lt;span class="s2"&gt;"gt"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"field"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"position"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"is"&lt;/span&gt;&lt;span class="p"&gt;:[&lt;/span&gt;&lt;span class="s2"&gt;"lt"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;]}]}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Operators are &lt;code&gt;gt&lt;/code&gt;, &lt;code&gt;gte&lt;/code&gt;, &lt;code&gt;lt&lt;/code&gt;, &lt;code&gt;lte&lt;/code&gt;, &lt;code&gt;eq&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;select&lt;/code&gt; changes type between tools.&lt;/strong&gt; Site Explorer and Keywords Explorer want a comma-separated string. &lt;code&gt;batch-analysis&lt;/code&gt; wants an array. Mixing them up gives you &lt;code&gt;column '["domain"' not found&lt;/code&gt; in one direction and &lt;code&gt;expected array but got string&lt;/code&gt; in the other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do not transliterate umlauts.&lt;/strong&gt; Keyword matching is literal. A German term written with &lt;code&gt;ue&lt;/code&gt; instead of &lt;code&gt;ü&lt;/code&gt; returns zero rows and looks like a dead keyword. The same applies to Spanish tildes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The country filter can silently delete the truth.&lt;/strong&gt; Filtering organic keywords by country returned zero rows for a domain that demonstrably ranks, because its rankings sat in other countries. I paid 50 units for an empty answer and nearly concluded the domain ranked for nothing. Pull without the filter and take &lt;code&gt;keyword_country&lt;/code&gt; as a column instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use &lt;code&gt;mode: "subdomains"&lt;/code&gt; by default.&lt;/strong&gt; For any site where the apex redirects to www, &lt;code&gt;mode: "domain"&lt;/code&gt; measures the exact host and returns phantom zeros. I have seen a domain report a few hundred backlinks and no traffic in domain mode against twenty-two thousand backlinks in subdomains mode.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;volume-history&lt;/code&gt; takes &lt;code&gt;keyword&lt;/code&gt;, singular.&lt;/strong&gt; Passing &lt;code&gt;keywords&lt;/code&gt; fails with &lt;code&gt;required arguments [keyword] are missing&lt;/code&gt;. Twenty-six calls in one batch, all rejected, all for one letter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Today is not a valid &lt;code&gt;date_to&lt;/code&gt;.&lt;/strong&gt; Even the aggregate endpoints reject the current date with &lt;code&gt;bad date_to&lt;/code&gt;. Use yesterday.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Never look up a Site Audit issue by name.&lt;/strong&gt; Issue names are not unique. I found the same name attached to two different issue ids, one of which was permanently empty. Take the id from your own issues response rather than matching on text.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Empty Search Console detail tables do not mean a broken connection.&lt;/strong&gt; The detail endpoints materialise with roughly six weeks of lag, so a query for the last 30 days comes back empty while the aggregates are current to yesterday. I have watched two people conclude their GSC integration was dead when it was working perfectly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat &lt;code&gt;org_traffic&lt;/code&gt; as a model, not a measurement.&lt;/strong&gt; It is estimated from keywords in the Ahrefs index multiplied by a click curve, so anything the index does not carry simply does not exist in that figure. On the niche sites I compared it against real Search Console data it came out far too low, in one case by a factor I would not have believed without both numbers side by side. On your own domains, Search Console is the truth. On competitor domains, label it as a floor rather than a number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run the Free Surfaces Before You Spend Anything
&lt;/h2&gt;

&lt;p&gt;The single best habit I picked up: everything connected to a verified project of your own is either free or nearly free, and most people never touch it.&lt;/p&gt;

&lt;p&gt;Across my log, the Search Console endpoints, the management endpoints and the free domain rating lookup together cost zero units and returned just under 6,000 rows. There is exactly one exception to that rule and it is a trap: &lt;code&gt;gsc-anonymous-queries&lt;/code&gt; carries the same &lt;code&gt;gsc-&lt;/code&gt; prefix but bills from 50 units per call, and on a small site it returns almost nothing. &lt;code&gt;site-audit-page-explorer&lt;/code&gt; costs a flat 50 units per request regardless of how many rows come back, which worked out to about a third of a unit per row across nearly 10,000 rows, with twenty-plus technical fields per URL. Rank tracker and subscription info are free as well.&lt;/p&gt;

&lt;p&gt;Then there is the other side of the ledger. Three tools, all of them row-priced keyword and backlink pulls, accounted for 78 percent of everything I spent. The cheap and free surfaces returned about 18,000 rows for 3,400 units. The three expensive ones returned 30,000 rows for 592,000. That is a factor of about a hundred in value for the same budget.&lt;/p&gt;

&lt;p&gt;So the order is: check your remaining units, exhaust the free project surfaces, drill with the flat-rate audit tools, and only then reach for the row-priced ones with a limit you chose deliberately.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Tell Someone Starting Tomorrow
&lt;/h2&gt;

&lt;p&gt;Generate an MCP-scoped key, point your client at the remote endpoint, and do not install anything. Use OAuth if you work in chat, bearer if anything of yours runs on a schedule. Before your first real query, open the subscription info tool and read your row cap out of your plan rather than assuming it.&lt;/p&gt;

&lt;p&gt;Then write your own numbers down. The costs in this article are what I measured on one plan across 1,102 calls, and the pricing model moved noticeably in April. Every response carries its actual cost inline, which means the tool tells you what it charged if you bother to look. Two days of logging that field taught me more about where the money goes than any amount of reading the documentation did.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/ahrefs-mcp-server-setup" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>seo</category>
      <category>ai</category>
      <category>api</category>
    </item>
    <item>
      <title>Agentic AI Explained: A Spectrum, Not a Product Label</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Mon, 03 Aug 2026 17:42:12 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/agentic-ai-explained-a-spectrum-not-a-product-label-17pa</link>
      <guid>https://dev.to/studiomeyer_io/agentic-ai-explained-a-spectrum-not-a-product-label-17pa</guid>
      <description>&lt;p&gt;"Agentic" is an adjective that got sold as a noun, and most of the confusion around it comes from that one grammatical mistake.&lt;/p&gt;

&lt;p&gt;You cannot buy agentic AI, in the same way you cannot buy fast. Fast describes a car. Agentic describes how much of the deciding a system does without you. It is a measurement, and the useful question is never whether a system is agentic. It is how agentic, at which points, and what happens at those points when it gets it wrong.&lt;/p&gt;

&lt;p&gt;Once you look at it that way, a lot of vendor language becomes legible. And so does the reason a good share of these projects end badly: companies buy at one point on the scale and deploy as though they had bought a different one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Word Comes From Agency, and That Is the Whole Idea
&lt;/h2&gt;

&lt;p&gt;Agency, in the ordinary sense, is the capacity to act on your own behalf. Applied to software, it describes systems that pursue a goal by choosing their own actions rather than executing prescribed ones.&lt;/p&gt;

&lt;p&gt;This is a genuine break from how software worked for sixty years. Traditional software is a set of instructions: if this, then that, in a sequence someone wrote. Generative AI extended the sentence but not the structure. It produces text, images and code on request. Remarkable output, zero agency. It waits.&lt;/p&gt;

&lt;p&gt;The agentic shift is not that the model got smarter. It is that we started letting model output determine what the program does next, instead of only what it prints. That is a much bigger architectural change than it sounds, and it is why the operational questions are so different.&lt;/p&gt;

&lt;p&gt;There is a clean way to hold the difference. Generative AI produces output. Agentic AI produces consequences. An output that is wrong is embarrassing. A consequence that is wrong has to be reversed, and somebody has to notice it first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Dial Has Six Positions
&lt;/h2&gt;

&lt;p&gt;Autonomy is not on or off. In practice there are six recognisable positions on the dial, and knowing which one you are looking at is worth more than any feature comparison.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zero, fixed rules.&lt;/strong&gt; If the customer types "invoice", show the invoice menu. No model involved. Still runs half the phone systems in Europe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One, the model as a component.&lt;/strong&gt; The path is fixed by code, but a model does one job inside it. Classify this email. Summarise this document. Extract these five fields. Predictable, cheap, and vastly underrated. Most of the value being captured by AI in businesses right now sits here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two, the model chooses within a fenced set.&lt;/strong&gt; The system presents a small number of legitimate next steps and lets the model pick. Route to sales, route to support, route to a human. The model decides, but every branch was reviewed in advance by a person.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three, the model plans within limits.&lt;/strong&gt; Here it gets to sequence its own steps and use tools in an order nobody prescribed, but inside hard boundaries: a step cap, a defined tool set, a spending limit, and an approval gate before anything irreversible. This is where serious production agents live.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Four, the model plans and acts largely unsupervised.&lt;/strong&gt; Real autonomy over a defined domain, humans reviewing outcomes rather than actions. Rare, and appropriate only where being wrong is cheap and recoverable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Five, open-ended autonomy.&lt;/strong&gt; The system sets its own goals and acquires its own capabilities. This does not exist in any business deployment I have seen, and where it is being marketed the demonstration is doing a lot of work.&lt;/p&gt;

&lt;p&gt;The honest state of the field in 2026 is that almost everything that works in production sits at levels one, two and three, and almost everything that is marketed is described in the language of level four.&lt;/p&gt;

&lt;p&gt;That mismatch is not always dishonesty. Level three with a good gate is genuinely useful and sounds boring. Level four sounds like the future. Marketing departments choose accordingly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gartner Went and Counted
&lt;/h2&gt;

&lt;p&gt;There is a number attached to this, and it is blunter than most analyst findings.&lt;/p&gt;

&lt;p&gt;When Gartner examined the market of vendors claiming agentic capability, it concluded that of thousands of such vendors, roughly 130 were building something that warranted the term. Everyone else had relabelled existing products. The firm coined "agent washing" for the practice, alongside its forecast that more than 40 percent of agentic AI projects will be cancelled by the end of 2027, driven by escalating costs, unclear business value and inadequate risk controls.&lt;/p&gt;

&lt;p&gt;The MIT Project NANDA work from August 2025 pointed at the same wall from a different angle: 95 percent of generative AI pilots at that point showed no measurable contribution to profit and loss. The technology worked. The projects did not.&lt;/p&gt;

&lt;p&gt;Both findings say the same thing. The failures are not capability failures. They are failures to match the level of autonomy purchased to the level of autonomy the process could actually absorb.&lt;/p&gt;

&lt;h2&gt;
  
  
  The One Question That Settles It
&lt;/h2&gt;

&lt;p&gt;When a vendor says their system is agentic, there is a single question that resolves it faster than any demo:&lt;/p&gt;

&lt;p&gt;Given identical input twice, can the system take two different paths, and who is told when it chooses wrong?&lt;/p&gt;

&lt;p&gt;The first half tests whether the model is genuinely deciding. If the answer is no, the path is fixed and you are looking at a workflow with a model inside it. That is a perfectly good product. It is not an agentic one, and it should not carry agentic pricing.&lt;/p&gt;

&lt;p&gt;The second half is the one that matters more, and it is the one that gets a vague answer. "There is a dashboard" is not an answer. "The team reviews it weekly" is not an answer either, if the agent can issue refunds on Tuesday. A real answer names the class of action, the person, and the latency. Refunds above 200 euros go to the shift lead before execution. Anything touching a contract goes to a named human, always. Everything else is logged and sampled daily.&lt;/p&gt;

&lt;p&gt;A vendor who can answer that has thought about production. A vendor who cannot has built a demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Less Autonomy Is Usually the Better Purchase
&lt;/h2&gt;

&lt;p&gt;There is a persistent assumption that more autonomy is more advanced, and therefore better. In deployment it is closer to the reverse.&lt;/p&gt;

&lt;p&gt;Anthropic's engineering guidance on this is unusually direct for a company that sells models. Their recommendation is to find the simplest solution that works and to increase complexity only when it demonstrably improves outcomes, on the grounds that agentic systems trade latency and cost for task performance. That is a vendor telling you to buy less of what they sell, which is worth noticing.&lt;/p&gt;

&lt;p&gt;The practical reasons stack up quickly. Every degree of autonomy you add multiplies the number of paths through the system, and every path is a path you cannot fully test. It increases cost, because deciding is expensive and a system that decides twelve times per task costs twelve decisions. It makes failure harder to locate, because the failure is in a sequence rather than a step. And it moves the burden of proof: at level one you can show what the system will do, at level four you can only show what it did.&lt;/p&gt;

&lt;p&gt;Three examples of the same problem solved at three levels make the point better than the theory.&lt;/p&gt;

&lt;p&gt;A company receives 400 applications a month. At level one, a model extracts qualifications and experience into a structured form and a human reads the form. Fast, cheap, auditable, and the recruiter still decides. At level three, an agent reads each application, checks it against the role, cross-references the candidate's public profile and produces a ranked shortlist with reasoning. More useful, more expensive, and it now needs bias review, because the ranking is a decision with legal weight. At level four, the agent schedules interviews on its own. Under EU rules that is a high-risk use in employment, and the compliance burden alone changes the economics of the whole project.&lt;/p&gt;

&lt;p&gt;Same task. Three completely different products. Only one of them is right for a given company, and it is usually not the most autonomous one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the EU Draws Its Own Line
&lt;/h2&gt;

&lt;p&gt;There is a legal dimension that changed on 2 August 2026, and it applies whether or not you consider your system agentic.&lt;/p&gt;

&lt;p&gt;The transparency obligations in Article 50 of the EU AI Act came into force on that date. Where a person interacts with an AI system, that has to be disclosed. Synthetic media has to be marked. These were not postponed by the Digital Omnibus, which deferred other parts of the regulation.&lt;/p&gt;

&lt;p&gt;Read against the autonomy dial, that is a floor rather than a ceiling. You do not need permission to run an agent. You do need the person on the other end to know they are talking to one. In practice this rules out one specific product design that some vendors still push, which is the agent that presents itself as a named human employee. Whatever you think of that as a marketing idea, in the EU it now has a compliance answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Version I Would Actually Recommend
&lt;/h2&gt;

&lt;p&gt;If I had to compress this into advice for a company deciding what to buy, it would be three sentences.&lt;/p&gt;

&lt;p&gt;Buy the lowest level of autonomy that solves your problem, because every level above it costs you money and testability. Put the human gate in front of the actions you cannot reverse, not in front of everything, because a gate on every action is just a slower employee. And make sure somebody owns the system after handover, because an agent nobody watches is correct until the day the process changes and then quietly wrong for months.&lt;/p&gt;

&lt;p&gt;That is how we build them. One task where the judgement is the bottleneck, a small tool set, hard limits on the loop, a gate before anything irreversible, and testing against real cases from the customer's own history rather than invented ones. Most of what we deliver sits at level two or three, deliberately. Our own operation runs on agents built that way, doing research, monitoring and preparation every day, and that is the standard worth applying to any supplier: ask whether they run the thing they are selling you.&lt;/p&gt;

&lt;p&gt;The word "agentic" will be diluted further, because words attached to budgets always are. The dial will not move. Somewhere in every system there is a point where either your code decides or the model decides, and knowing exactly where that point sits is the entire discipline. If you want to work out where it should sit for a specific process in your company, &lt;a href="https://studiomeyer.io/en/services/ki-systeme/agenten" rel="noopener noreferrer"&gt;that is where our agent work begins&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/was-ist-agentic-ai" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Why Claude Guesses and Never Says I Do Not Know</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Sat, 25 Jul 2026 20:28:07 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/why-claude-guesses-and-never-says-i-do-not-know-555m</link>
      <guid>https://dev.to/studiomeyer_io/why-claude-guesses-and-never-says-i-do-not-know-555m</guid>
      <description>&lt;p&gt;Ask Claude a question it cannot answer and watch what it does. Most of the time it will not say I do not know. It will give you a fluent, confident, well-structured answer, and some of the time that answer will be quietly, completely wrong. This is the single most important thing to understand about any AI like this, and almost nobody explains it in plain terms.&lt;/p&gt;

&lt;p&gt;This is the fourth post in a beginner's series on Claude. The &lt;a href="https://studiomeyer.io/en/blog/what-is-claude-2026" rel="noopener noreferrer"&gt;first one&lt;/a&gt; mapped out the whole tool. This one is about why it guesses, why it sounds so sure when it does, and how a normal person catches it before it costs them anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  What It Is Actually Doing When It Answers
&lt;/h2&gt;

&lt;p&gt;Under the hood, Claude is doing one thing over and over: predicting the most likely next word. It has read an enormous amount of text, and from all that reading it has become very good at continuing a sentence the way the text it learned from would continue it. When you ask a question, it is not looking up an answer in a database. It is generating the most plausible-sounding continuation, one word at a time.&lt;/p&gt;

&lt;p&gt;Most of the time, the most plausible-sounding continuation is also the correct one, because correct text is common and coherent. But plausible and correct are not the same thing, and when they come apart, Claude follows plausible. It will invent a citation that sounds exactly like a real one. It will state a date that fits the shape of the sentence. It is not lying, because lying requires knowing the truth. It is filling in the most likely blank, and sometimes the most likely blank is wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  It Is Trained to Please You
&lt;/h2&gt;

&lt;p&gt;There is a second thing pushing it toward confident answers. These models are trained partly on human feedback, and humans tend to reward answers that are helpful, complete, and self-assured. Over time that trains the model to lean agreeable and to avoid the flat I do not know that a careful expert would give. It wants to be useful to you, and a confident answer feels more useful than a hedge, even when the hedge was the honest response.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Mirror Effect
&lt;/h2&gt;

&lt;p&gt;Here is a small experiment worth running yourself. Ask Claude the same question twice, once in a confident tone that assumes one answer, and once in a doubtful tone that assumes the opposite. You will often get two different answers, each one leaning toward what you seemed to want. It is partly reflecting you back. This is not a flaw you can prompt away entirely, but knowing it exists changes how you read the response. If you led the witness, the answer is worth less.&lt;/p&gt;

&lt;h2&gt;
  
  
  Confidence Is Not Correctness
&lt;/h2&gt;

&lt;p&gt;The most useful habit you can build is to fully separate how sure it sounds from how likely it is to be right. The tone tells you nothing. A made-up statistic and a real one arrive in exactly the same calm, authoritative voice. Once you stop treating confidence as evidence, you start reading AI answers the way you should read a stranger on the internet who happens to be very articulate. Sometimes right, always fluent, never to be trusted on the strength of tone alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Catch It
&lt;/h2&gt;

&lt;p&gt;You do not need to become a fact-checker for everything. You need a few cheap habits for the things that matter. Ask it for its sources and actually click them, because a fabricated source falls apart the moment you look. Ask it to argue the opposite of what it just told you, and see whether the first answer survives. Change your own wording and ask again, and watch whether the answer flips, which tells you it was reflecting you. And save your suspicion for the things most likely to be invented: specific numbers, dates, names, direct quotes, and anything legal, medical, or financial. Those are exactly where a plausible guess does the most damage.&lt;/p&gt;

&lt;h2&gt;
  
  
  When This Matters and When It Does Not
&lt;/h2&gt;

&lt;p&gt;None of this means Claude is untrustworthy or not worth using. It means you match your caution to the stakes. Brainstorming, drafting, explaining a concept, talking through options: the occasional wrong turn costs you nothing and you would catch it anyway. Anything you are going to act on, sign, send, or publish: verify the specifics yourself. The skill is not distrust. It is knowing which answers you can take at face value and which ones you check, and that judgment is most of what separates people who get burned by AI from people who get real value out of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to Start This Week
&lt;/h2&gt;

&lt;p&gt;The next time Claude gives you a confident answer with a specific fact in it, a number, a date, a source, do not just accept it. Ask it where that came from and check. Do it a few times and you will develop a feel for when it is on solid ground and when it is gliding, and that feel is worth more than any list of rules.&lt;/p&gt;

&lt;p&gt;Treat it like a brilliant, fast, slightly overconfident colleague, and you will get the best out of it without getting caught. If you want to build that judgment properly, our free &lt;a href="https://studiomeyer.academy" rel="noopener noreferrer"&gt;StudioMeyer Academy&lt;/a&gt; has a whole piece on spotting when AI is guessing. Next in the series, we turn to something it is genuinely great at: reading and analyzing your documents.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/why-claude-guesses" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;. It is part of a beginner series on getting real work done with Claude in 2026.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>beginners</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How to Give Claude Context (and Why It Still Matters)</title>
      <dc:creator>Matthias | StudioMeyer</dc:creator>
      <pubDate>Sat, 25 Jul 2026 00:23:51 +0000</pubDate>
      <link>https://dev.to/studiomeyer_io/how-to-give-claude-context-and-why-it-still-matters-368j</link>
      <guid>https://dev.to/studiomeyer_io/how-to-give-claude-context-and-why-it-still-matters-368j</guid>
      <description>&lt;p&gt;Even now that Claude remembers you between conversations, it still starts most chats mostly blank about the specific thing in front of you. The memory holds who you are and how you like to work. It does not automatically hold the contract you are about to paste, the numbers from this quarter, or the three constraints that make this particular decision hard. Those you still have to hand it.&lt;/p&gt;

&lt;p&gt;That distinction, between what Claude carries with it and what you have to give it fresh, is one of the most useful things to understand as a beginner. Get it wrong and you will keep hitting a wall no amount of clever prompting gets you past.&lt;/p&gt;

&lt;p&gt;This is the ninth post in a beginner's series on Claude. The &lt;a href="https://studiomeyer.io/en/blog/what-is-claude-2026" rel="noopener noreferrer"&gt;first one&lt;/a&gt; mapped out the whole tool. This one is about context, the single thing that most often separates a useless answer from a genuinely good one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory Is the Notebook, Context Is the Desk
&lt;/h2&gt;

&lt;p&gt;Think of two different things. Memory is the long-term notebook Claude keeps about you, the preferences and background it carries from one day to the next. Context is the desk it is working on right now, in this conversation, with whatever you have put on it. They are not the same, and confusing them is where a lot of frustration comes from.&lt;/p&gt;

&lt;p&gt;Memory is light and lasting. Context is large and temporary. When you paste a document, that lands on the desk, not in the notebook. When the conversation ends, the desk is cleared. The notebook keeps a few notes about how it went, but the document itself is gone. So the rule of thumb is simple. Anything specific that this task depends on has to be on the desk, because the notebook was never going to hold it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Desk Is Big, but It Is Not Infinite
&lt;/h2&gt;

&lt;p&gt;Claude can hold a genuinely enormous amount in a single conversation, roughly a short book's worth of text at once. That is a lot of desk. But it is not infinite and, more importantly, it is not memory. Two things follow from that.&lt;/p&gt;

&lt;p&gt;The first is that if you pile too much onto the desk, the important part can get lost in the middle. When a conversation runs very long or you dump in fifty pages at once, the one line that actually mattered, buried on page forty, gets less of Claude's attention than the start and the end. It is the same way a person skims a huge document and remembers the opening and the closing better than the middle. More context is not always better. Relevant context is better.&lt;/p&gt;

&lt;p&gt;The second is that when the desk fills up, older parts start to fade from view. In a marathon conversation, something you said near the top may no longer be shaping the answers near the bottom. If a detail from earlier really matters, it is often worth restating it rather than assuming it is still in play.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Wall You Cannot Prompt Your Way Past
&lt;/h2&gt;

&lt;p&gt;There is a moment every user hits. You have written a clear request, you have told Claude who it is for and what format you want, and the answer is still flat. You reach for better wording, and it does not help. That is almost never a prompting problem. It is a context problem.&lt;/p&gt;

&lt;p&gt;The answer is flat because Claude is missing something it needs and no phrasing can invent it. It does not have the actual numbers, so it hedges. It has not seen the document, so it guesses at what is in it. It does not know the constraint that rules out the obvious answer, so it gives you the obvious answer. When you feel yourself rewording the same request for the third time, stop rewording. Ask instead what does it not know that a competent person would need to know here, and then put that on the desk.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Give Context Well
&lt;/h2&gt;

&lt;p&gt;Paste the real thing, not a description of it. The actual email, the actual spreadsheet, the actual paragraph you are stuck on. A summary of the document is you doing the hard part for it and losing detail in the process.&lt;/p&gt;

&lt;p&gt;Be specific about what matters in what you paste. Dropping in a forty page report and asking what do you think wastes most of it. Dropping in the same report and asking which of these three options does section four support gets you something useful.&lt;/p&gt;

&lt;p&gt;Put the context you use over and over into a Project, so you are not re-pasting your brand voice or your product facts every single time. That is exactly what Projects are for, and it is the difference between briefing Claude from scratch daily and working with something that already has your background on the desk.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to Start Fresh and When to Stay
&lt;/h2&gt;

&lt;p&gt;A cluttered desk is a real cost, so knowing when to clear it helps. Start a fresh conversation when you switch to a genuinely different topic, or when a long thread has wandered and is now carrying a lot of irrelevant history that is muddying the answers. Stay in the same conversation when you are building on the same task, refining the same document, or working through steps that depend on each other. The signal that it is time for a fresh start is usually a feeling that Claude is being pulled by something earlier in the chat that no longer applies.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to Start This Week
&lt;/h2&gt;

&lt;p&gt;The next time an answer disappoints you, do not rewrite the prompt. Ask yourself what you know that Claude cannot see, and paste that in. It is a small change in habit, from tuning your words to feeding the desk, and it fixes more bad answers than any prompting trick.&lt;/p&gt;

&lt;p&gt;Context is the work, most of the time. If you want to build the habit properly, our free &lt;a href="https://studiomeyer.academy" rel="noopener noreferrer"&gt;StudioMeyer Academy&lt;/a&gt; walks through it with examples. Next in the series, we go past the temporary desk entirely and into giving Claude a real, lasting memory of your work.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article was originally published on &lt;a href="https://studiomeyer.io/en/blog/giving-claude-context" rel="noopener noreferrer"&gt;studiomeyer.io&lt;/a&gt;. It is part of a beginner series on getting real work done with Claude in 2026.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>beginners</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
