<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vinny Barreca</title>
    <description>The latest articles on DEV Community by Vinny Barreca (@phantom-byte).</description>
    <link>https://dev.to/phantom-byte</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3962211%2F52789747-268f-4ecf-a2e9-76d76b3978e9.jpg</url>
      <title>DEV Community: Vinny Barreca</title>
      <link>https://dev.to/phantom-byte</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/phantom-byte"/>
    <language>en</language>
    <item>
      <title>Anthropic just paid Akamai $11.6 billion over seven years, and the bet is on CPUs, not GPUs</title>
      <dc:creator>Vinny Barreca</dc:creator>
      <pubDate>Sat, 26 Sep 2026 19:10:59 +0000</pubDate>
      <link>https://dev.to/phantom-byte/anthropic-just-paid-akamai-116-billion-over-seven-years-and-the-bet-is-on-cpus-not-gpus-3n3f</link>
      <guid>https://dev.to/phantom-byte/anthropic-just-paid-akamai-116-billion-over-seven-years-and-the-bet-is-on-cpus-not-gpus-3n3f</guid>
      <description>&lt;p&gt;Everyone argues about which GPU buys the frontier. Meanwhile, the biggest cloud deal of the year is for the boring chip that runs your agent's tool calls, browsing, and code execution.&lt;/p&gt;

&lt;p&gt;Agents do not live inside the model forward pass. They live in the harness around it, and that harness is CPU work. That is why a lab would lock up seven years of general-purpose compute while everyone else fights over HBM.&lt;/p&gt;

&lt;p&gt;We covered this before it was news: &lt;a href="https://articles.phantom-byte.com/your-agent-runs-on-a-cpu-nvidia-just-built-one-for-it.html" rel="noopener noreferrer"&gt;the agent workload already moved to the CPU&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>claude</category>
      <category>gpu</category>
    </item>
    <item>
      <title>I Scanned 300 Small Business Domains. 243 Had Broken Mail Records.</title>
      <dc:creator>Vinny Barreca</dc:creator>
      <pubDate>Sun, 20 Sep 2026 16:37:35 +0000</pubDate>
      <link>https://dev.to/phantom-byte/i-scanned-300-small-business-domains-243-had-broken-mail-records-3j04</link>
      <guid>https://dev.to/phantom-byte/i-scanned-300-small-business-domains-243-had-broken-mail-records-3j04</guid>
      <description>&lt;p&gt;I ran a mail-record scanner over 300 small business domains, the kind of local shops that send invoices, quotes, and appointment reminders out of their own domain.&lt;/p&gt;

&lt;p&gt;243 of them, 81 percent, had at least one defect.&lt;/p&gt;

&lt;p&gt;Not a judgment call. A measured defect: no SPF record, an SPF ending in ~all, a DMARC record that enforces nothing, or a DKIM key that has been revoked. The breakdown: 120 domains with no DMARC record at all, 112 with SPF ending in ~all, 73 running DMARC at p=none, 66 with no SPF record at all, and 48 with a DMARC record that collects no reports.&lt;/p&gt;

&lt;p&gt;Eighty-one percent. That is not a rounding error. That is the default state of small business email.&lt;/p&gt;

&lt;p&gt;THE ONE RECORD NOBODY PUBLISHES&lt;/p&gt;

&lt;p&gt;Here is the one I see most and the one almost nobody understands.&lt;/p&gt;

&lt;p&gt;A restaurant domain had this in DNS:&lt;/p&gt;

&lt;p&gt;v=SPF1 include:spf.easywp.com ~all&lt;/p&gt;

&lt;p&gt;and this at _dmarc:&lt;/p&gt;

&lt;p&gt;v=DMARC1; p=none;&lt;/p&gt;

&lt;p&gt;The SPF is fine. The DKIM is fine. But the DMARC policy says p=none, which tells every receiver on earth: if mail claiming to be from this domain fails both checks, do nothing. Deliver it anyway.&lt;/p&gt;

&lt;p&gt;p=none is not enforcement. It is a shrug written in DNS.&lt;/p&gt;

&lt;p&gt;The owner thinks they have DMARC because the record exists. They have the appearance of DMARC. Any spammer with a cheap mail server can send "invoices" from their domain, and every receiver will wave it through, because the record they published literally instructs receivers to accept the failure.&lt;/p&gt;

&lt;p&gt;Worse: this particular record had no rua tag, so even the reporting, the only reason p=none exists as a testing stage, was switched off. The domain was reporting nothing to nobody.&lt;/p&gt;

&lt;p&gt;THE RECORD THAT DOES THE OPPOSITE OF WHAT YOU THINK&lt;/p&gt;

&lt;p&gt;Second most common: SPF ending in ~all.&lt;/p&gt;

&lt;p&gt;v=spf1 include:secureserver.net ~all&lt;/p&gt;

&lt;p&gt;~all is soft fail. It tells receivers: mail from an IP not on this list is suspicious, but you should probably accept it anyway. That is a suggestion. Receivers treat a soft fail as a yellow light and wave most of it through while flagging you for suspicious behavior.&lt;/p&gt;

&lt;p&gt;The fix is -all. Hard fail. Any mail from an unauthorized IP is refused.&lt;/p&gt;

&lt;p&gt;So why does ~all exist everywhere? Because owners are afraid of locking themselves out. Legitimate fear, wrong remedy. You do not flip to -all on day one. You publish DMARC with a rua reporting address first, at p=none, and let the reports tell you every service actually sending as your domain. You will find one or two you forgot about. Then you authorize them, and then you flip.&lt;/p&gt;

&lt;p&gt;The staging is the fix. Nobody does the staging, so everybody sits at ~all forever.&lt;/p&gt;

&lt;p&gt;THE DOMAIN WITH NOTHING&lt;/p&gt;

&lt;p&gt;And then there is the domain with no SPF and no DMARC at all.&lt;/p&gt;

&lt;p&gt;Empty. Both records. No MX record either, so the domain receives no mail and sends with no authentication whatsoever.&lt;/p&gt;

&lt;p&gt;Anyone on earth can send mail as this domain. Your bank, your customers, your own employees, and nobody receiving it has any instruction on what is real. That is not a deliverability problem. That is an open door.&lt;/p&gt;

&lt;p&gt;This was not a rare find. 66 of the 300 domains had no SPF. 120 had no DMARC. If you are one of them, the question is not whether someone is spoofing you yet. The question is what stops them when they start.&lt;/p&gt;

&lt;p&gt;WHY MAIL-TESTER SAYS 10 AND GMAIL SAYS SPAM&lt;/p&gt;

&lt;p&gt;The classic trap. You run mail-tester, it says 10/10, and your invoices still land in junk.&lt;/p&gt;

&lt;p&gt;Mail-tester grades the message. Gmail grades the domain. A perfect score on a single test message tells you nothing about your domain reputation, and it says nothing about whether your DKIM selector is actually the one your provider signs with.&lt;/p&gt;

&lt;p&gt;Half the "10/10 and still spamming" questions I see are a provider mismatch. The domain carries DKIM records for a platform the owner stopped using two years ago, and there are no records for the platform they actually send from today. The stale sender sits in SPF, burning trust. Nobody notices because nothing enforces alignment. A p=none DMARC hides it by design.&lt;/p&gt;

&lt;p&gt;You cannot diagnose that from a message grader. You diagnose it from the records.&lt;/p&gt;

&lt;p&gt;WHAT TO DO TODAY&lt;/p&gt;

&lt;p&gt;Three checks, ten minutes, no tools you have to pay for.&lt;/p&gt;

&lt;p&gt;One. Look up your domain's TXT record. If SPF does not exist or ends in ~all or +all, that is your problem. +all is worse than missing: it authorizes the entire internet.&lt;/p&gt;

&lt;p&gt;Two. Look up _dmarc.yourdomain.com. If the record does not exist, publish one at p=none with a rua pointing at a mailbox you actually check. That is behavior-identical to what you have now; it cannot break sending, and it starts the reporting engine.&lt;/p&gt;

&lt;p&gt;Three. Once a few weeks of reports show every legitimate sender aligning, tighten. ~all to -all. Then p=none to p=quarantine. Then, after another clean window, p=reject. Each stage is its own decision.&lt;/p&gt;

&lt;p&gt;Never jump straight to p=reject. That is how a fix becomes an outage, and it is the reason most owners freeze at p=none and call it done.&lt;/p&gt;

&lt;p&gt;THE UNCOMFORTABLE QUESTION&lt;/p&gt;

&lt;p&gt;I fixed my own domain the same way: staged DMARC, watched the aggregate reports from Google land in my inbox for a day, confirmed every message passing dkim and spf with zero failure reasons, then tightened to -all and p=reject. The Google report confirmed the fix end-to-end. That is the standard, and it is not a high bar.&lt;/p&gt;

&lt;p&gt;So: is your domain one of the 243 or one of the 57?&lt;/p&gt;

&lt;p&gt;You do not know until you look. Go look.&lt;/p&gt;

&lt;p&gt;If you would rather hand the DNS surgery to someone who does it every day, I fix these records the same day, flat $500, and verify the result with real delivered mail: &lt;a href="https://phantom-byte.com/email-deliverability-rescue" rel="noopener noreferrer"&gt;PhantomByte Email Deliverability Rescue Service&lt;/a&gt;&lt;/p&gt;

</description>
      <category>email</category>
      <category>deliverability</category>
      <category>spf</category>
      <category>dns</category>
    </item>
    <item>
      <title>Your Agent's Tool Descriptions Are Costing You 66% Accuracy</title>
      <dc:creator>Vinny Barreca</dc:creator>
      <pubDate>Tue, 15 Sep 2026 18:56:49 +0000</pubDate>
      <link>https://dev.to/phantom-byte/your-agents-tool-descriptions-are-costing-you-66-accuracy-74f</link>
      <guid>https://dev.to/phantom-byte/your-agents-tool-descriptions-are-costing-you-66-accuracy-74f</guid>
      <description>&lt;p&gt;A $4 experiment rewrote three strings per MCP tool and moved SQLite success from 34% to 100%. The model never changed.&lt;/p&gt;

&lt;p&gt;The failure was never the model.&lt;/p&gt;

&lt;p&gt;SQLite strict success rate: 34%. That is a failing grade on a benchmark nobody was running.&lt;/p&gt;

&lt;p&gt;Toolmetry, an open-source project that measures how well agents use MCP tools, took existing MCP server tool descriptions and had an LLM rewrite them. The strings that tell an agent what each tool does and how to call it. That was the entire change.&lt;/p&gt;

&lt;p&gt;The model did not change. The harness did not change. Only the descriptions changed.&lt;/p&gt;

&lt;p&gt;After the rewrite, SQLite hit 100%. Total API spend for the whole experiment was $4.&lt;/p&gt;

&lt;p&gt;This is an interface problem, not a model problem. Tool descriptions are the contract between your agent and the outside world. When the contract is ambiguous, the agent breaks the terms every single time.&lt;/p&gt;

&lt;p&gt;The three failure archetypes&lt;/p&gt;

&lt;p&gt;Toolmetry found three patterns that were killing agent success rates before the rewrite.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Wrong tool confusion. Two tools in the same MCP server have overlapping descriptions, so the agent picks Tool A when it needs Tool. B. It looks like a bug. It is bad documentation.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ritual extra calls. The agent makes a call it believes is a prerequisite, burning tokens and latency for nothing.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Deprecated parameter patterns. The description references parameters or usage patterns that no longer work. The agent follows the instructions faithfully and produces errors.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each one is fixable with better words. Not better models. Better words.&lt;/p&gt;

&lt;p&gt;Wrong tool confusion in practice&lt;/p&gt;

&lt;p&gt;Before the rewrite, the SQLite server's read and write query tools were described almost identically. Most failures were the agent calling a read tool when it needed a write tool.&lt;/p&gt;

&lt;p&gt;Two tools that sound similar to a human sound identical to an agent. After the rewrite, the query tool description stated explicitly what it could and could not do. Success jumped to 100%.&lt;/p&gt;

&lt;p&gt;If two tools could be confused by a smart human, they will be confused by your agent.&lt;/p&gt;

&lt;p&gt;Ritual extra calls and the implied prerequisite&lt;/p&gt;

&lt;p&gt;Agents were calling a list_tables function before every query because the query tool's description said "first enumerate available tables." Implied, not required.&lt;/p&gt;

&lt;p&gt;Before:&lt;/p&gt;

&lt;p&gt;Tool: query&lt;/p&gt;

&lt;p&gt;Description: Execute an SQL query against the database. First enumerate&lt;br&gt;
available tables using list_tables, then construct your query against&lt;br&gt;
those tables.&lt;/p&gt;

&lt;p&gt;The agent called list_tables before every query, even when it already knew the schema from the previous call. Two extra API calls per interaction, doubled latency, zero benefit.&lt;/p&gt;

&lt;p&gt;After:&lt;/p&gt;

&lt;p&gt;Tool: query&lt;/p&gt;

&lt;p&gt;Description: Execute an SQL query against the database. Available tables:&lt;br&gt;
users, orders, products, invoices, and audit_log. Use standard SQL syntax.&lt;br&gt;
This tool handles both reads and writes. No need to call list_tables first.&lt;/p&gt;

&lt;p&gt;The ritual calls stopped immediately. The table list was right there in the description. No extra round trips, no wasted tokens.&lt;/p&gt;

&lt;p&gt;Remove implied prerequisites from your descriptions. If a step is not required, do not imply that it is.&lt;/p&gt;

&lt;p&gt;Deprecated parameters&lt;/p&gt;

&lt;p&gt;The git server had deprecated parameter references baked into its descriptions. Agents were passing old parameter names and getting failures. The documentation was lying about how the API worked.&lt;/p&gt;

&lt;p&gt;After the rewrite, all parameters matched current names. Success went from 75% to 96.7%.&lt;/p&gt;

&lt;p&gt;Version your tool descriptions. When you change an API, update the description in the same commit. A stale description is worse than no description. &lt;/p&gt;

&lt;p&gt;With no description, the agent guesses. With a stale description, the agent confidently does the wrong thing.&lt;/p&gt;

&lt;p&gt;The before and after numbers&lt;/p&gt;

&lt;p&gt;Server  Before  After   Improvement&lt;/p&gt;

&lt;p&gt;SQLite  34.0%   100%    +66.0 points&lt;br&gt;
Memory  61.8%   96.4%   +34.5 points&lt;br&gt;
Git 75.0%   96.7%   +21.7 points&lt;/p&gt;

&lt;p&gt;$4 of API spend for the entire run.&lt;/p&gt;

&lt;p&gt;Now compare that against upgrading to a frontier model. A cheaper model with broken descriptions still fails. An expensive model with broken descriptions still fails. The model is not the bottleneck. The interface is.&lt;/p&gt;

&lt;p&gt;Audit your MCP server in under an hour.&lt;/p&gt;

&lt;p&gt;This costs less than $5 in API spend and takes about an hour.&lt;/p&gt;

&lt;p&gt;List every tool in your MCP server and print each description.&lt;/p&gt;

&lt;p&gt;Read pairs of descriptions side by side. If two could be confused by a smart human, they will be confused by the agent.&lt;/p&gt;

&lt;p&gt;Check for implied prerequisites. Does any description say "first do X" when X is not actually required?&lt;/p&gt;

&lt;p&gt;Check for deprecated parameters. Diff your current API against the description text.&lt;/p&gt;

&lt;p&gt;Rewrite the descriptions with an LLM. Feed it the tool name, the current description, and the real API signature. Ask for a clear, non-overlapping description.&lt;/p&gt;

&lt;p&gt;Re-run your agent test suite and compare success rates.&lt;/p&gt;

&lt;p&gt;Toolmetry ships three commands to automate the loop: measure, optimize, and proxy. The proxy rewrites descriptions on the fly, so you can test improvements before committing anything.&lt;/p&gt;

&lt;p&gt;Why this matters more than model selection&lt;/p&gt;

&lt;p&gt;This is the fourth data point in a pattern that keeps repeating. Three separate papers already showed orchestration design beating model selection by 10x on token cost. &lt;/p&gt;

&lt;p&gt;Structural scaffolding reduces failure rates under a fixed model. Graph-structured context beats flat text. And now tool descriptions alone move a server from 34% to 100%.&lt;/p&gt;

&lt;p&gt;How you structure the agent's control flow, context, and tool interfaces matters more than which model you use. That is a direct challenge to the scaling; all you need is orthodoxy.&lt;/p&gt;

&lt;p&gt;The competitive advantage is shifting from model ownership to orchestration design. The people who understand that will win. The people who keep throwing bigger models at broken interfaces will keep getting 34% success rates and wondering why.&lt;/p&gt;

&lt;p&gt;The bottom line:&lt;/p&gt;

&lt;p&gt;Your tool descriptions are three strings per tool. They cost nothing to write and everything to get wrong.&lt;/p&gt;

&lt;p&gt;Wrong tool confusion, extra ritual calls, and deprecated parameter patterns exist in every MCP server that has never been audited. They exist in yours.&lt;/p&gt;

&lt;p&gt;Audit them today. It takes an hour. It costs $4. It might double your agent's success rate.&lt;/p&gt;

&lt;p&gt;Originally published at &lt;a href="https://articles.phantom-byte.com/tool-descriptions-cost-66-accuracy.html" rel="noopener noreferrer"&gt;articles.phantom-byte.com&lt;/a&gt;. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://phantom-byte.com" rel="noopener noreferrer"&gt;PhantomByte&lt;/a&gt; teaches you to build real AI infrastructure: local AI stacks, autonomous agents, multi-agent orchestration, and custom tools. Step-by-step tutorials you download, follow, and deploy.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
    </item>
    <item>
      <title>The $750 Copilot: Why Your AI Dependency Just Got a Price Tag</title>
      <dc:creator>Vinny Barreca</dc:creator>
      <pubDate>Mon, 01 Jun 2026 07:44:18 +0000</pubDate>
      <link>https://dev.to/phantom-byte/the-750-copilot-why-your-ai-dependency-just-got-a-price-tag-39jn</link>
      <guid>https://dev.to/phantom-byte/the-750-copilot-why-your-ai-dependency-just-got-a-price-tag-39jn</guid>
      <description>&lt;p&gt;Your GitHub Copilot bill just went from $29 to $750/month.&lt;/p&gt;

&lt;p&gt;Microsoft taught developers to "vibe code" (write less and let AI do the work).&lt;/p&gt;

&lt;p&gt;Build entire workflows around that subsidy.&lt;/p&gt;

&lt;p&gt;Then they installed a toll booth.&lt;/p&gt;

&lt;p&gt;This is not a pricing change. It is a dependency trap snapping shut.&lt;/p&gt;

&lt;p&gt;The $750 Copilot: the first honest price tag in a subsidized industry. &lt;a href="https://articles.phantom-byte.com/the-750-copilot-why-your-ai-dependency-just-got-a-price-tag.html" rel="noopener noreferrer"&gt;Full breakdown and the local-first alternative&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>opensource</category>
      <category>news</category>
    </item>
  </channel>
</rss>
