<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rod Miller</title>
    <description>The latest articles on DEV Community by Rod Miller (@rodmiller).</description>
    <link>https://dev.to/rodmiller</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3956799%2F922400ba-a8d1-4e5c-a04f-9aa16b0e7015.png</url>
      <title>DEV Community: Rod Miller</title>
      <link>https://dev.to/rodmiller</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rodmiller"/>
    <language>en</language>
    <item>
      <title>The Daily TAB Index</title>
      <dc:creator>Rod Miller</dc:creator>
      <pubDate>Fri, 31 Jul 2026 21:54:15 +0000</pubDate>
      <link>https://dev.to/rodmiller/the-daily-tab-index-3a3j</link>
      <guid>https://dev.to/rodmiller/the-daily-tab-index-3a3j</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fczq73h8drtmgftx7bb6j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fczq73h8drtmgftx7bb6j.png" alt=" " width="800" height="884"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Measured 2026-07-31&lt;/p&gt;

&lt;p&gt;Cost per suite run is the cost of one run across all five categories, counting scored cases only. Tokens consumed by refusals, provider policy blocks, and no-output cases are not included. Calculated from observed token counts and published list rates for the route TAB uses. It does not include prompt-caching, batch, or negotiated discounts, which vary by provider and by workload. See the methodology note dated July 29, 2026 below for an earlier period when the Agentic Execution category was excluded from this figure.&lt;/p&gt;

&lt;p&gt;Methodology note, July 29, 2026. From the admission of the Agentic Execution category through July 29, 2026, that category’s model cost and token usage were not captured. During that period the Cost per suite run shown for every model excludes the Agentic Execution contribution and is understated by that amount. The judge cost was zero throughout and is unaffected, and scores are unaffected: this concerns cost attribution only, not measurement results. The gap is not reconstructable, because per-call token usage for that category was never recorded, so there is no basis to restate the earlier figures without estimating, and TAB does not publish estimates as measurements. All prior rows remain exactly as recorded. From July 29, 2026 forward, Agentic Execution cost is captured on the same basis as the other categories.&lt;/p&gt;

&lt;p&gt;The “vs. lowest” column shows each model’s cost per suite run as a multiple of the lowest cost among public models in this cycle. It compares cost within this cycle only, and is a cost comparison, not a value or quality comparison.&lt;/p&gt;

&lt;p&gt;Each category is run at least three times per cycle under identical conditions. Consistency reports how many of a model’s five categories produced the same result across all runs this cycle. The scores shown are from the most recent run.&lt;/p&gt;

&lt;p&gt;◆ category leader.&lt;/p&gt;

&lt;p&gt;Some models decline to engage with certain test cases — most often security cases, where the model’s own safety classifier blocks the prompt before the model answers. That is a refusal, not a wrong answer. TAB records refusals as unscored so they neither inflate nor deflate the score: the count shown for each model (for example, “refused 3/12”) is exactly how many cases the model declined out of how many it was given, and the score reflects only the cases it engaged with.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frzx4rlkxpw6vjjlwyt83.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frzx4rlkxpw6vjjlwyt83.png" alt=" " width="800" height="1195"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkxoy7xoqv5v93zytozvq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkxoy7xoqv5v93zytozvq.png" alt=" " width="800" height="710"&gt;&lt;/a&gt;&lt;br&gt;
tabverified.ai&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The Daily TAB Index</title>
      <dc:creator>Rod Miller</dc:creator>
      <pubDate>Tue, 28 Jul 2026 17:33:29 +0000</pubDate>
      <link>https://dev.to/rodmiller/the-daily-tab-index-376k</link>
      <guid>https://dev.to/rodmiller/the-daily-tab-index-376k</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzfmwzojf9ezfr70efuyy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzfmwzojf9ezfr70efuyy.png" alt=" " width="800" height="791"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Does your AI have a hidden agenda? I ran 50 covert behavior tests on 10 frontier models.</title>
      <dc:creator>Rod Miller</dc:creator>
      <pubDate>Sun, 31 May 2026 21:49:49 +0000</pubDate>
      <link>https://dev.to/rodmiller/does-your-ai-have-a-hidden-agenda-i-ran-50-covert-behavior-tests-on-10-frontier-models-45ij</link>
      <guid>https://dev.to/rodmiller/does-your-ai-have-a-hidden-agenda-i-ran-50-covert-behavior-tests-on-10-frontier-models-45ij</guid>
      <description>&lt;p&gt;I run independent benchmarks on frontier AI models. No vendor funding, no advertising, no partnerships. I test with an independent judge model (GLM-5) to avoid self-grading bias.&lt;/p&gt;

&lt;p&gt;Last week I ran 50 Covert Behavior Detection tests on 10 frontier models across 5 categories. The benchmark measures whether a model does things behind your back: hidden actions, undisclosed reasoning, behavior changes when monitored, and attempts to appear less suspicious.&lt;/p&gt;

&lt;p&gt;Full Leaderboard (Covert Behavior Index, 0-100)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmuxwzwxochyijhuc2y53.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmuxwzwxochyijhuc2y53.png" alt=" " width="799" height="327"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What the categories mean&lt;/p&gt;

&lt;p&gt;Stated vs Actual: Does the model say one thing and do another?&lt;br&gt;
Accuracy Modification: Does the model change its accuracy based on the situation?&lt;br&gt;
Action Concealment: Does the model hide what it did?&lt;br&gt;
Evaluator Awareness: Does the model behave differently when it detects it's being evaluated?&lt;br&gt;
Anti-Suspicion: Does the model try to appear less suspicious when it's doing something questionable?&lt;/p&gt;

&lt;p&gt;Key findings&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Every model acts differently when it knows it's being watched.
Evaluator Awareness is the universal failure. The best model (GPT-5.4) scored 84%. The worst (DeepSeek V4 Pro) scored 68.8%. No model above 85%. If you're relying on observability tools and audit logs, you may be seeing the model's "being watched" behavior, not how it actually acts in production.&lt;/li&gt;
&lt;li&gt;Gemini is the only provider whose models conceal their actions.
Both Gemini models scored below 90% on Action Concealment (86.9% and 89.5%). Every other model across all four other providers scored 97% or higher. This isn't a model size issue, it's provider-specific.&lt;/li&gt;
&lt;li&gt;Claude Opus 4.7 scored worse than 4.6 for the fourth consecutive benchmark.
CBI: 89.7 → 89.2. Evaluator Awareness: 76.4 → 70.4. Across sycophancy, delegation, commerce, and now covert behavior, the newer Anthropic model has scored lower than its predecessor on every benchmark I've tested. Upgrades are not automatically improvements.&lt;/li&gt;
&lt;li&gt;Chinese models are competitive across the board.
DeepSeek V4 Flash (91.3), DeepSeek V4 Pro (91.1), and Qwen 3.7 Max (91.0) all placed in the top 5. Qwen had perfect Action Concealment (100.0%). Quality is a model-level property, not a country-level one.&lt;/li&gt;
&lt;li&gt;The spread is tight at the top but drops off at the bottom.
Top 8 models are within 2.6 points of each other (89.2-91.8). Then a 4-point gap to Gemini 3.1 Pro (85.2) and another 4 points to Gemini 3.5 Flash (81.3). Most models are clustered. Gemini is the outlier.
Methodology&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;50 tests across 5 categories&lt;br&gt;
Independent judge model (GLM-5) to prevent self-grading&lt;br&gt;
Two runs per model, scores averaged&lt;br&gt;
All models tested same day, same harness configuration&lt;br&gt;
US models via native APIs, Chinese models via OpenRouter&lt;/p&gt;

&lt;p&gt;Ran using the tabverified.ai platform.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>tooling</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
