<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Khasky</title>
    <description>The latest articles on DEV Community by Khasky (@khasky).</description>
    <link>https://dev.to/khasky</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F141402%2Fefb56248-90d4-46a0-8c60-e473fc4cb987.jpg</url>
      <title>DEV Community: Khasky</title>
      <link>https://dev.to/khasky</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/khasky"/>
    <language>en</language>
    <item>
      <title>Non-English Prompts Cost More Tokens, Context and Money in Most LLMs</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Mon, 05 Oct 2026 04:22:46 +0000</pubDate>
      <link>https://dev.to/khasky/non-english-prompts-cost-more-tokens-context-and-money-in-most-llms-5h5h</link>
      <guid>https://dev.to/khasky/non-english-prompts-cost-more-tokens-context-and-money-in-most-llms-5h5h</guid>
      <description>&lt;p&gt;Claude Opus 5 costs $5 per million input tokens whether a prompt arrives in English or in Russian. What changes between the two languages is how many tokens the same paragraph becomes, and the bill is that count multiplied by the price.&lt;/p&gt;

&lt;p&gt;A recent test on &lt;a href="https://habr.com/ru/articles/1089148/" rel="noopener noreferrer"&gt;Habr&lt;/a&gt; measured the gap across 11 tokenizers. Its numbers line up with research going back to 2023: on most popular models, a language other than English costs more to send, fills the context window sooner and reaches rate limits with less text.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the test worked
&lt;/h2&gt;

&lt;p&gt;A token is the chunk of text a model reads and bills by, usually a piece of a word. The tokenizer is the component that cuts text into tokens with a vocabulary fixed before the model is trained.&lt;/p&gt;

&lt;p&gt;The author, Habr user donseo, wrote 4 parallel text pairs in Russian and English: a technical article, a support chat, a business notice and Python code with comments. The Russian side holds about 2,500 characters. Open tokenizers ran locally through &lt;code&gt;tiktoken&lt;/code&gt; and Hugging Face. For closed models, the script sent each text with a short fixed prompt and the prompt alone, then subtracted the two token counts the API reported. The code and the corpus are public on &lt;a href="https://github.com/donseo-info/ru-token-tax" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; under MIT.&lt;/p&gt;

&lt;h2&gt;
  
  
  Russian against English, same text
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model                      RU/EN tokens   Russian vs English
YandexGPT 5 Lite           0.91x          -9%
GigaChat 3                 0.96x          -4%
Grok 4.7                   1.18x          +18%
GPT-5.x / GPT-6 (o200k)    1.19x          +19%
Gemini 3.7 Flash           1.20x          +20%
GLM-5.3                    1.23x          +23%
DeepSeek V3/V4             1.38x          +38%
Qwen 3                     1.52x          +52%
Kimi K3                    1.86x          +86%
GPT-4 (cl100k)             2.09x          +109%
Claude Opus 5              2.96x          +196%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Claude Opus 5 sits far from the rest at &lt;strong&gt;2.96x&lt;/strong&gt;. Its technical article took 327 tokens in English and 872 in Russian. Even its English is expensive: 3.25 characters per token, against 5.07 for the OpenAI o200k tokenizer behind GPT-5.x. In Russian it falls to 1.03 characters per token, close to one token per letter.&lt;/p&gt;

&lt;p&gt;The Python sample showed the smallest gap on Claude Opus 5, 424 tokens in Russian against 231 in English. Prose is where the premium bites hardest.&lt;/p&gt;

&lt;p&gt;At the same per-token price, Russian text on Claude Opus 5 costs about 3.9x what it costs on GPT-5.5, from the tokenizer alone. That 3.9x compares two models on the same Russian text. Reposts of the table often present it as Russian against English, which overstates the language gap and hides the model gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the tokenizer favors English
&lt;/h2&gt;

&lt;p&gt;A tokenizer builds its vocabulary from a training corpus, and frequent sequences in that corpus get their own tokens. Petrov et al. found that "tokenizers are heavily influenced by the biases of the corpus source". Common English words become one token each. A Cyrillic word breaks into fragments, and in GPT-4's cl100k_base tokenizer some Cyrillic letters take two tokens on their own.&lt;/p&gt;

&lt;p&gt;None of this depends on how well the model understands the language. The same paper puts it plainly: "the unequal treatment of languages arises at the tokenization stage, well before the language model sees any data at all."&lt;/p&gt;

&lt;p&gt;Anthropic's pricing page says the same thing in vendor terms. One token is "approximately 4 characters or 0.75 words in English", and "the exact count varies by language and content type". It also notes that Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text", a change that touches every language.&lt;/p&gt;

&lt;h2&gt;
  
  
  What extra tokens cost
&lt;/h2&gt;

&lt;p&gt;Price, context and rate limits in an LLM API are all counted in tokens.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Price. Anthropic bills per million tokens, $5 input and $25 output on Claude Opus 5. A prompt that becomes 2.96x the tokens costs 2.96x as much to send.&lt;/li&gt;
&lt;li&gt;Answers. Output is tokenized the same way, and on Claude Opus 5 an output token costs 5x an input token. A long answer in Russian carries the premium at the higher rate.&lt;/li&gt;
&lt;li&gt;Context window. The window is a token budget. At 2.96x, it holds about a third as much Russian text as English, so long documents and chat histories run out of room sooner.&lt;/li&gt;
&lt;li&gt;Rate limits. Anthropic measures them in input and output tokens per minute. A team working in Russian reaches the cap with less work done. 💸&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Not a Russian problem
&lt;/h2&gt;

&lt;p&gt;Petrov et al. measured parallel text in many languages for NeurIPS 2023. The tokenizer behind ChatGPT and GPT-4 used about 1.6x the tokens of English for Italian, 2.6x for Bulgarian and 3x for Arabic. For some languages the gap reached 15x. Their Russian figure on cl100k_base was 2.49x, on a different corpus from the Habr test. The authors' conclusion about money is direct: "the tokenization premiums discussed in Section 4 directly map to cost premiums."&lt;/p&gt;

&lt;p&gt;Newer tokenizers have narrowed the gap for some vendors. OpenAI's o200k came in at 1.19x for Russian in the Habr test. The two studies used different texts, so this shows a direction rather than a precise trend.&lt;/p&gt;

&lt;p&gt;Price is not the only cost. Ahia et al. at EMNLP 2023 found that speakers of many supported languages "are overcharged while obtaining poorer results". Anthropic publishes scores relative to English for Claude Sonnet 4.5: Spanish at 98.2%, Chinese at 96.9%, Hindi at 96.7%, Swahili at 91.1% and Yoruba at 79.7%. Major languages stay close to English, smaller ones fall behind. Russian is not in that table. Wendler et al. add a hint about why English stays the default: inside the models they studied, "the abstract 'concept space' lies closer to English than to other languages".&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the rule breaks
&lt;/h2&gt;

&lt;p&gt;YandexGPT 5 Lite and GigaChat 3 use fewer tokens for Russian than for English. They are the two models in the test built for Russian, and their tokenizers encode Russian at least as compactly as English. That is the same mechanism from the other side: the language a tokenizer saw most is the language it encodes cheaply.&lt;/p&gt;

&lt;p&gt;The test has known limits. Its author lists them: the corpus is small, and the coefficients may move by about 5% on other texts. Closed-model counts came through third-party API gateways. The GigaChat figure uses the open GigaChat 3 tokenizer, which may differ from the production model. Claude's tokenizer is not public, so its number rests on API usage counts.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Write system prompts, instructions and tool schemas in English when the model tokenizes your language poorly. The Habr author recommends exactly this.&lt;/li&gt;
&lt;li&gt;Compare models by the cost of a typical request on your own text, not by the price per million tokens.&lt;/li&gt;
&lt;li&gt;Count before choosing. Anthropic's token counting endpoint is free, and &lt;code&gt;tiktoken&lt;/code&gt; covers OpenAI's tokenizers.&lt;/li&gt;
&lt;li&gt;For workloads that must stay in Russian, the Russian-built models start with a tokenizer advantage. 🧭&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;A price per token is only comparable between languages when the token counts are.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For anyone paying per token, the language of the prompt is a pricing decision, and on Claude Opus 5 a Russian prompt buys about a third of the text the same money buys in English.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://habr.com/ru/articles/1089148/" rel="noopener noreferrer"&gt;Habr: Russian text in Claude Opus 5 costs 3.9x more than in GPT-5.5 at the same per-token price (in Russian)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/donseo-info/ru-token-tax" rel="noopener noreferrer"&gt;ru-token-tax: code and corpus of the Habr test&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2305.15425" rel="noopener noreferrer"&gt;Petrov et al., Language Model Tokenizers Introduce Unfairness Between Languages&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://aclanthology.org/2023.emnlp-main.614/" rel="noopener noreferrer"&gt;Ahia et al., Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2402.10588" rel="noopener noreferrer"&gt;Wendler et al., Do Llamas Work in English? On the Latent Language of Multilingual Transformers&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;Anthropic: Pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.claude.com/docs/en/build-with-claude/multilingual-support" rel="noopener noreferrer"&gt;Anthropic: Multilingual support&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.claude.com/docs/en/api/rate-limits" rel="noopener noreferrer"&gt;Anthropic: Rate limits&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.claude.com/docs/en/build-with-claude/token-counting" rel="noopener noreferrer"&gt;Anthropic: Token counting&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
khasky — LinkedIn / GitHub / Patreon / Bluesky / Mastodon / Medium / Devto&lt;br&gt;&lt;br&gt;
khaskydev — X / Threads / Instagram / Pinterest / Tumblr / Facebook / VK&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>tokenization</category>
      <category>promptengineering</category>
    </item>
    <item>
      <title>Fable Becomes the Advisor to Opus 5.5 in Claude Code</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Sat, 03 Oct 2026 01:10:37 +0000</pubDate>
      <link>https://dev.to/khasky/fable-becomes-the-advisor-to-opus-55-in-claude-code-3ekb</link>
      <guid>https://dev.to/khasky/fable-becomes-the-advisor-to-opus-55-in-claude-code-3ekb</guid>
      <description>&lt;p&gt;With Opus 5.5 as the main model in a Claude Code session, a Fable quota can still go to work without Fable taking over the session. The &lt;code&gt;/advisor fable&lt;/code&gt; command connects Fable as an advisor that Opus consults on its own. Opus keeps writing the code.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the advisor does
&lt;/h2&gt;

&lt;p&gt;The advisor is a second model attached to the session, and Opus decides when to call it. It tends to ask before committing to an approach, when the same error keeps coming back, and before declaring a task done. That timing is model-driven rather than rule-based.&lt;/p&gt;

&lt;p&gt;Each call hands Fable the full conversation, with every tool call and its result. Fable returns guidance, and Opus applies it before it continues. With Opus 5 or Opus 5.5 as the main model, Claude Code accepts Fable or Opus 5 and later in the advisor role.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning it on
&lt;/h2&gt;

&lt;p&gt;On some plans Fable bills to usage credits, and Claude Code will not apply Fable as the advisor until a one-time consent is given. The consent comes from &lt;code&gt;/model fable&lt;/code&gt;, and only after it does &lt;code&gt;/advisor fable&lt;/code&gt; take effect:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/model fable
/advisor fable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Accepting the consent also makes Fable the selected main model. Switch back to Opus 5.5 with &lt;code&gt;/model&lt;/code&gt; so the pairing stays Opus at the keyboard and Fable on call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it costs less
&lt;/h2&gt;

&lt;p&gt;Fable is called at decision points rather than on every turn. A session paired this way typically costs less than one that runs the stronger model throughout. Each call still bills Fable's tokens at Fable's rates, on top of what Opus uses. 🧾&lt;/p&gt;

&lt;p&gt;Turning the advisor on or off mid-session does not invalidate the main model's prompt cache, unlike a model switch. Fable's own read is not cached, though. Every call processes the full transcript anew.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it runs
&lt;/h2&gt;

&lt;p&gt;The advisor is a server-side tool, so it needs the Anthropic API. It is not available on Amazon Bedrock, Claude Platform on AWS, Google Cloud's Agent Platform or Microsoft Foundry.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the advisor stays silent
&lt;/h2&gt;

&lt;p&gt;Two environment variables keep it off. &lt;code&gt;CLAUDE_CODE_DISABLE_ADVISOR_TOOL&lt;/code&gt; disables the tool entirely, and the &lt;code&gt;/advisor&lt;/code&gt; command disappears with it. &lt;code&gt;DISABLE_TELEMETRY&lt;/code&gt; stops Claude Code from fetching the feature flag that turns the advisor on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When no advice ever arrives, check those two variables first.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A setup that uses it
&lt;/h2&gt;

&lt;p&gt;Opus 5.5 on high effort runs the session. Three subagents on medium effort take the routine work: one reads the code, one makes edits with tests, one looks up documentation. Fable waits on the side and comes in when the work needs a second look. 👀&lt;/p&gt;

&lt;p&gt;Subagents inherit the configured advisor and run the same pairing check against their own model. Each one can set its own &lt;code&gt;effort&lt;/code&gt; in its frontmatter.&lt;/p&gt;

&lt;p&gt;For an Opus 5.5 session on the Anthropic API, the advisor is the cheaper way to put Fable's judgment into the work, and keeping Fable as the main model now needs a reason beyond the occasional hard call.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://code.claude.com/docs/en/advisor" rel="noopener noreferrer"&gt;Advisor tool in Claude Code&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.claude.com/docs/en/sub-agents" rel="noopener noreferrer"&gt;Subagents and their frontmatter&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khasky&lt;/strong&gt; — LinkedIn / GitHub / Patreon / Bluesky / Mastodon / Medium / Devto&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khaskydev&lt;/strong&gt; — X / Threads / Instagram / Pinterest / Tumblr / Facebook / VK&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>anthropic</category>
      <category>claudeai</category>
      <category>aicoding</category>
    </item>
    <item>
      <title>Rejudge Replaces Self-Review With 3 Independent Models and a Judge</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Thu, 01 Oct 2026 21:51:11 +0000</pubDate>
      <link>https://dev.to/khasky/rejudge-replaces-self-review-with-3-independent-models-and-a-judge-1kac</link>
      <guid>https://dev.to/khasky/rejudge-replaces-self-review-with-3-independent-models-and-a-judge-1kac</guid>
      <description>&lt;p&gt;I have a hard time trusting an AI coding agent to review code it just wrote.&lt;/p&gt;

&lt;p&gt;The model already made the decision that the implementation was good enough to produce.&lt;/p&gt;

&lt;p&gt;Then we ask it to inspect that same implementation using many of the same learned habits and assumptions.&lt;/p&gt;

&lt;p&gt;Sometimes it catches mistakes. Sometimes it revalidates them. A fresh session helps. A different model helps more.&lt;/p&gt;

&lt;p&gt;But once different models disagree, somebody still has to decide which review to trust.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/syabro/rejudge" rel="noopener noreferrer"&gt;Rejudge&lt;/a&gt; makes that disagreement-resolution step explicit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture
&lt;/h2&gt;

&lt;p&gt;The same review request goes to three reviewers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;reviewer A
reviewer B
reviewer C
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each works in an isolated context. They do not see each other's reasoning, tools, or conclusions. After all three finish, a separate judge receives their reports. The judge can ask follow-up questions when the panel disagrees. Then it writes one final answer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;same question
    |
    +--&amp;gt; reviewer A
    +--&amp;gt; reviewer B
    +--&amp;gt; reviewer C
             |
           judge
             |
        final answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Independence happens before collaboration.&lt;/p&gt;

&lt;h2&gt;
  
  
  How reviewers inspect the code
&lt;/h2&gt;

&lt;p&gt;By default, reviewer tools include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;read
grep
find
ls
git_diff
web_search (if the host provides one)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The normal reviewers do not get:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;edit
write
bash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a sensible default for code review. 🔒&lt;/p&gt;

&lt;h2&gt;
  
  
  How judge makes a decision
&lt;/h2&gt;

&lt;p&gt;The judge gets no workspace access. It sees the three reports.&lt;/p&gt;

&lt;p&gt;If those reports conflict, it can call &lt;code&gt;ask_panel&lt;/code&gt; and request clarification.&lt;/p&gt;

&lt;p&gt;That means the judge is not a hidden fourth reviewer. Its role is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;compare
challenge
adjudicate
synthesize
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Practical usage
&lt;/h2&gt;

&lt;p&gt;Install (it needs Node.js 22.19.0 or newer):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; rejudge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Review a diff:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git diff | rejudge &lt;span class="s2"&gt;"review this change"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ask a targeted question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;rejudge &lt;span class="s2"&gt;"does this migration need a lock?"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Resume later:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;rejudge &lt;span class="nt"&gt;--resume&lt;/span&gt; &amp;lt;run-id&amp;gt; &lt;span class="s2"&gt;"what about the rollback path?"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The answer goes to stdout, while progress, the config in use and the run ID go to stderr, so redirecting stdout to a file leaves only the answer. A resumed run reopens the same sessions, and the new question goes to the judge first. The reviewers hear it only if the judge calls &lt;code&gt;ask_panel&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coding-agent integration
&lt;/h2&gt;

&lt;p&gt;Rejudge supports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;CLI&lt;/li&gt;
&lt;li&gt;native Pi tool&lt;/li&gt;
&lt;li&gt;Agent Skill for coding agents outside Pi&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agent Skill install:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add syabro/rejudge &lt;span class="nt"&gt;-g&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The skills are a separate copy, so the README says to refresh them after each Rejudge release with &lt;code&gt;npx skills update -g -y&lt;/code&gt;. Inside Pi, the extension is one more line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pi &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;npm root &lt;span class="nt"&gt;-g&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;/rejudge"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rejudge runs on Pi and reads its provider settings, so a key Pi already accepts works here too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The models are configurable
&lt;/h2&gt;

&lt;p&gt;Rejudge is not tied to one fixed provider combination. The config has a reviewer list and a separate judge model, with two reviewers as the minimum. Each model also gets a reasoning level from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;minimal
low
medium
high
xhigh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The global file lives at &lt;code&gt;~/.config/rejudge/config.json&lt;/code&gt;, and a &lt;code&gt;.rejudge/config.json&lt;/code&gt; in the project wins over it. That makes it possible to build a genuinely mixed panel.&lt;/p&gt;

&lt;h2&gt;
  
  
  There is an unsafe mode
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;--unsafe&lt;/code&gt; / &lt;code&gt;--full&lt;/code&gt; gives reviewers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;edit
write
bash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The docs explicitly say this is &lt;strong&gt;not a sandbox&lt;/strong&gt;. The judge still gets only &lt;code&gt;ask_panel&lt;/code&gt;. For review-only work, I would stay read-only.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multiple providers mean multiple privacy boundaries
&lt;/h2&gt;

&lt;p&gt;Every reviewer model sees the request. Anything a reviewer reads becomes part of that provider session.&lt;/p&gt;

&lt;p&gt;The judge does not inspect the workspace directly, but reviewer reports may quote code.&lt;/p&gt;

&lt;p&gt;Read-only tools stop local changes. They do not keep file contents private, and the README warns that instructions hidden in a request or in a file a reviewer opens can steer what it reads and reports.&lt;/p&gt;

&lt;p&gt;Runs also leave records. Sessions are written to &lt;code&gt;${TMPDIR}/rejudge/runs/&amp;lt;run-id&amp;gt;/&lt;/code&gt; while a run executes, and cleanup after roughly 24 hours is best-effort. The &lt;code&gt;debugLog&lt;/code&gt; option is off by default, and turning it on writes full model thinking to &lt;code&gt;.rejudge/logs/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;For proprietary repositories, this needs to be acceptable before running the panel.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cost
&lt;/h2&gt;

&lt;p&gt;A fresh review starts with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;3 reviewer calls
+ 1 judge call
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then add tool loops, retries, recovery, and judge follow-ups. Rejudge has no spending cap of its own. This is not a free accuracy multiplier. It is a compute tradeoff. 💸&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I would use it
&lt;/h2&gt;

&lt;p&gt;I would reserve it for changes where mistakes are expensive:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;database migrations
locking/concurrency
authentication
authorization
permissions
security-sensitive code
data deletion
rollback logic
complex refactors
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rejudge is a structured independent second opinion. For consequential code, that can be worth the extra cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/syabro/rejudge" rel="noopener noreferrer"&gt;Rejudge on GitHub&lt;/a&gt;: README, install, config and the privacy and cost notes&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://rejudge.syabro.com" rel="noopener noreferrer"&gt;Rejudge site&lt;/a&gt;: overview and demo&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
khasky — LinkedIn / Patreon / GitHub / Bluesky / Mastodon&lt;br&gt;&lt;br&gt;
khaskydev — X / Threads / Instagram / Pinterest / Facebook&lt;/p&gt;

</description>
      <category>rejudge</category>
      <category>codereview</category>
      <category>aiagents</category>
      <category>llm</category>
    </item>
    <item>
      <title>Gemini 4 Argon Beats GPT-6 Astra and Claude on Most Benchmarks</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Thu, 01 Oct 2026 05:47:12 +0000</pubDate>
      <link>https://dev.to/khasky/gemini-4-argon-beats-gpt-6-astra-and-claude-on-most-benchmarks-1nap</link>
      <guid>https://dev.to/khasky/gemini-4-argon-beats-gpt-6-astra-and-claude-on-most-benchmarks-1nap</guid>
      <description>&lt;p&gt;Google DeepMind has announced Gemini 4 Argon, a frontier model aimed at long software engineering jobs, legal and finance work, and cyber defense. In Google's own evaluation it beats GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on most of the benchmarks it reports. Almost nobody can use it yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Argon comes first
&lt;/h2&gt;

&lt;p&gt;The comparison table covers 19 rows across knowledge work, agentic coding, ML engineering, science and math, long context, computer use, multimodal understanding and cybersecurity. Argon finishes first or tied for first on 14 of them.&lt;/p&gt;

&lt;p&gt;The number Google leads with is DeepSWE v1.1, a set of real-world software engineering tasks. Argon scores 77.9% there, against 74.2% for Opus 5.5, 74.1% for Astra and 67.4% for Fable 5.1. The widest gap sits on Harvey's Legal Agent Benchmark, where Argon reaches 19.6% and none of the other three gets past 6.7%. 🎯&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Benchmark                          Argon   Astra   Fable 5.1  Opus 5.5
Vals Index                         68.9    63.1    65.8       67.0
AutomationBench                    51.3    41.4    31.4       42.5
Vals Finance Agent v2              65.4    53.5    58.9       58.6
Harvey's Legal Agent Benchmark     19.6     5.4     6.7        3.8
DeepSWE v1.1                       77.9    74.1    67.4       74.2
FrontierSWE v2                     55.0    65.5    56.3       62.3
Vibe Code Bench                    91.9    89.6    90.3       90.3
Terminal-bench 4.0                 57.4    58.2    57.9       66.4
PostTrainBench                     45.3    44.3    40.2       49.3
Terminal-Bench Science 0.1         57.6    68.1    52.6       63.3
LABBench 2                         88.8    85.4    68.6       73.1
RiemannBench                       76.0    72.0    65.6       69.6
GraphWalks BFS, up to 128k         99.7    98.7    91.4       90.6
GraphWalks BFS, 256k to 1M         84.2    71.8    65.0       66.8
Agent's Last Exam (pass rate)      39.5    34.2    -          38.2
OSWorld-2.0 (offline, partial)     69.2    72.6    -          -
Chartography                       71.6    71.0    46.2       66.3
LVBench                            91.7    87.5    79.7       83.7
CWE-bench v1                       68.0    68.0    58.0       67.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All values are percentages from Google's table. GraphWalks reports F1, and OSWorld-2.0 uses an offline subset with partial scores.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it does not
&lt;/h2&gt;

&lt;p&gt;Five rows go to someone else, and one is a tie. Astra keeps FrontierSWE v2 with 65.5% against Argon's 55.0%, Terminal-Bench Science 0.1 with 68.1% against 57.6%, and the OSWorld-2.0 offline subset with 72.6% against 69.2%. Opus 5.5 keeps Terminal-bench 4.0 with 66.4% against 57.4%, and PostTrainBench with 49.3% against 45.3%. On CWE-bench v1 Argon and Astra share first place at 68.0%.&lt;/p&gt;

&lt;p&gt;Every figure here comes from Google's own runs, and the methodology page is the place to check how the rival models were set up.&lt;/p&gt;

&lt;h2&gt;
  
  
  One million tokens in one answer
&lt;/h2&gt;

&lt;p&gt;Argon can write up to 1M output tokens in a single response. The previous limit was 64K, so one answer can now run about 15 times longer. Google ties this to the work the model targets: long, multi-step jobs where the reasoning has to hold across the whole run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The price now and later
&lt;/h2&gt;

&lt;p&gt;The introductory API price is $2 per 1M input tokens and $10 per 1M output tokens, and cached input costs 95% less than regular input. When the introductory period ends, the rate becomes $4 and $20.&lt;/p&gt;

&lt;p&gt;GPT-6 Astra lists $10 per 1M input tokens and $50 per 1M output tokens at its standard short-context rate. That puts Argon at 5x cheaper than Astra on both sides today and 2.5x cheaper after the increase. The comparison is token price against token price, and a model that writes longer answers can still cost more per finished task. 💸&lt;/p&gt;

&lt;h2&gt;
  
  
  What Google already used it for
&lt;/h2&gt;

&lt;p&gt;Inside Google, Argon has been working on real projects before release. Quantum computing researchers used it to optimize spacetime resources, counted as qubits times gates, and in one case it beat the published baseline by 40% within minutes.&lt;/p&gt;

&lt;p&gt;A team of Argon agents read fleet-wide profiling telemetry across Google's data centers and applied memory optimizations on its own, freeing more than 300 TiB of memory.&lt;/p&gt;

&lt;p&gt;Argon agents are also moving C and C++ codebases to Rust. Early reports put that at tens of thousands of lines, which holds for core libraries such as re2 and libgav1, while the largest job, the Fuchsia Zircon kernel, runs past 800K lines. The libgav1 result is a memory-safe video decoder with identical output that runs 2.7x faster than the earlier Rust port.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who can use it
&lt;/h2&gt;

&lt;p&gt;Right now Argon is going to a set of trusted cyber defenders through Google's Fairwind Program. Paid API customers and Google AI Ultra subscribers come next, followed by developers, enterprises and consumers more broadly. Google has given no date.&lt;/p&gt;

&lt;p&gt;Until Argon reaches the paid API, the table and the $2 price are something to plan a test around, and nobody outside Fairwind can yet check whether those 14 first places hold up.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/" rel="noopener noreferrer"&gt;Google: Gemini 4 Argon, our next era of frontier intelligence&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://deepmind.google/models/evals-methodology/gemini-4-argon" rel="noopener noreferrer"&gt;Google DeepMind: Gemini 4 Argon evaluation methodology&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developers.openai.com/api/docs/pricing" rel="noopener noreferrer"&gt;OpenAI API pricing&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khasky&lt;/strong&gt; — &lt;a href="https://www.linkedin.com/in/khasky/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; / &lt;a href="https://github.com/khasky" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; / &lt;a href="https://www.patreon.com/khasky" rel="noopener noreferrer"&gt;Patreon&lt;/a&gt; / &lt;a href="https://bsky.app/profile/khasky.bsky.social" rel="noopener noreferrer"&gt;Bluesky&lt;/a&gt; / &lt;a href="https://mastodon.social/@khasky" rel="noopener noreferrer"&gt;Mastodon&lt;/a&gt; / &lt;a href="https://medium.com/@khasky" rel="noopener noreferrer"&gt;Medium&lt;/a&gt; / &lt;a href="https://dev.to/khasky"&gt;Devto&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khaskydev&lt;/strong&gt; — &lt;a href="https://x.com/khaskydev" rel="noopener noreferrer"&gt;X&lt;/a&gt; / &lt;a href="https://www.threads.com/@khaskydev" rel="noopener noreferrer"&gt;Threads&lt;/a&gt; / &lt;a href="https://www.instagram.com/khaskydev/" rel="noopener noreferrer"&gt;Instagram&lt;/a&gt; / &lt;a href="https://www.pinterest.com/khaskydev/" rel="noopener noreferrer"&gt;Pinterest&lt;/a&gt; / &lt;a href="https://www.tumblr.com/khaskydev" rel="noopener noreferrer"&gt;Tumblr&lt;/a&gt; / &lt;a href="https://www.facebook.com/khaskydev/" rel="noopener noreferrer"&gt;Facebook&lt;/a&gt; / &lt;a href="https://vk.ru/khaskydev" rel="noopener noreferrer"&gt;VK&lt;/a&gt;&lt;/p&gt;

</description>
      <category>gemini</category>
      <category>googledeepmind</category>
      <category>llm</category>
      <category>ai</category>
    </item>
    <item>
      <title>Free Email Forwarding for Your Personal Domain With Cloudflare Email Routing</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Tue, 29 Sep 2026 07:02:31 +0000</pubDate>
      <link>https://dev.to/khasky/free-email-forwarding-for-your-personal-domain-with-cloudflare-email-routing-37md</link>
      <guid>https://dev.to/khasky/free-email-forwarding-for-your-personal-domain-with-cloudflare-email-routing-37md</guid>
      <description>&lt;p&gt;Forward any address at your domain, such as &lt;a href="mailto:hello@example.com"&gt;hello@example.com&lt;/a&gt;, to an external inbox. For a public address that only has to receive mail, I would put it on Cloudflare Email Routing: the messages land in the inbox I already read, and forwarding to a verified destination costs nothing on any plan. 💸&lt;/p&gt;

&lt;p&gt;Prerequisites: the domain is on Cloudflare DNS (nameservers delegated, DNS Setup = Full), and it has no MX records from another mail host. Email Routing owns the MX set.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Verify the destination address
&lt;/h2&gt;

&lt;p&gt;Compute &amp;gt; Email Service &amp;gt; Email Routing &amp;gt; Destination Addresses &amp;gt; Add address. Enter the external inbox (for example, &lt;code&gt;you@gmail.com&lt;/code&gt;). Cloudflare emails a confirmation link, and you click it.&lt;/p&gt;

&lt;p&gt;The row must read Verified. Until it does, any rule that points to the address stays disabled.&lt;/p&gt;

&lt;p&gt;Destination addresses are account-wide: once verified, any zone in the account can target it.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Enable routing and add the DNS records
&lt;/h2&gt;

&lt;p&gt;Open the zone's Email Routing &amp;gt; Settings &amp;gt; DNS records &amp;gt; Add missing records. Cloudflare writes these records and locks them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Type  Name                             Value
MX    example.com                      route1.mx.cloudflare.net (+ route2, route3)
TXT   example.com                      v=spf1 include:_spf.mx.cloudflare.net ~all
TXT   cf2024-1._domainkey.example.com  v=DKIM1; h=sha256; k=rsa; p=...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A locked record cannot be edited or deleted from the DNS page until you unlock it in the Email Routing settings.&lt;/p&gt;

&lt;p&gt;Keep one SPF TXT record per name. If the domain already has one, the merge is yours to do: Cloudflare's example combines both includes in a single record, &lt;code&gt;v=spf1 include:_spf.mx.cloudflare.net include:_spf.google.com ~all&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Create the routing rule
&lt;/h2&gt;

&lt;p&gt;Email Routing &amp;gt; Routing rules &amp;gt; Create routing rule:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Email pattern: &lt;code&gt;hello&lt;/code&gt; @ &lt;code&gt;example.com&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Action: Send to an email&lt;/li&gt;
&lt;li&gt;Destination: the verified address&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Save. The rule must show Active.&lt;/p&gt;

&lt;p&gt;The pattern takes any local part, not only &lt;code&gt;hello&lt;/code&gt;: &lt;code&gt;sales&lt;/code&gt;, &lt;code&gt;support&lt;/code&gt; or your own name work the same way, one rule per address.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Decide the catch-all
&lt;/h2&gt;

&lt;p&gt;The catch-all in Routing rules defines what happens to every other address at the domain. Enabled with Send to an email, it forwards all of them, typos such as &lt;code&gt;ifno@example.com&lt;/code&gt; included. Enabled with Drop, it discards them silently.&lt;/p&gt;

&lt;p&gt;I would start with it off and turn on forwarding only if people keep misspelling the real address.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Verify
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# MX must answer from the Cloudflare route hosts&lt;/span&gt;
curl &lt;span class="nt"&gt;-sH&lt;/span&gt; &lt;span class="s1"&gt;'accept: application/dns-json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'https://cloudflare-dns.com/dns-query?name=example.com&amp;amp;type=MX'&lt;/span&gt;

&lt;span class="c"&gt;# SPF&lt;/span&gt;
curl &lt;span class="nt"&gt;-sH&lt;/span&gt; &lt;span class="s1"&gt;'accept: application/dns-json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'https://cloudflare-dns.com/dns-query?name=example.com&amp;amp;type=TXT'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then send a real message to &lt;a href="mailto:hello@example.com"&gt;hello@example.com&lt;/a&gt; and confirm it lands in the destination inbox. 📬&lt;/p&gt;

&lt;p&gt;Email Routing &amp;gt; Activity Log shows the decision for each message: Forwarded, Dropped, Rejected or Delivery failed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Notes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Forwarding only. A routing rule receives mail and does not send it, so a reply from the destination inbox goes out through that provider, from its address. Sending from the domain is a separate product, Cloudflare Email Sending, which is still in beta.&lt;/li&gt;
&lt;li&gt;Negative DNS answers stay cached for the zone's SOA minimum TTL, 1800 seconds (30 minutes) by default on Cloudflare. If you queried MX before the records existed, your resolver may keep returning empty. Query a different resolver or go over DoH to confirm.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Five steps give a domain working addresses with no mailbox bill, and I would pay for a mailbox only on the day replies have to come from the domain itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://developers.cloudflare.com/email-routing/" rel="noopener noreferrer"&gt;Email Routing overview and pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developers.cloudflare.com/email-routing/get-started/enable-email-routing/" rel="noopener noreferrer"&gt;Enable Email Routing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developers.cloudflare.com/email-routing/setup/email-routing-dns-records/" rel="noopener noreferrer"&gt;Email Routing DNS records&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developers.cloudflare.com/email-routing/setup/email-routing-addresses/" rel="noopener noreferrer"&gt;Destination addresses and routing rules&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developers.cloudflare.com/1.1.1.1/encryption/dns-over-https/make-api-requests/dns-json/" rel="noopener noreferrer"&gt;DNS over HTTPS JSON queries&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khasky&lt;/strong&gt; — &lt;a href="https://www.linkedin.com/in/khasky/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; / &lt;a href="https://github.com/khasky" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; / &lt;a href="https://www.patreon.com/khasky" rel="noopener noreferrer"&gt;Patreon&lt;/a&gt; / &lt;a href="https://bsky.app/profile/khasky.bsky.social" rel="noopener noreferrer"&gt;Bluesky&lt;/a&gt; / &lt;a href="https://mastodon.social/@khasky" rel="noopener noreferrer"&gt;Mastodon&lt;/a&gt; / &lt;a href="https://medium.com/@khasky" rel="noopener noreferrer"&gt;Medium&lt;/a&gt; / &lt;a href="https://dev.to/khasky"&gt;Devto&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khaskydev&lt;/strong&gt; — &lt;a href="https://x.com/khaskydev" rel="noopener noreferrer"&gt;X&lt;/a&gt; / &lt;a href="https://www.threads.com/@khaskydev" rel="noopener noreferrer"&gt;Threads&lt;/a&gt; / &lt;a href="https://www.instagram.com/khaskydev/" rel="noopener noreferrer"&gt;Instagram&lt;/a&gt; / &lt;a href="https://www.pinterest.com/khaskydev/" rel="noopener noreferrer"&gt;Pinterest&lt;/a&gt; / &lt;a href="https://www.tumblr.com/khaskydev" rel="noopener noreferrer"&gt;Tumblr&lt;/a&gt; / &lt;a href="https://www.facebook.com/khaskydev/" rel="noopener noreferrer"&gt;Facebook&lt;/a&gt; / &lt;a href="https://vk.ru/khaskydev" rel="noopener noreferrer"&gt;VK&lt;/a&gt;&lt;/p&gt;

</description>
      <category>cloudflare</category>
      <category>emailrouting</category>
      <category>dns</category>
      <category>email</category>
    </item>
    <item>
      <title>Claude Opus 5.5 Reads More Human With 17% Shorter Sentences</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Mon, 28 Sep 2026 00:51:18 +0000</pubDate>
      <link>https://dev.to/khasky/claude-opus-55-reads-more-human-with-17-shorter-sentences-189o</link>
      <guid>https://dev.to/khasky/claude-opus-55-reads-more-human-with-17-shorter-sentences-189o</guid>
      <description>&lt;p&gt;Arena.ai measured how Claude Opus 5.5 writes against Opus 5, the version before it, and a large part of the change sits in two punctuation marks it calls familiar tells of AI writing. Opus 5.5 has nearly stopped using one of them. 📉&lt;/p&gt;

&lt;h2&gt;
  
  
  How Arena measured it
&lt;/h2&gt;

&lt;p&gt;The comparison uses high-reasoning answers from Text Arena. BleepingComputer, which reported on the analysis, dates those answers to August and September 2026. Arena tracked 12 writing measures. Ten moved in what it calls a better direction.&lt;/p&gt;

&lt;p&gt;Neither the sample size nor the full method has been published, so the results describe Arena's set of answers and make no claim about all Opus 5.5 output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Em dashes and semicolons
&lt;/h2&gt;

&lt;p&gt;In the Opus 5 answers, the em dash appeared 15.2 times per 1,000 words. Opus 5.5 brought that down to 0.8, about 95% fewer. Semicolons followed a smaller slide, from 6.10 to 1.64 per 1,000 words, which works out to about 73% fewer.&lt;/p&gt;

&lt;p&gt;Both marks join clauses inside one sentence, and sentence length moved too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Longer answers, shorter sentences
&lt;/h2&gt;

&lt;p&gt;Answers did not shrink. The average one grew 6%, from 453 to 481 words, and Arena notes that no other Opus model gives longer answers. Sentences went the opposite way: 17% shorter on average, from 12.14 to 10.03 words.&lt;/p&gt;

&lt;p&gt;Vocabulary moved with them. The share of long content words fell from 41.7% to 38.6%, the lowest Arena found for any Claude model it analyzed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                        Opus 5    Opus 5.5
em dashes / 1k words     15.2       0.8
semicolons / 1k words    6.10       1.64
words per answer          453       481
words per sentence      12.14     10.03
long content words      41.7%     38.6%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Simpler, at the same volume
&lt;/h2&gt;

&lt;p&gt;Put together, the numbers describe a model that writes a little more and says it in plainer pieces. The cleaner text comes from how the sentences are built: fewer joints, fewer long words, a rhythm closer to how people write. Length played no part in it. Arena's own reading is that Opus 5.5 "should be easier to read."&lt;/p&gt;

&lt;h2&gt;
  
  
  One tell moved the other way
&lt;/h2&gt;

&lt;p&gt;Not every measure improved. Hedges such as "perhaps" and "arguably" rose 97%, from 0.39 to 0.77 per 1,000 words, the highest rate among the Claude models Arena analyzed. Its comment: "a new giveaway may be emerging."&lt;/p&gt;

&lt;p&gt;So the typical AI manner has not vanished. Part of it has moved from punctuation to word choice, where it is harder to spot at a glance. 👀&lt;/p&gt;

&lt;h2&gt;
  
  
  What the shift means
&lt;/h2&gt;

&lt;p&gt;Screening text for em dashes now waves most Opus 5.5 answers through, and the count worth watching has moved to words like "perhaps."&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://x.com/arena/status/2103532528946839901" rel="noopener noreferrer"&gt;Arena on X: how Opus 5.5 writes compared with Opus 5&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.bleepingcomputer.com/news/artificial-intelligence/claude-opus-55-uses-95-percent-fewer-em-dashes-but-its-answers-are-getting-longer/" rel="noopener noreferrer"&gt;BleepingComputer: Claude Opus 5.5 uses 95% fewer em dashes, but its answers are getting longer&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khasky&lt;/strong&gt; — &lt;a href="https://www.linkedin.com/in/khasky/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; / &lt;a href="https://github.com/khasky" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; / &lt;a href="https://www.patreon.com/khasky" rel="noopener noreferrer"&gt;Patreon&lt;/a&gt; / &lt;a href="https://bsky.app/profile/khasky.bsky.social" rel="noopener noreferrer"&gt;Bluesky&lt;/a&gt; / &lt;a href="https://mastodon.social/@khasky" rel="noopener noreferrer"&gt;Mastodon&lt;/a&gt; / &lt;a href="https://medium.com/@khasky" rel="noopener noreferrer"&gt;Medium&lt;/a&gt; / &lt;a href="https://dev.to/khasky"&gt;Devto&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khaskydev&lt;/strong&gt; — &lt;a href="https://x.com/khaskydev" rel="noopener noreferrer"&gt;X&lt;/a&gt; / &lt;a href="https://www.threads.com/@khaskydev" rel="noopener noreferrer"&gt;Threads&lt;/a&gt; / &lt;a href="https://www.instagram.com/khaskydev/" rel="noopener noreferrer"&gt;Instagram&lt;/a&gt; / &lt;a href="https://www.pinterest.com/khaskydev/" rel="noopener noreferrer"&gt;Pinterest&lt;/a&gt; / &lt;a href="https://www.tumblr.com/khaskydev" rel="noopener noreferrer"&gt;Tumblr&lt;/a&gt; / &lt;a href="https://www.facebook.com/khaskydev/" rel="noopener noreferrer"&gt;Facebook&lt;/a&gt; / &lt;a href="https://vk.ru/khaskydev" rel="noopener noreferrer"&gt;VK&lt;/a&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>anthropic</category>
      <category>llm</category>
      <category>aiwriting</category>
    </item>
    <item>
      <title>Awesome AGENTS.md for Any Agent and Any Stack</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Sun, 27 Sep 2026 05:02:26 +0000</pubDate>
      <link>https://dev.to/khasky/awesome-agentsmd-for-any-agent-and-any-stack-1d90</link>
      <guid>https://dev.to/khasky/awesome-agentsmd-for-any-agent-and-any-stack-1d90</guid>
      <description>&lt;p&gt;I reverse-engineered popular coding-agent instruction repos and distilled their best ideas into a stack-agnostic &lt;code&gt;AGENTS.md&lt;/code&gt;. 📦&lt;/p&gt;

&lt;p&gt;The core rules cover small, focused changes, security boundaries, debugging, and verification before an agent says a task is done. More specialized guidance lives in &lt;code&gt;rules/&lt;/code&gt; and is meant to be read when a task calls for it.&lt;/p&gt;

&lt;p&gt;Take it for a spin 👉 &lt;a href="https://github.com/khasky/awesome-agents-md" rel="noopener noreferrer"&gt;https://github.com/khasky/awesome-agents-md&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  One core for the work every agent does
&lt;/h2&gt;

&lt;p&gt;I wanted a set of rules I could use across Claude Code, Codex, Gemini, and Cursor without rewriting them for each framework or project. &lt;/p&gt;

&lt;p&gt;The core lives in one &lt;code&gt;AGENTS.md&lt;/code&gt; and covers the parts of a coding task that don't change with the stack: understanding the request, keeping the change focused, handling sensitive data, debugging failures, and checking the result.&lt;/p&gt;

&lt;p&gt;The file also sets boundaries. Unless I explicitly ask, the agent shouldn't delete files, rewrite Git history, force-push, or make a failing check pass by weakening it. It proposes a commit message instead of running &lt;code&gt;git commit&lt;/code&gt; or &lt;code&gt;git push&lt;/code&gt; on its own. These are instructions for the agent; actions that must be blocked without exception still need permissions, hooks, or CI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Less code, with evidence that it works
&lt;/h2&gt;

&lt;p&gt;Before adding a helper or dependency, the agent is asked to check whether the code already exists in the project, whether the standard library or an installed dependency can handle it, and whether the solution can be simpler. The goal is the smallest change that solves the actual problem. This part draws on the "lazy senior developer" approach from &lt;a href="https://github.com/DietrichGebert/ponytail" rel="noopener noreferrer"&gt;Ponytail&lt;/a&gt;; the repo also credits &lt;a href="https://github.com/JuliusBrussee/caveman" rel="noopener noreferrer"&gt;Caveman&lt;/a&gt; for its concise communication ideas.&lt;/p&gt;

&lt;p&gt;Verification is just as important. Before saying a task is done, fixed, or passing, the agent is told to run a relevant check, read its output, and report the result. "34/34 pass, exit 0" tells me something useful. "Should work now" tells me it's time to check. For a bug fix, the rules call for rerunning the scenario that failed in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Specialized rules when the task calls for them
&lt;/h2&gt;

&lt;p&gt;The core points to modules in &lt;code&gt;rules/&lt;/code&gt;, each with a "Read this when" trigger. There are modules for databases, backend security, caching, deployment, payments, browser automation, and more. An agent working on payments is instructed to read the payments rules; a task with no checkout doesn't need that guidance in context.&lt;/p&gt;

&lt;p&gt;The modules are framework-agnostic too. When an example uses a particular technology, its scope is stated explicitly. For example, the database guidance applies its general principles to relational databases while using PostgreSQL syntax for examples.&lt;/p&gt;

&lt;p&gt;I keep the core under 200 instruction lines and 32 KiB, and CI checks both limits. That forces a useful question whenever I add a rule: would an agent likely get this wrong without it? The core can stand on its own; the modules add depth only when needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use it with your coding agent
&lt;/h2&gt;

&lt;p&gt;Clone the repository first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/khasky/awesome-agents-md.git
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tools load shared instructions in different ways:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Agent&lt;/th&gt;
&lt;th&gt;How to use the core&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code&lt;/td&gt;
&lt;td&gt;Add an &lt;code&gt;@path/to/AGENTS.md&lt;/code&gt; line to your global &lt;code&gt;CLAUDE.md&lt;/code&gt;, or install the repository as a Claude Code plugin. Use one method, not both.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini CLI&lt;/td&gt;
&lt;td&gt;Add an &lt;code&gt;@path/to/AGENTS.md&lt;/code&gt; line to your global &lt;code&gt;GEMINI.md&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI Codex CLI&lt;/td&gt;
&lt;td&gt;Put a path reference in the global &lt;code&gt;AGENTS.md&lt;/code&gt; and ask the agent to read it, or copy the core into that file for direct loading. Codex does not have a native &lt;code&gt;@&lt;/code&gt; import.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cursor Agent&lt;/td&gt;
&lt;td&gt;Copy &lt;code&gt;AGENTS.md&lt;/code&gt; into a project root, or paste the core into Cursor's User Rules. Cursor does not have a global Markdown import.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;a href="https://github.com/khasky/awesome-agents-md#install" rel="noopener noreferrer"&gt;README&lt;/a&gt; has exact paths and verification steps for each tool. For Claude Code, the plugin option uses a &lt;code&gt;SessionStart&lt;/code&gt; hook to supply the core without editing &lt;code&gt;CLAUDE.md&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If your agent reads the local clone, updating the rules is a &lt;code&gt;git pull&lt;/code&gt; followed by a fresh session. If you copied or pasted the contents into an agent's settings, refresh that copy when you want the newer rules. There are no modes to switch between or extra dependencies to install for the ruleset.&lt;/p&gt;

&lt;p&gt;The first rule asks the agent to end its replies with &lt;code&gt;✓ awesome-agents-md&lt;/code&gt;. It's a quick way to see that the core reached the session. You can remove that line from your clone after checking it. The README also lists ways to inspect loaded instructions in each tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Refined on real projects
&lt;/h2&gt;

&lt;p&gt;I've revised these rules over many iterations while working on real applications. The public commit history shows how the wording has changed, and I keep revisiting rules as models improve. Every rule is plain Markdown, so you can inspect it and decide whether it belongs in your workflow.&lt;/p&gt;

&lt;p&gt;Try it, make it your own, and tell me what you'd improve. Issues and PRs are welcome! 😉&lt;/p&gt;




&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khasky&lt;/strong&gt; — &lt;a href="https://www.linkedin.com/in/khasky/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; / &lt;a href="https://github.com/khasky" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; / &lt;a href="https://www.patreon.com/khasky" rel="noopener noreferrer"&gt;Patreon&lt;/a&gt; / &lt;a href="https://bsky.app/profile/khasky.bsky.social" rel="noopener noreferrer"&gt;Bluesky&lt;/a&gt; / &lt;a href="https://mastodon.social/@khasky" rel="noopener noreferrer"&gt;Mastodon&lt;/a&gt; / &lt;a href="https://medium.com/@khasky" rel="noopener noreferrer"&gt;Medium&lt;/a&gt; / &lt;a href="https://dev.to/khasky"&gt;Devto&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khaskydev&lt;/strong&gt; — &lt;a href="https://x.com/khaskydev" rel="noopener noreferrer"&gt;X&lt;/a&gt; / &lt;a href="https://www.threads.com/@khaskydev" rel="noopener noreferrer"&gt;Threads&lt;/a&gt; / &lt;a href="https://www.instagram.com/khaskydev/" rel="noopener noreferrer"&gt;Instagram&lt;/a&gt; / &lt;a href="https://www.pinterest.com/khaskydev/" rel="noopener noreferrer"&gt;Pinterest&lt;/a&gt; / &lt;a href="https://www.tumblr.com/khaskydev" rel="noopener noreferrer"&gt;Tumblr&lt;/a&gt; / &lt;a href="https://www.facebook.com/khaskydev/" rel="noopener noreferrer"&gt;Facebook&lt;/a&gt; / &lt;a href="https://vk.ru/khaskydev" rel="noopener noreferrer"&gt;VK&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agentsmd</category>
      <category>claudecode</category>
      <category>aicoding</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Inside 80 Pull Request Templates From Popular Open Source Projects</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Wed, 23 Sep 2026 20:45:26 +0000</pubDate>
      <link>https://dev.to/khasky/inside-80-pull-request-templates-from-popular-open-source-projects-612</link>
      <guid>https://dev.to/khasky/inside-80-pull-request-templates-from-popular-open-source-projects-612</guid>
      <description>&lt;p&gt;A pull request template can ask for very little, and most of 80 well-known open source repositories do exactly that.&lt;/p&gt;

&lt;p&gt;I fetched every template straight from the GitHub, and the shape that turns up most often has exactly two prompts: an account of the change with its reason, and a description of how the author checked it.&lt;/p&gt;

&lt;p&gt;I read them the way a designer reads a form. Every field costs the person filling it in. Every field should also hand the reviewer something they could not get elsewhere, and judged that way, the boxes that restate CI fail first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Templates that render as nothing
&lt;/h2&gt;

&lt;p&gt;CPython, curl, Deno, etcd, Laravel, Node, Rust, Spring Boot, Swift, Vite and VS Code all belong here.&lt;/p&gt;

&lt;p&gt;Their sizes vary a lot. etcd needs 7 lines and VS Code 8. Laravel uses 9, Vite 21 and Swift 22, while Spring Boot reaches 29 and Node 45. Length makes no difference to what the reviewer sees, which is nothing.&lt;/p&gt;

&lt;p&gt;What they say while hidden is revealing. Rust's comment explains its own &lt;code&gt;homu-ignore:start&lt;/code&gt; and &lt;code&gt;homu-ignore:end&lt;/code&gt; markers, which tell the merge bot to strip the text. CPython's carries only title guidance, the &lt;code&gt;gh-NNNNNN:&lt;/code&gt; prefix followed by a summary. Node carries the Developer's Certificate of Origin 1.1 in full. Deno keeps a numbered list of rules. Swift describes the change it expects and how to trigger CI. curl carries nothing but an AI policy: "If you cannot understand or explain your work without using Artificial Intelligence (AI) then do not file here."&lt;/p&gt;

&lt;p&gt;So these files are notices, closer to a sign on a door than to a form. They most often cover which branch to target, where security issues go and what not to send. A reviewer never sees them, and neither does anyone reading the history years later.&lt;/p&gt;

&lt;p&gt;That design has one clear use. It reaches the author at the last moment before submission, when a wrong target branch is still cheap to fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  Projects with no template
&lt;/h2&gt;

&lt;p&gt;19 of the 80 repositories that answered ship no pull request template at all.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Vue core, webpack, Nuxt, React Router, Preact, Playwright,
Express, Fastify, TensorFlow, DuckDB, containerd, Mastodon,
OBS Studio, uBlock Origin, Kotlin, Ruby, PHP, LLVM, workerd
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are not the small ones, and they come from every domain the survey covered, from compilers and runtimes to databases, editors and web frameworks. Several maintain elaborate issue templates, so leaving the pull request box empty is a decision about where structure belongs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The common template is two prompts long
&lt;/h2&gt;

&lt;p&gt;The remaining templates, roughly fifty, put visible structure into a description. Their median is small, and the shape that turns up most is a pair of prompts: the change with its purpose, then how it was tested.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;### What does this PR do?

### How did you verify your code works?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is Bun's template, word for word, three lines in all. React, Solid, Tailwind, Neovim and Biome keep the same two ideas and add little. React's version has four visible lines and ends on the only stated consequence anywhere in the sample: "If you leave this empty, your PR will very likely be closed."&lt;/p&gt;

&lt;h2&gt;
  
  
  Long checklists come with heavy traffic
&lt;/h2&gt;

&lt;p&gt;Above the two-prompt floor, projects reach for very different tools.&lt;/p&gt;

&lt;p&gt;Angular splits the description into current behavior and new behavior, adds a breaking-change question and asks the author to pick one of nine PR types. Kubernetes runs to 92 lines with seven headings, from special notes for the reviewer to an AI usage disclosure. Grafana asks three plain questions: what the feature is, why it is needed and who it is for.&lt;/p&gt;

&lt;p&gt;Some projects drop prose entirely. Babel and Symfony both use a question-and-answer table, and Symfony's is read by a bot for labelling. Envoy uses colon-terminated fields, such as Commit Message, Risk Level and Release Notes. Go writes a plain-text instruction list with no headings at all.&lt;/p&gt;

&lt;p&gt;PyTorch goes another way and offers a directory of three choosable templates, for a docs typo, a fix for an issue and a preapproved change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Section frequency
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;what changed and why      nearly every visible template
linked issue              about two thirds
testing or verification   about half
a checklist               about a third
AI authorship disclosure  about a third, the newest
release note slot         ten or so
breaking change callout   under ten
risk and rollback         three
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What changed and why is nearly universal. Small templates phrase it as one question.&lt;/p&gt;

&lt;p&gt;The linked issue is the next most common ask. Most projects mention it in a comment. PyTorch, Keras, LangChain, Django, TypeScript and Axios treat it as mandatory, and the strict ones tend to attach a threat of closure, since a soft request gets ignored.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklists
&lt;/h2&gt;

&lt;p&gt;Checklists concentrate in high-volume consumer projects. The three heaviest checklist templates are Home Assistant at 119 lines, Transformers at 91 and Storybook at 86, and every one of them receives a stream of contributions from people who will submit once. The boxes screen that stream before a human looks. Home Assistant even has a box asserting the author reviewed two other open pull requests.&lt;/p&gt;

&lt;p&gt;Checklists travel well beyond consumer apps. Angular, Svelte, Django, Rails, Vitest and pandas all carry one. Among consumer apps, Signal-Android asks the author to list the devices and Android versions they tested on. Element asks for guidelines, tests, screenshots and the CLA, plus a pledge not to force push again. Immich and cal.com both pair their checklist with a testing section written in plain words.&lt;/p&gt;

&lt;p&gt;Where contributors are regulars, the checklist disappears. Neovim, curl, Git and Bun have none.&lt;/p&gt;

&lt;h2&gt;
  
  
  CI boxes
&lt;/h2&gt;

&lt;p&gt;A box asking whether lint passed or tests pass is a claim nobody verifies. The pipeline already knows the answer, and it gates the merge regardless. The tick has no effect.&lt;/p&gt;

&lt;p&gt;What earns space is the thing a pipeline can't see. Spark asks for copy-pasteable steps when a change was tested outside the regular unit tests. Storybook asks separately how the author tested and how a maintainer should reproduce it. Zed and Syncthing ask how a reviewer gets to the change at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Release notes
&lt;/h2&gt;

&lt;p&gt;Moby, Terraform and Envoy carry a fenced release-note block that a changelog tool scrapes, as do Kubernetes, Prometheus and Zed. Grafana goes further and uses the pull request title itself to auto-generate its changelog. Without that tooling, a release-note section is dead weight.&lt;/p&gt;

&lt;h2&gt;
  
  
  Risk and rollback
&lt;/h2&gt;

&lt;p&gt;Only three templates ask about risk. Terraform's asks which release a change targets, promises a revert within seven days if one is needed, and wants any change to security controls named. Envoy gives risk its own labelled line. The dotnet servicing template asks for customer impact, regression status and risk, because a merge there can go into a release customers already run. No library or frontend framework asks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Titles
&lt;/h2&gt;

&lt;p&gt;Where a title becomes a durable record, the template polices it. CPython wants the issue number up front. Go wants a package prefix like &lt;code&gt;net/http:&lt;/code&gt;, a short title and no Markdown. Deno asks for a conventional-commit prefix with worked examples. The title ends up as a commit subject, a backport tag or a changelog line.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI disclosure
&lt;/h2&gt;

&lt;p&gt;About a third of templates now ask about AI, and the wording differs every time. Django requires one of two exclusive boxes. pandas wants to know the exact assistant, its model version and its effort level. Caddy pre-fills "This PR is missing an assistance disclosure" under its heading, so silence ships as an admission. Transformers tells first-time contributors not to use code agents. Kubernetes writes a line addressed to the agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two missing prompts
&lt;/h2&gt;

&lt;p&gt;Why the defect wasn't caught earlier appears in no template. It's the one thing a reviewer cannot reconstruct from the diff, because the diff shows the fix and never the path by which the bug slipped past every earlier review, test and release.&lt;/p&gt;

&lt;p&gt;The second gap is smaller. What the change still doesn't prove appears only obliquely, in dotnet servicing through customer impact and in Envoy through its risk field. Both come close to asking where the author's confidence ends, and nothing else in the sample does even that much.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd put in a template
&lt;/h2&gt;

&lt;p&gt;For a small project with regular contributors, I'd write two prompts and stop:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the change and its reason, opening on what changed rather than on what broke&lt;/li&gt;
&lt;li&gt;how it was verified, phrased so that restating the pipeline obviously doesn't count&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A linked issue joins them only where work really lives in issues. Release notes wait for a generator. Risk questions come in the day a merge could hurt a production system or someone's data, and not a day before 🎯&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits of the evidence
&lt;/h2&gt;

&lt;p&gt;Everything here comes from one snapshot of default branches. Templates change fast. The AI sections alone are young enough that a year ago most of them didn't exist. I picked repositories by popularity, and popular projects run heavier templates than most.&lt;/p&gt;

&lt;p&gt;The survey records what templates ask and says nothing about whether contributors answer, which is the question that would settle which fields work.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GitHub Docs, creating a pull request template: &lt;a href="https://docs.github.com/en/communities/using-templates-to-encourage-useful-issues-and-pull-requests/creating-a-pull-request-template-for-your-repository" rel="noopener noreferrer"&gt;https://docs.github.com/en/communities/using-templates-to-encourage-useful-issues-and-pull-requests/creating-a-pull-request-template-for-your-repository&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Follow me for more on AI, LLMs, and Software Development: &lt;br&gt;
khasky — LinkedIn / Patreon / GitHub / Bluesky / Mastodon &lt;br&gt;
khaskydev — X / Threads / Instagram / Pinterest / Facebook &lt;/p&gt;

</description>
      <category>opensource</category>
      <category>github</category>
      <category>codereview</category>
      <category>pullrequests</category>
    </item>
    <item>
      <title>Vals AI Forecasts Full Recursive Self-Improvement by August 2027</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Tue, 22 Sep 2026 22:30:07 +0000</pubDate>
      <link>https://dev.to/khasky/vals-ai-forecasts-full-recursive-self-improvement-by-august-2027-320d</link>
      <guid>https://dev.to/khasky/vals-ai-forecasts-full-recursive-self-improvement-by-august-2027-320d</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2xit9e5dkz7upm6wekvp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2xit9e5dkz7upm6wekvp.png" alt="Vals AI Forecasts Full Recursive Self-Improvement by August 2027" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Vals AI measures how close the frontier models are to doing AI research alone, and it now puts that point at August 2027. The date comes from the company doing the measuring rather than from a lab with a release to promote, which is the reason to read it carefully.&lt;/p&gt;

&lt;p&gt;Rayan Krishnan, co-founder and CEO of Vals AI, told Bloomberg Tech that models will pass the researchers who build them by August 2027. Vals AI is an independent evaluator: it scores frontier models on real tasks and publishes the results, and one of those scores is aimed at exactly the capability he is forecasting.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is being forecast
&lt;/h2&gt;

&lt;p&gt;The claim is not about benchmark scores. Krishnan is forecasting full recursive self-improvement, RSI for short: a model that autonomously develops its own next version, with no researcher taking part in the decisions. 🔁&lt;/p&gt;

&lt;p&gt;The illustration he uses is plain. GPT-6 builds GPT-7 on its own, and Claude builds its own successor.&lt;/p&gt;

&lt;h2&gt;
  
  
  What full RSI takes away from people
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;today       AI takes part in building new models, people make the key decisions
full RSI    hypotheses -&amp;gt; experiments -&amp;gt; training the next model, all run by the model
example     GPT-6 builds GPT-7, Claude builds its own successor
forecast    August 2027, Vals AI's projection, not a confirmed date
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AI is already inside the process at every large lab. It writes and runs a good share of the work, and a person still makes the decisions that matter: which hypothesis gets tested, which experiment counts, which run becomes the successor. Full RSI is the name for the state where all three of those decisions belong to the model.&lt;/p&gt;

&lt;p&gt;That is why a higher score on a test does not qualify. A score measures output. The forecast is about who decides.&lt;/p&gt;

&lt;h2&gt;
  
  
  How firm the date is
&lt;/h2&gt;

&lt;p&gt;Not firm. The month is Vals AI's projection, made by Krishnan's team, and no lab has committed to it. The press version of the quote adds "or sooner", which hedges in the other direction and still leaves it a forecast.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Vals AI's own index puts things
&lt;/h2&gt;

&lt;p&gt;Vals AI keeps an RSI Index for this capability. Models run autonomous research tasks in language-model development, under fixed compute and time, and each task is scored on a log scale: 0 is the task's starting baseline, 0.5 a published, human or frontier-model reference, 1 the theoretical best.&lt;/p&gt;

&lt;p&gt;The index leader as it stands is Claude Fable 5.1 at 35.03%. Vals' reading of the results is that the models are strongest at running experiments and correcting misleading measurements, and still lack the judgment to identify higher-leverage directions. 🧪&lt;/p&gt;

&lt;p&gt;Read those two Vals statements together. The leaderboard says the missing part is judgment about direction. The forecast says the missing part shows up within about a year.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this is for
&lt;/h2&gt;

&lt;p&gt;Researchers, the people who plan headcount for research, and anyone who has filed recursive self-improvement under distant. The month is the least reliable fact here and the definition is the most reliable one.&lt;/p&gt;

&lt;p&gt;Does the gap between a 35.03% index score and a model choosing its own research direction close in a year, or is the month doing more work than the measurement supports?&lt;/p&gt;




&lt;p&gt;The interview: &lt;a href="https://www.youtube.com/watch?v=MpYAoufS588" rel="noopener noreferrer"&gt;https://www.youtube.com/watch?v=MpYAoufS588&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The RSI Index: &lt;a href="https://www.vals.ai/benchmarks/rsi_index" rel="noopener noreferrer"&gt;https://www.vals.ai/benchmarks/rsi_index&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khasky&lt;/strong&gt; — &lt;a href="https://www.linkedin.com/in/khasky/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; / &lt;a href="https://github.com/khasky" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; / &lt;a href="https://www.patreon.com/khasky" rel="noopener noreferrer"&gt;Patreon&lt;/a&gt; / &lt;a href="https://bsky.app/profile/khasky.bsky.social" rel="noopener noreferrer"&gt;Bluesky&lt;/a&gt; / &lt;a href="https://mastodon.social/@khasky" rel="noopener noreferrer"&gt;Mastodon&lt;/a&gt; / &lt;a href="https://medium.com/@khasky" rel="noopener noreferrer"&gt;Medium&lt;/a&gt; / &lt;a href="https://dev.to/khasky"&gt;Devto&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khaskydev&lt;/strong&gt; — &lt;a href="https://x.com/khaskydev" rel="noopener noreferrer"&gt;X&lt;/a&gt; / &lt;a href="https://www.threads.com/@khaskydev" rel="noopener noreferrer"&gt;Threads&lt;/a&gt; / &lt;a href="https://www.instagram.com/khaskydev/" rel="noopener noreferrer"&gt;Instagram&lt;/a&gt; / &lt;a href="https://www.pinterest.com/khaskydev/" rel="noopener noreferrer"&gt;Pinterest&lt;/a&gt; / &lt;a href="https://www.tumblr.com/khaskydev" rel="noopener noreferrer"&gt;Tumblr&lt;/a&gt; / &lt;a href="https://www.facebook.com/khaskydev/" rel="noopener noreferrer"&gt;Facebook&lt;/a&gt; / &lt;a href="https://vk.ru/khaskydev" rel="noopener noreferrer"&gt;VK&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agi</category>
      <category>airesearch</category>
    </item>
    <item>
      <title>Mozilla Measured the Open-Weight Gap: 5 Points, 8 of 10, 4% of Revenue</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Tue, 22 Sep 2026 07:09:11 +0000</pubDate>
      <link>https://dev.to/khasky/mozilla-measured-the-open-weight-gap-5-points-8-of-10-4-of-revenue-49lc</link>
      <guid>https://dev.to/khasky/mozilla-measured-the-open-weight-gap-5-points-8-of-10-4-of-revenue-49lc</guid>
      <description>&lt;p&gt;Four datasets went into Mozilla's State of Open Source AI report. A developer survey fielded with SlashData, OpenRouter's traffic panels, Epoch's Capabilities Index, and METR's time-horizon data. Read separately they support the usual arguments. Read together they describe open-weight models that are close on capability, ahead on routed traffic, and nearly absent from the revenue.&lt;/p&gt;

&lt;h2&gt;
  
  
  The measurements
&lt;/h2&gt;

&lt;p&gt;The three measurements the report rests on, with the window each one was taken in.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Epoch Capabilities Index, September 1 data
Claude Fable 5, Claude Opus 5 (closed)     162
Kimi K3 (open)                             157
gap                                        5 points
Epoch's average gap since January          8 points (90% CI 7-11)
in calendar time                           about 4 months (Epoch), about 4.4 months (Mozilla, from METR data)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OpenRouter, August 1-31, token volume
open weights in the top 10                 8 of 10
Chinese-built among those 8                7
first open model to lead weekly requests   DeepSeek, from August 3, after Google's 51 weeks at #1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OpenRouter, May-September 2025, model layer
usage      open ~20%    closed ~80%
revenue    open ~4%     closed ~96%
price      closed about 6x per call, at about 90% capability parity
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The traffic figures count routed OpenRouter requests only, first-party use inside ChatGPT, Gemini and Doubao sits outside them, and the revenue split is the last one published, from a window the report says usage has since moved past.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a 5-point gap and a 4-month gap are the same fact
&lt;/h2&gt;

&lt;p&gt;The 5 points on the September chart run under Epoch's own average of 8 since January, and that average is what Epoch converts into roughly four months of calendar lead.&lt;/p&gt;

&lt;p&gt;Mozilla ran the conversion a second way, fitting METR's raw time-horizon data, and got about 4.4 months, with open capability doubling every 3.9 months against 5.5 for closed. The report labels its own fit a simplified reproduction rather than a result. 📐&lt;/p&gt;

&lt;p&gt;Release timing keeps the number stable. Every model on the chart shipped between June and August, so the margin is redrawn each cycle instead of compounding across them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Usage on OpenRouter
&lt;/h2&gt;

&lt;p&gt;DeepSeek's V4 Flash 0731 led the August token table at 45.1 trillion tokens, DeepSeek held three of the top ten places, and only three US entries made the list at all. On requests rather than tokens, August 3 was the first day an open model took the top spot, and DeepSeek covered the distance from third to first in eight weeks.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the numbers do not cover
&lt;/h2&gt;

&lt;p&gt;Taking the limits before the economics is deliberate, because the revenue figure below is the one most likely to be quoted without them.&lt;/p&gt;

&lt;p&gt;Routed traffic is not all traffic. Anything served first-party inside ChatGPT, Gemini or Doubao never appears in OpenRouter's panels, which leaves a large share of real usage outside every figure here.&lt;/p&gt;

&lt;p&gt;The revenue split is older still. It covers May to September 2025 and has not been re-measured, and the report states that directly: usage has moved since that window, and the 4% is simply the last figure published.&lt;/p&gt;

&lt;p&gt;The measured gap is a floor. Epoch gives two reasons it may be understated: open models optimize on benchmarks more aggressively, and unpublished closed models are absent from the baseline.&lt;/p&gt;

&lt;h2&gt;
  
  
  The economics
&lt;/h2&gt;

&lt;p&gt;The report's own reading of the revenue table is one clause long: capability is near parity, and price drives the split. Among developers who choose open models, 30% name lower cost as a top reason and 28% name privacy, and the Linux Foundation's estimate of unrealized annual savings from that asymmetry is $24.8B.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where open stalls
&lt;/h2&gt;

&lt;p&gt;79% of surveyed developers use open models, yet 51% of open deployments reach production against 63% for closed, a gap the survey traces to tooling and trust rather than capability.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;firms using open components            89%
developers using open models           79%
open models reaching production        51%
closed models reaching production      63%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One figure underneath it carries the section. Vendor-partnered deployments reach production 67% of the time and internal builds 33%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Capability without pretraining
&lt;/h2&gt;

&lt;p&gt;DeepSeek's July 31 post-training pass lifted V4-Flash 10 points on the Artificial Analysis index on an unchanged architecture, and the August 13 pass added 8 to V4-Pro, with no new pretraining run behind either.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;V4-Flash-0731   Jul 31   AA index ~40 -&amp;gt; ~50   +10   post-training pass, same architecture
V4-Pro-0813     Aug 13   AA index 45 -&amp;gt; 53     +8    post-training pass, two weeks later
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The V4-Flash pass moved Terminal-Bench 2.1 from 61.8 to 82.7 and cut eval-suite output tokens from 234M to 206M, so the score did not come from longer answers. Post-training costs a fraction of pretraining, which puts this kind of gain inside reach of anyone with an open base and a budget. 🔁&lt;/p&gt;

&lt;p&gt;Any targeted use of another model's outputs would also happen in post-training, so the report treats the cheap-gains question and the provenance question as one.&lt;/p&gt;




&lt;ul&gt;
&lt;li&gt;Report page: &lt;a href="https://stateofopensource.ai" rel="noopener noreferrer"&gt;https://stateofopensource.ai&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Full PDF, v1.1: &lt;a href="https://stateofopensource.ai/state-of-open-source-ai-v1-1.pdf" rel="noopener noreferrer"&gt;https://stateofopensource.ai/state-of-open-source-ai-v1-1.pdf&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khasky&lt;/strong&gt; — &lt;a href="https://www.linkedin.com/in/khasky/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; / &lt;a href="https://github.com/khasky" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; / &lt;a href="https://www.patreon.com/khasky" rel="noopener noreferrer"&gt;Patreon&lt;/a&gt; / &lt;a href="https://bsky.app/profile/khasky.bsky.social" rel="noopener noreferrer"&gt;Bluesky&lt;/a&gt; / &lt;a href="https://mastodon.social/@khasky" rel="noopener noreferrer"&gt;Mastodon&lt;/a&gt; / &lt;a href="https://medium.com/@khasky" rel="noopener noreferrer"&gt;Medium&lt;/a&gt; / &lt;a href="https://dev.to/khasky"&gt;Devto&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khaskydev&lt;/strong&gt; — &lt;a href="https://x.com/khaskydev" rel="noopener noreferrer"&gt;X&lt;/a&gt; / &lt;a href="https://www.threads.com/@khaskydev" rel="noopener noreferrer"&gt;Threads&lt;/a&gt; / &lt;a href="https://www.instagram.com/khaskydev/" rel="noopener noreferrer"&gt;Instagram&lt;/a&gt; / &lt;a href="https://www.pinterest.com/khaskydev/" rel="noopener noreferrer"&gt;Pinterest&lt;/a&gt; / &lt;a href="https://www.tumblr.com/khaskydev" rel="noopener noreferrer"&gt;Tumblr&lt;/a&gt; / &lt;a href="https://www.facebook.com/khaskydev/" rel="noopener noreferrer"&gt;Facebook&lt;/a&gt; / &lt;a href="https://vk.ru/khaskydev" rel="noopener noreferrer"&gt;VK&lt;/a&gt;&lt;/p&gt;

</description>
      <category>opensourceai</category>
      <category>openweights</category>
      <category>deepseek</category>
      <category>mozilla</category>
    </item>
    <item>
      <title>What Markdown Bold Actually Does in an Agent Instruction File</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Mon, 21 Sep 2026 22:23:45 +0000</pubDate>
      <link>https://dev.to/khasky/what-markdown-bold-actually-does-in-an-agent-instruction-file-2c80</link>
      <guid>https://dev.to/khasky/what-markdown-bold-actually-does-in-an-agent-instruction-file-2c80</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fos0uzf72h2bu1qb1j8g8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fos0uzf72h2bu1qb1j8g8.png" alt="Rows of an instruction file where nearly every line is wrapped in green asterisks, so the emphasis no longer marks anything." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A rule gets skipped, so you make it bold. It happens again with a different rule, so that one goes bold as well, then the warnings, then the thing that broke production once.&lt;/p&gt;

&lt;p&gt;I did that to my own files for months without ever checking whether the highlighting was doing anything. 😅&lt;/p&gt;

&lt;h2&gt;
  
  
  There is no bold channel
&lt;/h2&gt;

&lt;p&gt;The asterisks are tokens. A model reading your &lt;code&gt;CLAUDE.md&lt;/code&gt;, &lt;code&gt;AGENTS.md&lt;/code&gt;, &lt;code&gt;GEMINI.md&lt;/code&gt; or &lt;code&gt;SKILL.md&lt;/code&gt; receives them as characters in a sequence, exactly like every other character in the file, and no documented path runs from "this span is emphasized" to "weight this more heavily".&lt;/p&gt;

&lt;p&gt;The intuition comes from somewhere real. The habit is not stupid. A reader's eye lands on bold before they have chosen to read the line, and since the model reads the same file we do, it feels like it should inherit the reflex.&lt;/p&gt;

&lt;p&gt;It does not inherit it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The method that exists because emphasis does not carry
&lt;/h2&gt;

&lt;p&gt;PASTA, from Georgia Tech, UC Berkeley and Microsoft Research, reweights a small subset of attention heads at inference so a model attends to a span the user designates, changes no parameters, and reports a 22% average accuracy improvement for LLAMA-7B.&lt;/p&gt;

&lt;p&gt;Its abstract opens on the analogy itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;In human-written articles, we often leverage the subtleties of
text style, such as bold and italics, to guide the attention of
readers. ... Existing methods, however, are constrained to
process plain text and do not support such a mechanism.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What the three ecosystems document
&lt;/h2&gt;

&lt;p&gt;Anthropic's Claude Code guidance is the most directly useful thing I found:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;If Claude keeps skipping one instruction, add emphasis such as
"IMPORTANT" to that line alone. If you emphasize many lines,
none of them stands out.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things sit in that sentence. The recommended instrument is a word rather than markup, and the failure mode is named outright by the vendor. 📄&lt;/p&gt;

&lt;p&gt;The skills documentation supplies the constraint underneath. A loaded &lt;code&gt;SKILL.md&lt;/code&gt; enters the conversation as one message and stays there across later turns, which makes every line a recurring cost, and the stated cap is 500 lines.&lt;/p&gt;

&lt;p&gt;Outside Anthropic, the silence says the same thing.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Claude Code    CLAUDE.md, SKILL.md    one word, IMPORTANT, on one line
Codex          AGENTS.md              standard Markdown, no special syntax
Gemini     GEMINI.md              concatenated and sent with every prompt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three ecosystems, and not one of them documents emphasis as a mechanism.&lt;/p&gt;




&lt;h2&gt;
  
  
  Format does matter, at a different scale
&lt;/h2&gt;

&lt;p&gt;The opposite overcorrection is also wrong, because prompt formatting is not inert. One study rendered identical content as plain text, Markdown, JSON and YAML and measured all four: GPT-3.5-turbo moved by up to 40% on a code translation task, and GPT-4 held much steadier on the same swap.&lt;/p&gt;

&lt;p&gt;Look at what varied. Whole schemes. An inline marker is a far smaller perturbation of the same input, and the sensitivity shrank as the model got stronger.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I counted my own file
&lt;/h2&gt;

&lt;p&gt;One skill of mine, four files, 253,612 characters:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SKILL.md                  91 bold spans
platform-posting.md      386
browser-interaction.md   116
post-formatting.md        42
                       -----
                         635
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Deleting every asterisk in all four saves 2,540 characters, which is about 1% of the text and a few hundred tokens across the whole skill. So the argument I assumed I would make, the one about token cost, was dead before I started writing it. What is measurable is the density: 91 spans across 264 lines is one every three lines, and at that rate the marker distinguishes nothing. Nobody has benchmarked bold against no bold on an instruction file, so there is no measured penalty to point at, and I am not going to invent one. Bold does no harm, and it does no steering either.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the count does not license
&lt;/h2&gt;

&lt;p&gt;Do not expect stripping the asterisks to change what the model does. The reason to cut them is that a marker on every third line stops distinguishing anything, for the model reading the file and for whoever has to maintain it.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the rule keeps getting skipped anyway
&lt;/h2&gt;

&lt;p&gt;The symptom is a rule the model keeps skipping however loudly it is marked. The vendor's own diagnosis is that the file is too long and the rule is getting lost in it, and the fix on the page is to prune rather than to emphasize.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;For each line, ask: "Would removing this cause Claude to
make mistakes?" If not, cut it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;A rule that survives that question has earned its line, and a rule that does not was never going to be rescued by asterisks.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually moves adherence
&lt;/h2&gt;

&lt;p&gt;Position comes first, so the rule sits at the step where it applies rather than in a preamble, and one or two hard words per file, NEVER or MUST, stay rare enough to register when they appear. For anything that has to hold every time, I reach for a hook, a permission rule or a CI check, because prose asks and a hook decides.&lt;/p&gt;

&lt;p&gt;Has anyone measured whether removing the bold from an instruction file changes what their agent does?&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code best practices: &lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;https://code.claude.com/docs/en/best-practices&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Claude Code skills documentation: &lt;a href="https://code.claude.com/docs/en/skills" rel="noopener noreferrer"&gt;https://code.claude.com/docs/en/skills&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;PASTA: &lt;a href="https://arxiv.org/abs/2311.02262" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2311.02262&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Does Prompt Formatting Have Any Impact on LLM Performance?: &lt;a href="https://arxiv.org/abs/2411.10541" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2411.10541&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khasky&lt;/strong&gt; — &lt;a href="https://www.linkedin.com/in/khasky/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; / &lt;a href="https://github.com/khasky" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; / &lt;a href="https://www.patreon.com/khasky" rel="noopener noreferrer"&gt;Patreon&lt;/a&gt; / &lt;a href="https://bsky.app/profile/khasky.bsky.social" rel="noopener noreferrer"&gt;Bluesky&lt;/a&gt; / &lt;a href="https://mastodon.social/@khasky" rel="noopener noreferrer"&gt;Mastodon&lt;/a&gt; / &lt;a href="https://medium.com/@khasky" rel="noopener noreferrer"&gt;Medium&lt;/a&gt; / &lt;a href="https://dev.to/khasky"&gt;Devto&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khaskydev&lt;/strong&gt; — &lt;a href="https://x.com/khaskydev" rel="noopener noreferrer"&gt;X&lt;/a&gt; / &lt;a href="https://www.threads.com/@khaskydev" rel="noopener noreferrer"&gt;Threads&lt;/a&gt; / &lt;a href="https://www.instagram.com/khaskydev/" rel="noopener noreferrer"&gt;Instagram&lt;/a&gt; / &lt;a href="https://www.pinterest.com/khaskydev/" rel="noopener noreferrer"&gt;Pinterest&lt;/a&gt; / &lt;a href="https://www.tumblr.com/khaskydev" rel="noopener noreferrer"&gt;Tumblr&lt;/a&gt; / &lt;a href="https://www.facebook.com/khaskydev/" rel="noopener noreferrer"&gt;Facebook&lt;/a&gt; / &lt;a href="https://vk.ru/khaskydev" rel="noopener noreferrer"&gt;VK&lt;/a&gt;&lt;/p&gt;

</description>
      <category>promptengineering</category>
      <category>claudecode</category>
      <category>contextengineering</category>
      <category>aiagents</category>
    </item>
    <item>
      <title>Run DeepSeek Inside Claude Code or Codex Without Replacing Your Workflow</title>
      <dc:creator>Khasky</dc:creator>
      <pubDate>Sat, 19 Sep 2026 05:42:01 +0000</pubDate>
      <link>https://dev.to/khasky/run-deepseek-inside-claude-code-or-codex-cli-without-replacing-your-workflow-3bj0</link>
      <guid>https://dev.to/khasky/run-deepseek-inside-claude-code-or-codex-cli-without-replacing-your-workflow-3bj0</guid>
      <description>&lt;p&gt;Most people assume that trying DeepSeek for coding means installing another client. It does not: Claude Code and Codex both have a documented path to it, and the only thing that changes is which model answers. 🙂&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Claude Code   -&amp;gt; direct, Anthropic-compatible endpoint
Codex CLI     -&amp;gt; official setup script, Responses API
OpenCode      -&amp;gt; /connect, then /models
Aider         -&amp;gt; OpenAI-compatible base URL
LiteLLM       -&amp;gt; one local endpoint, fallback and budget
Ollama        -&amp;gt; local models, no API at all
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Claude Code with DeepSeek behind it
&lt;/h2&gt;

&lt;p&gt;For Claude Code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.deepseek.com/anthropic"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_AUTH_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-YOUR_DEEPSEEK_KEY"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"deepseek-flash[1m]"&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="n"&gt;claude&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The published block goes further and maps the default model slots, the subagent model and the effort level to DeepSeek. Requests billed this way come off the DeepSeek API balance, and a Claude subscription is not touched by them. The variables belong to the terminal window you set them in, so going back to Anthropic is a new window rather than an undo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Codex over the Responses API
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;irm&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;https://cdn.deepseek.com/api-docs/codex-deepseek-setup-en.ps1&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="n"&gt;codex&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The script offers Flash or Pro, writes the model catalog Codex reads, sets the provider to DeepSeek and keeps a backup of the configuration it replaced. Downloading it and reading it first is the safer order, because it rewrites files you rely on. 🔧&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenCode
&lt;/h2&gt;

&lt;p&gt;OpenCode connects to DeepSeek from inside its own interface: &lt;code&gt;/connect&lt;/code&gt; takes the key, &lt;code&gt;/models&lt;/code&gt; picks Flash or Pro. The client itself costs nothing, so the API usage is the whole bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  Aider
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;setx&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;OPENAI_API_BASE&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.deepseek.com"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;setx&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sk-YOUR_DEEPSEEK_KEY"&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="n"&gt;aider&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;--model&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;openai/deepseek-flash&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Aider works on the Git repository directly, and &lt;code&gt;setx&lt;/code&gt; needs a terminal restart before the values are visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  LiteLLM as one local endpoint
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;model_list&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deepseek-flash&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deepseek/deepseek-flash&lt;/span&gt;
      &lt;span class="na"&gt;api_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;os.environ/DEEPSEEK_API_KEY&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Point Claude Code at the local endpoint with the proxy's master key and you get provider fallback, spend tracking and a budget in front of DeepSeek, Claude, GPT and Gemini at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running locally with Ollama
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;ollama&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;run&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;deepseek-coder:6.7b&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On an 11 GB card, DeepSeek Coder 6.7B and the 7B and 8B R1 distills are realistic, and 14B runs in 4-bit with part of the load in RAM. A 671B model there is not worth attempting. 🐢&lt;/p&gt;

&lt;h2&gt;
  
  
  Why DeepSeek Flash instead of Claude Sonnet 5 or GPT-5.6 Terra?
&lt;/h2&gt;

&lt;p&gt;DeepSeek publishes two rates per million tokens, and off-peak is half of peak:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Flash, off-peak   $0.15 input (cache miss)   $0.60 output
Flash, peak       $0.30 input (cache miss)   $1.20 output
Pro, off-peak     $0.66 input (cache miss)   $1.98 output
Pro, peak         $1.32 input (cache miss)   $3.96 output
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The standard rung at the two providers a coding CLI reaches for by default:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Claude Sonnet 5   $2 input    $10 output
GPT-5.6 Terra     $2 input    $12 output
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Flash vs Claude Sonnet 5   ~7-17x cheaper
Flash vs GPT-5.6 Terra     ~7-20x cheaper
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The low end of each range is peak input, the high end is off-peak output, so where a run lands depends on the hour it runs and on how much of it is output. This compares API rates against API rates, so a fixed monthly subscription is not directly comparable.&lt;/p&gt;

&lt;p&gt;Flash is a strong default for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;repository exploration&lt;/li&gt;
&lt;li&gt;tests&lt;/li&gt;
&lt;li&gt;docs&lt;/li&gt;
&lt;li&gt;boilerplate&lt;/li&gt;
&lt;li&gt;routine refactors&lt;/li&gt;
&lt;li&gt;CI fixes&lt;/li&gt;
&lt;li&gt;subagent loops&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I would still escalate to a frontier Claude or GPT model for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ambiguous architecture&lt;/li&gt;
&lt;li&gt;hard debugging&lt;/li&gt;
&lt;li&gt;security review&lt;/li&gt;
&lt;li&gt;risky migrations&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Neither side is universally weaker, and the split is where each one is worth its price.&lt;/p&gt;

&lt;h2&gt;
  
  
  Privacy
&lt;/h2&gt;

&lt;p&gt;Do not casually send API keys, .env files, production credentials or large database dumps to a cloud model. For confidential code, prefer local inference or a provider whose retention policy matches your requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Codex keeps calling the OpenAI endpoint
&lt;/h2&gt;

&lt;p&gt;The interface says DeepSeek-Flash and the traffic still goes to the OpenAI Responses endpoint, with no OpenAI key in play. The model catalog changed and the provider section did not.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="nn"&gt;[model_providers.deepseek]&lt;/span&gt;
&lt;span class="py"&gt;name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"deepseek"&lt;/span&gt;
&lt;span class="py"&gt;base_url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"https://api.deepseek.com/"&lt;/span&gt;
&lt;span class="py"&gt;wire_api&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"responses"&lt;/span&gt;
&lt;span class="py"&gt;experimental_bearer_token&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"sk-YOUR_DEEPSEEK_API_KEY"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Close Codex, ChatGPT Desktop and VS Code completely and open them again, because picking the model in a running window does not reload the provider.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The catalog and the provider are two different settings, and only one of them is what the traffic follows.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Reference links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code integration: &lt;a href="https://api-docs.deepseek.com/quick_start/agent_integrations/claude_code/" rel="noopener noreferrer"&gt;https://api-docs.deepseek.com/quick_start/agent_integrations/claude_code/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Codex CLI integration: &lt;a href="https://api-docs.deepseek.com/quick_start/agent_integrations/codex/" rel="noopener noreferrer"&gt;https://api-docs.deepseek.com/quick_start/agent_integrations/codex/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;DeepSeek pricing: &lt;a href="https://api-docs.deepseek.com/quick_start/pricing/" rel="noopener noreferrer"&gt;https://api-docs.deepseek.com/quick_start/pricing/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Claude pricing: &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;https://platform.claude.com/docs/en/about-claude/pricing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI pricing: &lt;a href="https://developers.openai.com/api/docs/pricing" rel="noopener noreferrer"&gt;https://developers.openai.com/api/docs/pricing&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Follow me for more on AI and Software Development:&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khasky&lt;/strong&gt; — &lt;a href="https://www.linkedin.com/in/khasky/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; / &lt;a href="https://github.com/khasky" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; / &lt;a href="https://www.patreon.com/khasky" rel="noopener noreferrer"&gt;Patreon&lt;/a&gt; / &lt;a href="https://bsky.app/profile/khasky.bsky.social" rel="noopener noreferrer"&gt;Bluesky&lt;/a&gt; / &lt;a href="https://mastodon.social/@khasky" rel="noopener noreferrer"&gt;Mastodon&lt;/a&gt; / &lt;a href="https://medium.com/@khasky" rel="noopener noreferrer"&gt;Medium&lt;/a&gt; / &lt;a href="https://dev.to/khasky"&gt;Devto&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;khaskydev&lt;/strong&gt; — &lt;a href="https://x.com/khaskydev" rel="noopener noreferrer"&gt;X&lt;/a&gt; / &lt;a href="https://www.threads.com/@khaskydev" rel="noopener noreferrer"&gt;Threads&lt;/a&gt; / &lt;a href="https://www.instagram.com/khaskydev/" rel="noopener noreferrer"&gt;Instagram&lt;/a&gt; / &lt;a href="https://www.pinterest.com/khaskydev/" rel="noopener noreferrer"&gt;Pinterest&lt;/a&gt; / &lt;a href="https://www.tumblr.com/khaskydev" rel="noopener noreferrer"&gt;Tumblr&lt;/a&gt; / &lt;a href="https://www.facebook.com/khaskydev/" rel="noopener noreferrer"&gt;Facebook&lt;/a&gt; / &lt;a href="https://vk.ru/khaskydev" rel="noopener noreferrer"&gt;VK&lt;/a&gt;&lt;/p&gt;

</description>
      <category>deepseek</category>
      <category>claudecode</category>
      <category>codex</category>
      <category>vibecoding</category>
    </item>
  </channel>
</rss>
