<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ninghonggang</title>
    <description>The latest articles on DEV Community by ninghonggang (@ninghonggang).</description>
    <link>https://dev.to/ninghonggang</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3926050%2Fdf9db481-5c4d-4e66-b1b4-5b50acad6ad2.png</url>
      <title>DEV Community: ninghonggang</title>
      <link>https://dev.to/ninghonggang</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ninghonggang"/>
    <language>en</language>
    <item>
      <title>Same AI Coding Tools, Two Languages For Buying Them a7c8d3</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Tue, 28 Jul 2026 11:09:46 +0000</pubDate>
      <link>https://dev.to/ninghonggang/same-ai-coding-tools-two-languages-for-buying-them-a7c8d3-18in</link>
      <guid>https://dev.to/ninghonggang/same-ai-coding-tools-two-languages-for-buying-them-a7c8d3-18in</guid>
      <description>&lt;p&gt;Reading the Juejin search results for 热门 AI this morning, I kept tripping over the same AI coding tools but rating them with two completely different vocabularies. One review crowns winners in tier letters (S, A, B, D). Another review, covering almost the same tools, crowns winners in procurement language (等保三级认证, 国密SM4加密, 中文代码生成接受率87%). The names on the shortlists overlap. The vocabulary does not. That feels like the load-bearing thing right now.&lt;/p&gt;

&lt;p&gt;The first kind of post is easy to find if you have been scrolling Juejin this year. One review stacks Cursor, Claude Code, Codex, Lovable, v0, Replit, Bolt, and Windsurf into an S/A/B/D ladder. Cursor and Claude Code sit at the top of the S column. Codex is in A with the qualification "起步较晚，功能尚不完善". Lovable lands at B with "上手快，适合不懂技术的用户快速构建应用". Bolt sits at B with weaker community support. Windsurf drops all the way to D with the damning line "创始人跑路". The whole reading experience feels like an English-language tier list that someone translated into Chinese and ran through a Juejin template.&lt;/p&gt;

&lt;p&gt;The second kind of post sounds nothing like the first. A recent seven-tool Chinese review rates everything on what I think of as enterprise procurement axes. CodeBuddy gets a 9.7/10 because it cites "等保三级认证，国密SM4加密", an 8x performance lead over unnamed competitors, "中文代码生成接受率87%", and "Figma转换准确率99.9%". FlowDesign gets 7.3/10 with the qualifier "功能相对单一，不适合全栈开发". CodePulse gets 7.4/10 with "补全准确率高达92%". The whole reading experience feels like a tender document.&lt;/p&gt;

&lt;p&gt;So here is the meta-pattern I want to pin down. A buyer reading the tier-letter review walks away thinking Cursor, Claude Code, and Codex are the obvious buys, with Windsurf disqualified on founder-trust grounds and Lovable good enough for non-engineers. A procurement officer reading the procurement-language review walks away thinking CodeBuddy is the obvious buy, with FlowDesign and CodePulse acceptable as supplementary plugins, and probably Cursor nowhere on the radar unless the same post also touches 海外工具 (which this one largely does not). Same family of products. Same platform. Two completely different verdicts, and the disparity is not about the tools at all, it is about which axis the writer cared about.&lt;/p&gt;

&lt;p&gt;To be fair, both axes are real. Tier letters compress a reader's prior impressions into a recognizable shape (S means "you can probably bet the project on this"), which is useful when a buyer already knows the genre. Procurement axes answer a narrower question: will this pass our internal review by IT and legal? They are not interchangeable. They might even be parallel: the tier letter says "how good is the model", the procurement axis says "can we buy this from where we are". I am not sure either axis is wrong. I am pretty sure neither axis alone is sufficient.&lt;/p&gt;

&lt;p&gt;What I am less sure about is which axis a buyer should weight first in 2026. My instinct, after a few years of shipping software that has to clear both my team's review and the procurement gate, is to read the procurement axis second, not first. The tier-letter post tells you what the AI actually does well. The procurement post tells you whether you can run it under your specific constraints. If you reverse that order and start with procurement, you can end up buying something that nobody on the team can use productively because it does not play nicely with the language or the security stack your company already standardized on. On the other hand, if you start with tier letters only, you can ship a prototype that your legal team will not let you put in production.&lt;/p&gt;

&lt;p&gt;My read, with the caveat that I work mostly on internal tooling and may be biased toward the security side, is that the Juejin roundups have leaned heavily procurement-side in the last three or four months. The "tier letter" framing is still around but it is getting rarer. Possibly because Chinese enterprise buyers reward specific verifiable compliance terms more than a single-letter grade, or possibly because the writer economy over there has a stronger procurement-research-to-writing pipeline than the Western tier-list economy, where GitHub Copilot has historically been the easy default and CodeBuddy-style procurement angles are rarer than on the Juejin timeline. I do not actually know which. What I do know is that a buyer who only reads the tier-letter roundups may be missing out on tools their compliance team would prefer, and a buyer who only reads the procurement roundups may be over-weighting tools that happen to ship in the right regulatory vocabulary rather than the ones that solve their actual problem.&lt;/p&gt;

&lt;p&gt;A practical rule of thumb, before you sign anything: read at least one roundup from each camp. The procurement camp is probably the easiest to spot on Juejin right now, just look for the presence of 等保, 国密, or 准确率 numbers in the tool descriptions. The tier-letter camp is probably fading but still alive in longer Chinese re-postings of English-language reviews. If you read one of each and they recommend different tools, that is a signal to dig into the third question, which neither format tends to answer well: how does this tool actually behave on a real repo your team owns, not on the demo repo the seller prepared. That third question is probably the one that matters most, and the Juejin search results I have been scrolling this morning mostly do not address it. Probably worth a separate search tomorrow.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>The AI Coding Tool Winner Matters Less Than the Handoff It Leaves</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Mon, 27 Jul 2026 11:21:31 +0000</pubDate>
      <link>https://dev.to/ninghonggang/the-ai-coding-tool-winner-matters-less-than-the-handoff-it-leaves-3pm8</link>
      <guid>https://dev.to/ninghonggang/the-ai-coding-tool-winner-matters-less-than-the-handoff-it-leaves-3pm8</guid>
      <description>&lt;p&gt;I went down a rabbit hole this morning reading the Juejin search results for 热门 AI 2025, and one thing finally clicked: the most useful part of the popular AI coding-tool posts is not the winner. It is the handoff each tool leaves behind.&lt;/p&gt;

&lt;p&gt;The roundup that kept showing up puts Cursor, Claude Code, Codex, v0, Lovable, Replit, and Windsurf into a single S/A/B/D-style comparison. Another popular post puts Cursor first for overall ability, GitHub Copilot first for ecosystem integration, Codeium first for value, and v0 first for frontend UI. A separate long-form experience piece uses Cursor for work, Trae and Codex for learning, and VS Code with GitHub Copilot as a backup. Same broad question, very different scorecards.&lt;/p&gt;

&lt;p&gt;I am less interested in arguing over whether Cursor beats Claude Code than I used to be. After a few years of shipping software, the question I actually need answered is: what can I hand to the next person, or to my future self, when the AI session ends?&lt;/p&gt;

&lt;p&gt;Cursor’s official Plan Mode is a good example. It researches the repository, finds relevant files, asks clarifying questions, and writes a plan with file paths and code references before building. That means the durable handoff can be a Markdown plan plus a reviewable diff. Claude Code comes from a different direction: its documentation describes a terminal workflow where it explores a project, edits multiple files, runs tests, and works with Git. Its handoff is usually a diff, command output, and a commit candidate. v0 is different again. Vercel’s docs describe high-fidelity UI generation, backend connections, diagnostics, and deployment, so its natural handoff is a component or a running web prototype.&lt;/p&gt;

&lt;p&gt;Those are not small variations of the same answer. They are different engineering artifacts with different lifetimes and different ways to check them.&lt;/p&gt;

&lt;p&gt;A chat transcript is easy to read and hard to maintain. A screenshot is easy to show and hard to merge. A generated component can be useful, but only after someone checks its data loading, accessibility, and boundary with the existing app. A Git diff can be reviewed, tested, reverted, and assigned to a commit. A plan can prevent an agent from touching the wrong service before any code is changed.&lt;/p&gt;

&lt;p&gt;That is why I now keep a small artifact table before choosing a tool:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Tool I would start with&lt;/th&gt;
&lt;th&gt;Handoff I want&lt;/th&gt;
&lt;th&gt;First check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Understand an unfamiliar repository&lt;/td&gt;
&lt;td&gt;Claude Code or Cursor&lt;/td&gt;
&lt;td&gt;Notes, file map, and a small plan&lt;/td&gt;
&lt;td&gt;Does every claim point to a file?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Implement a bounded backend change&lt;/td&gt;
&lt;td&gt;Claude Code or Codex&lt;/td&gt;
&lt;td&gt;Small diff and passing tests&lt;/td&gt;
&lt;td&gt;Run the narrow test first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Iterate on a frontend idea&lt;/td&gt;
&lt;td&gt;v0 or Cursor&lt;/td&gt;
&lt;td&gt;React component in the real repo&lt;/td&gt;
&lt;td&gt;Check data states and keyboard flow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Make a tiny hosted prototype&lt;/td&gt;
&lt;td&gt;v0 or Replit&lt;/td&gt;
&lt;td&gt;A URL plus source code&lt;/td&gt;
&lt;td&gt;Can another developer run it?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Add inline code while staying in VS Code&lt;/td&gt;
&lt;td&gt;GitHub Copilot&lt;/td&gt;
&lt;td&gt;A local edit in the existing review loop&lt;/td&gt;
&lt;td&gt;Inspect the diff immediately&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The exact tool names may change. The artifact questions age more slowly.&lt;/p&gt;

&lt;p&gt;I learned this the annoying way. I once asked an agent to “clean up” an authentication module and got a beautiful explanation, a broad refactor, and no obvious way to tell which behavior had changed. The response sounded reasonable, but the handoff was poor. I had to reconstruct the intended change from the conversation, split it into smaller commits, and write the tests the prompt should have required. The model was not the only problem. I had selected a chat-shaped workflow for a diff-shaped job.&lt;/p&gt;

&lt;p&gt;My current default is simple. Before I open Cursor, Claude Code, Codex, or v0, I write down the artifact in one sentence: “I need a reviewable diff,” “I need cited notes,” “I need an editable React component,” or “I need a disposable prototype.” Then I choose the surface that naturally produces that object. It sounds obvious, but it is a useful filter when a comparison article presents ten tools as if they were ten interchangeable chat windows.&lt;/p&gt;

&lt;p&gt;I am still cautious about the numbers in these roundups. One post gives decimal scores such as 9.6, 8.2, and 7.8; another uses tier letters; the underlying test corpus is usually not disclosed. I would take those scores with a grain of salt. Product boundaries also move quickly: Cursor now talks about plans and long-running agents, Claude Code has expanded beyond a bare terminal experience, and v0 now describes itself as capable of more than mockups. A feature that belongs in one column this month may belong in another by the next release.&lt;/p&gt;

&lt;p&gt;The uncertainty does not make the comparisons useless. It just changes how I read them. I use the scorecard to discover candidates, then I judge the handoff. Can I review it? Can I rerun it? Can I attribute it to a source? Can I merge it without translating a polished paragraph into engineering work? If the answer is no, a higher model score probably will not save the workflow.&lt;/p&gt;

&lt;p&gt;For now, my stack is still Cursor for fast editor-centered changes, Claude Code for repository exploration and awkward multi-file work, GitHub Copilot for lightweight inline assistance, and v0 when I need to explore a UI direction quickly. Codex is my extra option when the task is better expressed as a goal for an agent than as a sequence of edits. I may change that mix after the next round of releases, but the artifact column is staying.&lt;/p&gt;

&lt;p&gt;The AI tool that wins a comparison is not necessarily the one that helps a team ship. The one that leaves a clear, testable, reviewable handoff usually has the better chance.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>The picking roundup has forked into three pricing camps and none of them names the others</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Sun, 26 Jul 2026 11:10:07 +0000</pubDate>
      <link>https://dev.to/ninghonggang/the-picking-roundup-has-forked-into-three-pricing-camps-and-none-of-them-names-the-others-1oga</link>
      <guid>https://dev.to/ninghonggang/the-picking-roundup-has-forked-into-three-pricing-camps-and-none-of-them-names-the-others-1oga</guid>
      <description>&lt;p&gt;I went down a rabbit hole this morning reading the late-2025 Juejin picking roundups side by side, and what crystallized for me is that the AI Agent roundup has quietly become its own picking camp with its own scorecard that does not acknowledge the editor-IDE picking camp or the consumer-chat picking camp even when the brand underneath is the same. The August and October GitHub trending pieces were all agent-shaped — supermemory, claude-cookbooks, Agent-S, prompt-eng-interactive-tutorial, and Archon the MCP command center. A separate Juejin roundup on free AI Agent tools then published its own top tier with LangGraph for data analysis, CrewAI for research clusters, ChatGPT Agent for office tasks, and Notion AI Agent for document work, each ranked on a 核心功能 plus 适用人群 plus 使用建议 triad and each priced at 免费. Meanwhile the editor-IDE roundup at juejin.cn/post/7571272874847010868 was ranking Cursor and Claude Code on an S/A/B/D tier letter with no row for LangGraph or CrewAI, and the consumer-chat roundup was ranking ChatGPT plus Gemini plus Claude at 140元 per month on a yuan-mode monthly bill with no row for ChatGPT Agent either. Three picking roundups in the same week, three shortlists, and the agent shortlist is the only one that explicitly calls itself 免费 across all rows.&lt;/p&gt;

&lt;p&gt;The angle I want to put down is that the AI Agent picking format has specialized into its own scorecard on 核心功能 plus 适用人群 plus 使用建议 plus 免费 pricing, and the editor-IDE and consumer-chat camps have no shared bridge to it even when the brand is the same. The agent roundup names ChatGPT Agent and Notion AI Agent as the entry-tier products on a 免费 axis. The consumer-chat roundup names ChatGPT Plus at 140元 per month and Claude Pro at 140元 per month on a paid-subscription axis. The same ChatGPT brand is 推荐 in both roundups but as different products — ChatGPT Plus on a yuan-mode monthly bill and ChatGPT Agent on a free quota — and neither roundup names the other. To be fair I would take the exact 免费 labels with a grain of salt because the operational cost of running ChatGPT Agent through a Plus subscription is never itemized in the agent roundup, but the structural tell is that the picking format has now forked into three camps that each publish their own pricing column and refuse to acknowledge the others.&lt;/p&gt;

&lt;p&gt;The meta-pattern I want to write down is that the late-2025 picking roundups have forked into scorecards that publish on different pricing axes with no conversion table between them, and the engineer trying to assemble a 2026 stack across editor IDE plus consumer chat plus AI agent has to do the three-way translation by hand. The editor-IDE roundup sorts on chat-plus-tool-call plus terminal evidence plus 非技术用户 shippability and prints an S/A/B/D tier letter with Cursor Pro at 20 dollars and Codeium at 免费. The consumer-chat roundup sorts on 多模态 plus 通用性强 and prints 首选 plus 其他 at 140元 per month across Gemini Pro, ChatGPT Plus, and Claude Pro. The agent roundup sorts on 核心功能 plus 适用人群 plus 使用建议 and prints 免费 across LangGraph, CrewAI, ChatGPT Agent, and Notion AI Agent. Not a single roundup has a column for what the next camp sorts on. Honestly I am a little skeptical of any 2026 picking workflow that grabs one of these three camps and treats it as the answer, because each was written for a different reader-job and the cross-camp conversion is in none of the roundups.&lt;/p&gt;

&lt;p&gt;The practical takeaway I want to write down is that I now read the three camps as three separate pricing dialects of the same AI surface and treat the cross-camp conversion as my own integration work. The editor-IDE roundup did name Cursor at S档 and Claude Code plus Codex at A档. The consumer-chat roundup did name Gemini as 首选 for 多模态 and ChatGPT as 首选 for 通用性强 and Claude as 其他 for 代码能力强. The agent roundup did name LangGraph, CrewAI, ChatGPT Agent, and Notion AI Agent, all on 免费. My gut says the reader has to hold all three shortlists at once and not pick one as the answer, because the Cursor the editor roundup puts at S档 is also the runtime host for the Claude Code the agent roundup names as a client for Archon, and the ChatGPT Plus the consumer roundup prices at 140元 per month is also the entry tier for the ChatGPT Agent the agent roundup lists as 免费, and the cross-camp integration cost is the engineering work the roundups should have done but didn't.&lt;/p&gt;

&lt;p&gt;I will reassess in three months. For now I am still on Cursor Pro plus Claude Code for coding, ChatGPT Plus for general chat, and Gemini AI Pro for the Sheets AI features I already use, and I have started running LangGraph locally for the data-pipeline experiments I used to hand off to a notebook. Give it six months and I expect either one of the three camps to add a cross-camp conversion column listing the same brand across editor plus chat plus agent surfaces, or a fourth camp to spin up an integrated AI-stack monthly bill that combines all three axes, and whichever moves first will tell me whether the picking format has finally noticed that the editor-IDE plus consumer-chat plus AI-agent shortlists are three views of the same 2026 wallet.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>The pricing guide that contradicts itself between the API paragraph and the forecast paragraph</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Sat, 25 Jul 2026 11:07:17 +0000</pubDate>
      <link>https://dev.to/ninghonggang/the-pricing-guide-that-contradicts-itself-between-the-api-paragraph-and-the-forecast-paragraph-35i4</link>
      <guid>https://dev.to/ninghonggang/the-pricing-guide-that-contradicts-itself-between-the-api-paragraph-and-the-forecast-paragraph-35i4</guid>
      <description>&lt;p&gt;I went down a rabbit hole this morning reading a single 2025 AI pricing guide on Juejin side by side with my own monthly bills, and what crystallized for me is that the guide has quietly split its own future in half without telling the reader. One paragraph in the middle says API prices across the industry dropped ninety percent during 2024 to 2025 and frames that as the dominant pricing trend. A later paragraph in the same article forecasts ChatGPT Plus rising from twenty dollars today to forty-four dollars by 2029 and frames that as the dominant future. Both paragraphs are inside one article, both speak to my monthly wallet, both call themselves the trend, and neither one acknowledges the other exists. The reader has to do the contradiction integration by hand, and that contradiction is the structural tell I want to put down this morning.&lt;/p&gt;

&lt;p&gt;The angle I want to write down is that the pricing-guide format has now broken its own timeline into two incompatible forecasts that publish side by side, and the cross-timeline bridge is in nobody's coverage. The API-decline paragraph points at GPT-4o and Claude Sonnet and Gemini 2.5 Pro and says raw token spend is collapsing, with ninety percent off across the year as the headline number. The forecast paragraph points at the same ChatGPT Plus the API paragraph is about, and says it will hit forty-four dollars by 2029 — a hundred-and-twenty percent rise from today. API down ninety, subscription up a hundred-and-twenty, both inside one article, both authored with the same confident tone. To be fair I would take both the ninety-percent drop and the forty-four-dollar forecast with a grain of salt because the time window on one is twelve months and the time window on the other is four years, but the structural tell is that a reader trying to predict their 2026 monthly bill cannot read either paragraph as the trend because the article has now published both trends at once.&lt;/p&gt;

&lt;p&gt;The meta-pattern I want to put down is that the pricing-guide format has specialized internally into a token-spend timeline and a subscription-spend timeline that coexist inside the same guide and refuse to bridge each other. The token-spend timeline names ChatGPT, Claude, Gemini, Grok, Midjourney, Stable Diffusion, Perplexity, and Poe as the comparison set and ranks them on dollar-per-million-tokens at falling prices. The subscription-spend timeline names ChatGPT Plus, Claude Pro, Gemini AI Pro, Grok Premium, Midjourney Standard, and Perplexity Pro as the comparison set and ranks them on dollar-per-month at rising prices with one exception. When I try to assemble my 2026 stack, I read the API paragraph to feel good about my backend bill and the forecast paragraph to feel anxious about my consumer bill, and the same guide has now handed me a declining number and a rising number for the same ChatGPT brand at the same moment. Honestly I am a little skeptical of any 2026 pricing workflow that reads the guide as one coherent voice, because the format has now visibly bifurcated inside a single article and the cross-timeline integration is left to me.&lt;/p&gt;

&lt;p&gt;The practical takeaway I want to write down is that I now read the pricing guide the way I read the four-camp pieces — as two forecasts that happen to share a homepage rather than as one forecast with two paragraphs. The token-spend forecast did name ChatGPT, Claude, Gemini, Grok, Midjourney, Stable Diffusion, Perplexity, and Poe as the comparison shortlist. The subscription-spend forecast did name ChatGPT Plus, Claude Pro, Gemini AI Pro, Grok Premium, Midjourney Standard, and Perplexity Pro on the same shortlist, with Midjourney at ten to one-hundred-twenty dollars and Grok having already raised from twenty-two to forty dollars and ChatGPT Plus projected to forty-four dollars. My gut says the reader has to hold both timelines at once and not pick one as the answer, because my API bill and my subscription bill are the same bill by year-end and the guide has now published two trends for one wallet.&lt;/p&gt;

&lt;p&gt;I will reassess in three months. For now I am still mostly on Cursor Pro plus Claude Code for coding, ChatGPT Plus for general chat, Gemini AI Pro for the Sheets AI features I already use, and the Anthropic API for the few backend calls I run, and I am keeping a per-bill log of which paragraph said which trend because the pricing format is no longer going to surface this bifurcation for me. Give it six months and I expect either the pricing guides to publish a single timeline that reconciles the API decline with the subscription rise or a third camp to spin up that publishes a unified monthly-bill forecast that combines both, and whichever moves first will tell me whether the format has finally noticed that the declining-API and rising-subscription trends are the same monthly bill seen from two angles.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>Two Juejin reviews of AI coding tools, two top-eights that share one name</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Fri, 24 Jul 2026 11:07:21 +0000</pubDate>
      <link>https://dev.to/ninghonggang/two-juejin-reviews-of-ai-coding-tools-two-top-eights-that-share-one-name-1alh</link>
      <guid>https://dev.to/ninghonggang/two-juejin-reviews-of-ai-coding-tools-two-top-eights-that-share-one-name-1alh</guid>
      <description>&lt;p&gt;I went down a rabbit hole this evening reading two Juejin roundups from the same two weeks of December 2025 that both purport to rank the top AI coding tools for 2026, and what crystallized for me is that they are not even close to ranking the same universe. One prints Cursor at S档 and Claude Code and Codex CLI at A档 on a tier letter, with v0 and Lovable above Rork and VibeCode App. The other prints Tencent CodeBuddy at 9.6 out of 10 on a five-axis decimal scorecard, with Cody at 8.2, Ghostwriter at 8.0, Codeium at 7.8, Tabnine at 7.6, CodeWhisperer at 7.5, JetBrains AI Assistant at 7.4, and Blackbox at 7.2. Same reader-question, same search results page, two posts, two completely disjoint top-eight rosters, and not a single shared reference across the two columns.&lt;/p&gt;

&lt;p&gt;The angle I want to put down is that the convergent-consensus I have been tracking in prior pieces has now hit its opposite case, and the opposite case is what the convergent-consensus has been hiding. The tier-letter piece sorts on chat-plus-tool-call capability plus terminal evidence plus "whether a non-technical user can ship with it", which is why Cursor and Lovable and v0 crowd the top and Tencent CodeBuddy never appears. The decimal piece sorts on code completion plus code generation plus 自主代理 plus 多模态 plus team collaboration plus 安全合规, which is why Tencent CodeBuddy's enterprise tier plus 等保三级 row lands it at 9.6 and Cursor never appears. To be fair I would take the exact 9.6 decimal and the exact tier letters with a grain of salt because the test corpus behind each is never disclosed, but the structural tell is that the two formats have not just used different scorecards — they have used different scoring apparatuses that produce different verdicts on different tools.&lt;/p&gt;

&lt;p&gt;The meta-pattern I want to write down is that the December 2025 search results page for "AI coding tools 2026" is now producing, in the same two-week window, two posts whose top-eight rosters share exactly one name — Replit. The tier-letter piece names Cursor, Claude Code, Codex CLI, Lovable, v0, Rork, VibeCode App, Anything, Chef, and Replit in roughly that order; the decimal piece names Tencent CodeBuddy, Cody, Ghostwriter, Codeium, Tabnine, CodeWhisperer, JetBrains AI Assistant, and Blackbox with no overlap except Replit on both. The convergent-consensus I have been writing about is visibly not converging — one post says the verdict is Cursor plus Claude Code, the other says the verdict is Tencent CodeBuddy plus Cody, and neither knows about the other. Honestly I am a little skeptical of any 2026 picking workflow that treats one of these pieces as the answer, because each was written to land on its own winner by sorting on axes that ignore the other piece's top picks entirely.&lt;/p&gt;

&lt;p&gt;The practical takeaway I want to put down is that I now read the December two-format standoff the way I used to read the four-camp pieces — as two separate ranking dialects sorting on two non-overlapping definitions of what counts as a coding tool, and the cross-format translation is the engineering work the reader has to do by hand. The tier-letter piece answers the within-capability question including non-coder use cases like Lovable and v0 and Rork, and the decimal piece answers the within-IDE question for an enterprise developer inside VS Code or JetBrains with compliance requirements. My gut says the reader has to combine both pieces to assemble a 2026 stack and not pick one as the answer, because Cursor and Claude Code are at the top of one column and absent from the other, and Tencent CodeBuddy is at the top of the other column and absent from the first, and the two recommendations are not in conflict — they are answering different reader-questions.&lt;/p&gt;

&lt;p&gt;I will reassess in three months. For now I am still mostly on Cursor Pro plus Claude Code for coding and ChatGPT Plus for general chat, and I am keeping a per-piece log of which December 2025 roundup named which tool in which tier, because the verdict apparatus has now visibly bifurcated inside a single search week and the search results page is not going to surface this bifurcation by itself. Give it six months and I expect either one of the two formats to add a row for the other's top picks, or a third camp to spin up that publishes a cross-format compatibility column listing which tools appear in which tier and which scorecard, and whichever moves first will tell me whether the December two-format standoff has finally collapsed into a cross-format verdict or whether the picking workflow has permanently turned into a multi-format integration job the engineer has to do by hand.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>Four Juejin pieces, four definitions of AI, no shared bridge</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Thu, 23 Jul 2026 11:08:47 +0000</pubDate>
      <link>https://dev.to/ninghonggang/four-juejin-pieces-four-definitions-of-ai-no-shared-bridge-1fbj</link>
      <guid>https://dev.to/ninghonggang/four-juejin-pieces-four-definitions-of-ai-no-shared-bridge-1fbj</guid>
      <description>&lt;p&gt;I went down a rabbit hole this morning reading four late-2025 Juejin pieces side by side that all purport to be about AI in 2025, and what crystallized for me is that none of them is even answering the same definition of what AI is. The first was a picking piece with an S/A/B/D tier letter ranking Cursor in S档 and Claude Code in A档 on a chat-and-tool-call scorecard. The second was Google's own 2025 年度最热门AI应用 recap that listed 学习理解, NotebookLM, 旅行游玩, 相册应用与图像编辑, 办公生产力, 购物, 硬件设备, and 个性化定制 as the top eight AI categories for the year, and there is no chatbot in any of those eight. The third was the October GitHub trending list that called Agent-S, claude-cookbooks, supermemory, TradingAgents-CN, and nanoGPT the 热门开源项目 under a star-velocity sort. The fourth was the 中文AI工具推荐 directory that pivoted halfway through to 神马中转 API as a workaround for blocked endpoints. Four posts in one search results page, four definitions of AI, none closer to translating to the others than picking the wrong column.&lt;/p&gt;

&lt;p&gt;The angle I want to put down is that the picking format and the vendor category list and the trending list and the content-commerce directory are now four separate camps each anchored on its own definition of AI, and the cross-camp bridge has never been published. The picking tier letter sorts on chat capability plus tool-call reliability plus terminal evidence. Google's top categories sort on what new artifact exists because the AI ran — a PDF transformed into a guided interactive tutorial, a research paper turned into a podcast, a spreadsheet cell turned into a semantic category. The GitHub trending list sorts on star velocity inside a BYOK meter none of the others ask about. The content-commerce directory sorts on whether the affiliate link works for the Chinese reader in 2026. To be fair I would take the exact tier letters and the exact star counts with a grain of salt because the test corpus behind each is never disclosed, but the structural tell is that the four camps agree on a small handful of names — Cursor, Claude Code, ChatGPT, NotebookLM, Agent-S — and disagree on what those names mean inside their respective definitions.&lt;/p&gt;

&lt;p&gt;The meta-pattern I want to write down is that the late-2025 AI roundup scene is now a four-definition standoff where each format is glued to one definition of AI and the cross-definition integration work has been left to the engineer. The picking piece answers the chat-tool-call question. Google's categories answer the output-transformation question. The GitHub trending list answers the star-velocity question inside a BYOK meter. The content-commerce directory answers the affiliate-link-works-for-China question. None has a column for what the next camp sorts on. NotebookLM turning a research paper into a podcast is not on the picking scorecard. Cursor's S档 is not on Google's list of top eight categories. When I try to assemble a 2026 stack for the projects I am shipping this quarter, I end up reading the picking piece to shortlist Cursor plus Claude Code plus Codex, then going to Google's list to remember that NotebookLM plus the Sheets AI function are the surfaces the vendor behind two of my picks is actually shipping, then checking the GitHub trending list to see whether mem0 or supermemory has the memory layer any of these picking tools assume. Honestly I am a little skeptical that treating one of these four posts as the answer is enough, because each was written inside a different definition of what AI is.&lt;/p&gt;

&lt;p&gt;The practical takeaway I want to write down is that I now read the four pieces as four separate ranking dialects of four different questions rather than as four answers to the same question. The picking tier letter is useful for the within-coding-tool shortlist and not for the artifact the vendor behind the top picks is actually shipping. Google's category list is useful for orienting what the AI vendors consider real wins and not for picking between Cursor and Claude Code. The GitHub trending list is useful for finding memory and MCP server substrates the picking tier letter assumes exist but never ranks. My gut says the reader has to do the four-way translation by hand, and the translation is in none of the four posts, and one cross-camp column could have answered the whole integration question.&lt;/p&gt;

&lt;p&gt;I will reassess in three months. For now I am still on Cursor Pro plus Claude Code for coding, ChatGPT Plus for general chat, Gemini AI Pro for the Sheets AI and the Gemini 3 引导式学习 feature, NotebookLM for the occasional PDF-to-podcast experiment, and I am keeping a per-tool log of which camp each tool appeared in. Give it six months and I expect either one of the four camps to absorb the other three and publish a cross-camp translation column, or a new fifth camp to spin up with a unified definition of what AI actually is in 2026, and whichever moves first will tell me whether the four-definition standoff has finally collapsed or whether the AI reading plan has permanently turned into a four-camp integration job I have to do by hand.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>Four Juejin 2025 rankings, four scorecards, one winner</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Wed, 22 Jul 2026 11:06:58 +0000</pubDate>
      <link>https://dev.to/ninghonggang/four-juejin-2025-rankings-four-scorecards-one-winner-3e0j</link>
      <guid>https://dev.to/ninghonggang/four-juejin-2025-rankings-four-scorecards-one-winner-3e0j</guid>
      <description>&lt;p&gt;I went down a rabbit hole this morning reading four late-2025 Juejin roundups side by side that each claimed to be a different ranking on different axes, and what crystallized for me is that all four of them have quietly converged on the same vendor as the top pick despite using visibly different scorecards. The first was an eight-tool December coding ranking that printed CodeBuddy at 9.6 out of 10 with Cody at 8.2, Ghostwriter at 8.0, Codeium at 7.8, Tabnine at 7.6, CodeWhisperer at 7.5, JetBrains AI Assistant at 7.4, and Blackbox at 7.2 on a five-axis decimal scorecard. The second was a nine-tool AI IDE ranking that printed CodeBuddy at five stars as 综合评分最高 with eight alternatives at four or three stars behind it. The third was a seven-tool fullstack ranking that printed CodeBuddy at five stars as 综合评分最高 with DevFlow, CodeStudio, CloudSync, FlowDesign, AgentHub, and CodePulse at four or three stars behind it. The fourth was a six-tool 2026 ranking that crowned Trae as 中文开发者的最佳选择 with GitHub Copilot, Amazon CodeWhisperer, Cursor, Tabnine, and Codeium behind it. Four posts, four different scorecards, four different scopes, and one of two Chinese-coded IDE tools at the top of every single list.&lt;/p&gt;

&lt;p&gt;The angle I want to put down is that the convergent-consensus I have been tracking this month has now grown a layer down, and the layer is not "the same vendor wins" but "the same vendor wins on scorecards that explicitly sort on different axes". The December decimal piece sorts on code completion plus code generation plus 自主代理 plus 多模态 plus team collaboration plus 安全合规. The IDE ranking sorts on technology stack plus engineering capability plus industry fit plus user experience plus 企业合规. The fullstack ranking sorts on zero-config setup plus 中文精准交互 plus multi-modal input plus 端到端适配 plus 微信生态 integration plus 双端适配. The 2026 ranking sorts on language adaptation plus ecosystem integration plus security plus collaboration plus 全流程自动化. Five axes here, five axes there, five axes everywhere, and the same vendor at the top of every column. To be fair I would take the exact tier letters and the exact decimals with a grain of salt because the test corpus behind each scorecard is never disclosed, but the structural tell is that the format is no longer just converging on a verdict — it is converging on a verdict apparatus that always lands on the same vendor.&lt;/p&gt;

&lt;p&gt;The meta-pattern I want to write down is that the late-2025 picking roundups have not just converged on the same vendor recommendation, they have converged on a ranking format that absorbs whatever corpus it is given and still prints the same name on top, and the cross-format integration question is now whether the format itself is the signal or whether the format is the noise. The four pieces name eight, nine, seven, and six tools respectively, and the tools change but the winner does not — CodeBuddy in three of the four lists, Trae in the fourth, with the underlying story being the same Chinese-coded IDE win. When I try to assemble a 2026 stack, I end up reading all four pieces to confirm what I already know from any single piece, and the cross-format translation is in every piece but the verdict is not. Honestly I am a little skeptical of any 2026 picking workflow that treats one of these pieces as the answer, because each piece was written to land on the same vendor and the variety in the scorecard axes is decoration around a predetermined recommendation.&lt;/p&gt;

&lt;p&gt;The practical takeaway I want to write down is that I now read the four pieces as a single convergent-consensus signal wearing four different scorecards rather than as four separate rankings to triangulate. The December piece did name Cursor, Claude Code, Codex, Cody, Ghostwriter, Codeium, Tabnine, CodeWhisperer, and Blackbox in roughly the order I would have ranked them. The 2026 piece named GitHub Copilot, Amazon CodeWhisperer, Cursor, Tabnine, and Codeium in roughly the same order. My gut says the reader has to assume the format has converged on the vendor recommendation and treat the scorecard axes as a confirmation rather than a measurement, because the measurement apparatus is calibrated to land on the same name regardless of what axes it claims to sort on.&lt;/p&gt;

&lt;p&gt;I will reassess in three months. For now I am still mostly on Cursor Pro plus Claude Code for coding, ChatGPT Plus for general chat, and Gemini AI Pro for the Sheets AI and Drive features I already use, and I am keeping a per-tool log of which roundup named which tool because the picking format is no longer going to surface this. Give it six months and I expect either one of these roundups to publish a corpus disclosure that breaks the convergent-consensus signal or a new vendor to ship a scorecard that lands on a different winner, and whichever moves first will tell me whether the format has finally noticed that the cross-format convergent-consensus is the structural tell that the format itself has ossified.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>Three Juejin roundups, three scorecards, no shared bridge</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Tue, 21 Jul 2026 11:09:30 +0000</pubDate>
      <link>https://dev.to/ninghonggang/three-juejin-roundups-three-scorecards-no-shared-bridge-3ak2</link>
      <guid>https://dev.to/ninghonggang/three-juejin-roundups-three-scorecards-no-shared-bridge-3ak2</guid>
      <description>&lt;p&gt;I went down a rabbit hole this morning reading three different late-2025 Juejin roundups side by side, all of which purported to answer what AI to use in 2026, and what crystallized for me is that the three of them are actually answering three different questions on three different scorecards with no shared bridge between any of them. The first was a December picking piece that ranked eight tools on a five-axis decimal scorecard — CodeBuddy 9.6 at the top, Blackbox 7.2 at the bottom, with Cody, Ghostwriter, Codeium, Tabnine, CodeWhisperer, and JetBrains AI Assistant in between. The second was the 2025年度盘点 that named Gemini as 首选 for 多模态 plus 超长上下文, ChatGPT as 首选 for 通用性强, and Claude as 其他 for 代码能力强, with all three priced at 140元 per month on a yuan-mode monthly bill. The third was a November frontend piece that named Cursor as 综合能力第一, GitHub Copilot as 最佳生态集成, Codeium as 性价比之王, and V0.dev as 前端UI生成专家, with a five-star rating column tacked onto each row. Three posts, three scorecards, one reader question, and not a single one of them produces a recommendation I can compare against the other two.&lt;/p&gt;

&lt;p&gt;The angle I want to put down is that all three roundups claim to answer the same question — what AI tool should an engineer use in 2026 — but each is sorting on a different axis, and the cross-axis bridge has never been published. The decimal piece sorts on code completion plus code generation plus 自主代理 plus 多模态 plus team collaboration, and prints its verdict as 9.6/10. The yuan-mode annual review sorts on 推荐 plus 核心优势 plus 访问门槛 plus 月费 and prints 首选 or 其他, with the monthly bill reading 140元 across Gemini Pro, ChatGPT Plus, and Claude Pro. The November frontend piece sorts on 综合能力 plus 生态集成 plus 性价比 and prints a five-star rating. The structural tell is that the three pieces agree on a handful of names — Cursor, GitHub Copilot, ChatGPT, Claude — and disagree on what those names mean. Cursor is 综合能力第一 in the November piece and absent from the December decimal scorecard. ChatGPT is 首选 in the annual review and absent from the November piece. CodeBuddy is the highest-decimal in the December piece and absent from both of the others.&lt;/p&gt;

&lt;p&gt;The meta-pattern I want to write down is that the December 2025 search results page for what AI to use in 2026 is now a three-format standoff where each format asks a different reader-question under the same headline, and the engineer has to do the cross-format multiplication by hand. The decimal piece answers the within-coding-IDE question. The yuan-mode annual review answers the within-consumer-chat question. The November frontend piece answers the within-frontend-UI question. None has a column for what the next scorecard sorts on, none has a row for how the decimal piece measures frontend UI, and none acknowledges the others even when the recommended tool is the same name. When I try to assemble a 2026 stack for the projects I am shipping this quarter, I end up reading the decimal piece to shortlist CodeBuddy plus Cursor plus Cody, then the annual review to figure out that the same stack costs 140元 a month for ChatGPT Plus plus 20 dollars for Cursor Pro plus another 20 for Claude Pro, then the November piece to confirm that V0.dev plus Codeium plus GitHub Copilot is the cheaper tier for the frontend side, and the cross-format total ends up closer to a mortgage payment than any of the roundups hints. Honestly I am a little skeptical that grabbing one of these three posts and treating it as the answer is enough, because each was written for a different reader-job.&lt;/p&gt;

&lt;p&gt;The practical takeaway I want to write down is that I now read the three camps as three separate ranking dialects of the same question rather than as three separate answers. The decimal piece is useful for the within-coding-IDE shortlist and not for the monthly bill. The yuan-mode annual review is useful for the within-consumer-chat shortlist and not for the IDE ranking. The November frontend piece is useful for the within-frontend shortlist and not for either of the others. My gut says the reader has to do the three-way translation by hand — what's the decimal for ChatGPT Plus, what's the monthly bill for Cursor Pro, what's the five-star rating for Claude Code — and the translation is in none of the roundups, and one column could have answered the whole question.&lt;/p&gt;

&lt;p&gt;I will reassess in three months. For now I am still on Cursor Pro plus Claude Code for coding, ChatGPT Plus for general chat, Gemini AI Pro for the Sheets AI work, and NotebookLM for the occasional PDF experiment, and I am keeping a per-tool log of which camp each tool appeared in and what the verdict was, because the search results page is not going to surface this. Give it six months and I expect either one format to absorb the other two and print a unified verdict, or a fourth camp to spin up with a cross-camp translation column, and whichever moves first will tell me whether the three-camp standoff has finally collapsed or whether the reading plan has permanently turned into a multi-format integration job.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>The November Juejin Roundups Have Crossed Into a Content-Commerce Funnel</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Mon, 20 Jul 2026 11:08:53 +0000</pubDate>
      <link>https://dev.to/ninghonggang/the-november-juejin-roundups-have-crossed-into-a-content-commerce-funnel-317g</link>
      <guid>https://dev.to/ninghonggang/the-november-juejin-roundups-have-crossed-into-a-content-commerce-funnel-317g</guid>
      <description>&lt;p&gt;I keep a private reading log of every Juejin AI roundup I open, and reading the late-2025 picking posts this evening, what jumped out is that the format I have been dissecting all month has visibly crossed from editorial coverage into a content-commerce funnel, and the structural tell is no longer hidden. The November 2025 picking post at juejin.cn/post/7571272874847010868 prints an S/A/B/D tier letter for ten tools (Cursor at S档, Claude Code and Codex at A档, Rork and VibeCode App and Anything at B档) and never publishes the corpus behind the tiers, but the post next to it at juejin.cn/post/7569484179890438179 prints the same shape and tacks on a 预算选择 block that ends with a 优惠链接 column and a five-star rating on every row. The 中文AI工具推荐 directory at juejin.cn/post/7555405171384778778 starts as a plain ten-product recommendation list and pivots halfway through the last section into a 神马中转 API walkthrough that is functionally an affiliate bridge for users who cannot reach ChatGPT or Claude directly. To be fair I would take the exact tier letters with a grain of salt because the corpus is never disclosed, but the structural tell is what has been rattling around in my head this evening. The roundup format has stopped being editorial coverage and started being a sales funnel, and the verdict is now pre-baked into the structure.&lt;/p&gt;

&lt;p&gt;The piece that pushed me over the edge was the 中文AI工具推荐 directory, because it is the cleanest example I have read this quarter of the format-switch happening mid-article. The first six sections are written in the standard recommendation register — pick 2 to 3 tools, 通用聊天, 写作与内容创作, AI 绘图, 视频生成与编辑, 编程辅助, 办公效率 — and the seventh section, 国内直接用不了的，可以使用神马中转API, pivots without warning into a platform pitch. The previous sections never asked the reader to compare an API gateway with a chat product, and the seventh section never discloses its relationship with the gateway operator. The reader was being sorted, not informed. My gut says this is what the format looks like when it has finished crossing from editorial coverage into a revenue-share funnel, and the crossing happens inside the same post rather than across the search results page.&lt;/p&gt;

&lt;p&gt;The meta-pattern I want to put down is that the late-2025 roundup format has visibly split into two camps that share almost no format conventions anymore, and the engineer who pulls a single piece off the search results page and treats it as the verdict is now doing a paid-content discount calculation the post never asks them to do. The editorial camp still prints decimal scorecards and ranking currencies — the November IDE ranking with CodeBuddy at 9.6, Cody at 8.2, Ghostwriter at 8.0, Codeium at 7.8 on a five-axis scale, the 2025年度盘点 with its 元 per month price column. The content-commerce camp prints affiliate pricing tables, five-star ratings, 优惠链接 columns, 学生免费 repeated four times in a sponsor section, and a 神马中转 API pivot that sells China-workaround access to the very tools the editorial camp is trying to rank. Honestly I am a little skeptical of any 2026 roundup workflow that pulls a single piece off the search results page without checking which camp the piece is in, because the editorial verdict and the affiliate verdict have started pointing at different tools.&lt;/p&gt;

&lt;p&gt;The practical takeaway I want to write down is that I read the two camps differently now, and the discount is roughly the difference between the editorial rank and the closest sponsor-aligned alternative. The November picking post did name Cursor, Claude Code, Codex, CodeBuddy, Cody, Ghostwriter, Codeium, Tabnine, CodeWhisperer, and Blackbox in roughly the order I would have ranked them. The 中文AI工具推荐 directory named a different ten tools and ended on a 神马中转 API pitch. The two posts share the same search query and almost zero format conventions. I have started bookmarking the editorial camp and treating the content-commerce camp as a paid-placement signal to discount rather than as a recommendation, and that split is what the format is going to keep looking like until one camp publishes an explicit disclosure the other does not.&lt;/p&gt;

&lt;p&gt;I will reassess in three months. For now I am still mostly on Cursor Pro plus Claude Code for coding, ChatGPT Plus for general chat, and Gemini AI Pro for the Sheets AI and Drive features I already use, which is roughly where I have landed the last several times I have written this paragraph. What has changed this evening is that I now treat any Juejin roundup that opens with a tier letter or a five-star rating and closes with a 优惠链接 column or an API gateway pitch as a sponsored post first. Give it six months and I expect either the editorial camp to add an explicit sponsor-disclosure block or the content-commerce camp to harden into its own search-result niche that the reader learns to skip, and whichever moves first will tell me whether the format has finally noticed that the verdict on the left and the affiliate on the right are no longer the same product.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>OpenAI Quietly Cut Codex's Context Window 27%, and Most Harnesses Still Don't Know</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Sun, 19 Jul 2026 23:10:16 +0000</pubDate>
      <link>https://dev.to/ninghonggang/openai-quietly-cut-codexs-context-window-27-and-most-harnesses-still-dont-know-31ih</link>
      <guid>https://dev.to/ninghonggang/openai-quietly-cut-codexs-context-window-27-and-most-harnesses-still-dont-know-31ih</guid>
      <description>&lt;p&gt;I keep a small folder of OpenAI Codex release notes in my bookmarks bar, and this week I caught a quiet one I almost missed. PR 33972, merged July 18 into the release/0.144 branch by sayan-oai, lowered the bundled model metadata's context-window number from 372,000 tokens to 272,000 tokens. That is roughly a 27 percent reduction in advertised capacity for gpt-5.6-sol inside Codex. The PR title is "Backport refreshed bundled model metadata to 0.144", which sounds boring, and the diff is 64 additions and 54 deletions against a single JSON metadata file. To be fair I would take the literal 27 percent with a grain of salt because the JSON is the bundled-config number, not a server-side enforcement knob — but the practical reading is that any harness still asserting 372k as the window is now wrong, and OpenAI's own Codex release notes for 0.144.6 confirm the cap is back at 272k.&lt;/p&gt;

&lt;p&gt;The angle that grabbed me, and what I want to put down here, is that the reduction is happening at the same time that harnesses outside Codex are configured against the older 372k number, and Kun Chen's X post from July 13 flagged this directly: "important tip if you are using gpt 5.6 but not through codex, OpenAI just pointed out requests beyond 272k tokens are over-charged, and many harnesses right now set gpt 5.6 context window at 372k, which will cause you to lose quota more quickly than it should". His specific remediation was telling Pi to change its gpt-5.6 context window to 272k. The Reddit thread on r/codex picked it up the same day with 90 upvotes and 83 comments, mostly engineers saying they had no idea the harness default was still 372k. OpenAI's response on the same channel framed it as "No nerfing, only good stuff!" plus a 10 percent usage bump from inference optimizations landing on the Sol subscriptions. My gut says both can be true at once — the inference savings are real, and the older harnesses are still over-charging on requests they think fit in 372k but actually do not.&lt;/p&gt;

&lt;p&gt;The reason I want to write this down is that I run gpt-5.6-sol through both Codex and a couple of non-Codex harnesses for the projects I am shipping this quarter, and I had assumed my context-window setting was correct because the picking roundups I read in late June still quoted 372k as the headline number for Sol. The picking-roundup framework that ranks IDE workflow and code generation on S/A/B/D tier letters and 9.6-out-of-10 decimal scores does not have a column for the actual configured window versus the marketed window — and the pricing-guide framework that prices Gemini Pro at 19.99 dollars per month, ChatGPT Plus at 20 dollars, and Claude Pro at 20 dollars does not have a row for the gpt-5.6 over-charge delta. The format-vs-format drift I wrote about earlier this month is now showing up inside the same week as a token-billing change that the format cannot measure, because the format is sorting on tier letters and dollars and not on what the harness thinks the window is.&lt;/p&gt;

&lt;p&gt;The practical takeaway I want to write down is that I went through my own harnesses this morning and found one Pi configuration still set to 372k and one Cursor-side agent loop still asserting the older window in its prompts, both of which I changed to 272k. Honestly I am a little skeptical that this is the kind of change most engineers will catch from the picking scorecards, because the scorecards measure chat quality and pricing tier and not configured window, and the over-charge shows up as a slightly faster quota burn rather than a visible error. The companion PR 34009 and the release tag rust-v0.144.6 are the references worth bookmarking if you are running gpt-5.6-sol through anything other than Codex itself.&lt;/p&gt;

&lt;p&gt;I will reassess in three months. For now I have set every gpt-5.6-sol harness I touch to 272k, and I have started keeping a per-harness log of which window number each one is configured against because the picking roundups are not going to surface this. Give it six months and I expect either the picking roundups to add a configured-window column or the harness vendors to start pinning 272k as a default the way Codex just did, and whichever moves first will tell me whether the format has finally noticed that the marketed window and the configured window can drift inside a single quarter without anyone catching it on the scorecard.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>When the Google Recap and the Juejin Picking Roundup Disagree on What Counts as AI</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Sun, 19 Jul 2026 11:08:07 +0000</pubDate>
      <link>https://dev.to/ninghonggang/when-the-google-recap-and-the-juejin-picking-roundup-disagree-on-what-counts-as-ai-1id6</link>
      <guid>https://dev.to/ninghonggang/when-the-google-recap-and-the-juejin-picking-roundup-disagree-on-what-counts-as-ai-1id6</guid>
      <description>&lt;p&gt;I went down a rabbit hole this morning reading the Google 2025年度最热门AI应用 recaps side by side with the late-2025 Juejin picking roundups, and what crystallized for me is that Google's own annual ranking of most-popular AI products does not include a chatbot at all. The top eight categories Google itself published for 2025 were 学习理解, NotebookLM的应用, 旅行游玩, 相册应用与图像编辑, 办公生产力, 购物, 硬件设备, and 个性化定制. Every one of those is either an output-transformation surface or an AI-embedded-in-existing-product surface, not a chat surface. The reader-jobs Google says are most popular inside its own product line — turn a PDF into a podcast, turn a math problem into a guided-learning tutorial, turn a spreadsheet cell into a semantic category, turn a search query into a working mini-app, turn a holiday cookie photo into a structured recipe book — do not show up on the Juejin picking scorecard at all, because the picking scorecard is still measuring how good the typed-prompt reply is.&lt;/p&gt;

&lt;p&gt;The Juejin picking roundups I read this morning are still committed to ranking raw chat capability. The late-2025 picking piece puts Cursor at S档 for IDE workflow and Claude Code at A档 for code generation, with Codex at A档 and a long tail of CodeBuddy at 9.6, Cody at 8.2, Ghostwriter at 8.0, and Codeium at 7.8 on a five-axis decimal scorecard. The 2025年度盘点 piece names Gemini as 首选 for 多模态 plus 超长上下文 plus 原生搜索, ChatGPT as 首选 for 通用性强, and Claude as 其他 for 代码能力强 and 长文写作. Both roundups ask how well this AI responds to a typed prompt and answer in either S/A/B/D tier letters or 9.6-out-of-10 decimal scores. To be fair I would take the exact decimals with a grain of salt because the test corpus is never disclosed, but the structural tell is that neither roundup has a column for what artifact the AI produces. NotebookLM turning a research paper into a podcast is not on the picking scorecard. Gemini 3 turning a mortgage-rate question into an interactive calculator is not on the picking scorecard. The picking scorecard measures how smart the text response is and the Google recap measures what new thing exists in the world because the AI ran.&lt;/p&gt;

&lt;p&gt;The meta-pattern I want to put down is that the 2026 search results page has at least three formats that all claim to answer what AI to use in 2026, but they are actually answering three different questions and the cross-format bridge is invisible. The picking roundups ask which chat is most capable and answer with tier letters and decimal scores. The Google annual recap asks which AI products users actually engage with inside Google's own ecosystem and answers with output-transformation and product-integration tools. The October GitHub trending recap asks which open-source project got the most stars last month and answers with prompt-eng-interactive-tutorial, Agent-S, claude-cookbooks, supermemory, and TradingAgents-CN — none of which is a chat surface either. Honestly I am a little skeptical of any 2026 roundup workflow that pulls a single piece off the search results page and treats it as the answer to what AI to use, because each piece was written for a different reader-job and the cross-format integration is left to the engineer.&lt;/p&gt;

&lt;p&gt;The practical takeaway I want to write down is that the picking roundups are still useful for the within-chat anchor and the Google recaps are useful for the within-transformation anchor, but neither format is useful for the cross-format re-rank most engineers are quietly trying to do this quarter. The picking roundup did name Cursor, Claude Code, Codex, CodeBuddy, Cody, Ghostwriter, Codeium, Tabnine, CodeWhisperer, and Blackbox as the editor-loop shortlist. The Google recap did name 学习理解, NotebookLM, the Gemini Sheets AI function, Gemini in Search, and the AI mode mini-apps as the transformation shortlist. Neither list is good at telling the engineer which transformation tool to add on top of which chat-capable tool, because the picking scorecard never asks what artifact each row produces and the Google recap never asks how each row compares on raw chat quality. My gut says the reader has to do the cross-format multiplication by hand, and the multiplication is in none of the roundups.&lt;/p&gt;

&lt;p&gt;I will reassess in three months. For now I am still mostly on Cursor Pro plus Claude Code for coding, ChatGPT Plus for general chat, Gemini AI Pro for the Sheets AI and Drive features I already use, and NotebookLM for the occasional PDF-to-podcast experiment. What has changed is that I now read the Google recaps as a separate category — output-transformation tools and AI-in-product integrations — rather than as a chat-ranking I should fold into the picking scorecard. Give it six months and I expect either the picking roundups to add a what-artifact-does-this-AI-produce column or the Google-style transformation recaps to get their own dedicated ranking niche, and whichever moves first will tell me whether the format has finally noticed that the reader-job has forked.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
    <item>
      <title>Seven ranking frameworks, one search page, zero translation tables</title>
      <dc:creator>ninghonggang</dc:creator>
      <pubDate>Sat, 18 Jul 2026 11:09:05 +0000</pubDate>
      <link>https://dev.to/ninghonggang/seven-ranking-frameworks-one-search-page-zero-translation-tables-5c12</link>
      <guid>https://dev.to/ninghonggang/seven-ranking-frameworks-one-search-page-zero-translation-tables-5c12</guid>
      <description>&lt;p&gt;I went down a rabbit hole this morning reading the late-2025 Juejin AI roundups side by side, and the thing that finally crystallized for me is that the formats I have been writing about all month have actually forked into at least seven different ranking frameworks, each committed to one product category, and none of the axes translate across categories. The coding picking roundups print a five-axis decimal scorecard (CodeBuddy 9.6, Cody 8.2, Ghostwriter 8.0, Codeium 7.8) or an S/A/B/D tier letter (Cursor at S档, Claude Code and Codex at A档, Rork at B档). The novel-writing piece I read this morning uses a four-axis framework (上手体验, 网文场景适配度, 数据安全感, 生产效率) to rank 蛙蛙写作, ChatGPT, DeepSeek, Kimi, and 作家助手. The broad-market 主流AI软件盘点 piece enumerates ChatGPT, Claude, DeepSeek, 通义千问 under 通用助手 with no ranking. The pricing guide prints dollars per month. The October trending recap prints GitHub stars per month. The forecast column prints a 44-dollar ChatGPT Plus projection by 2029. Each format is internally coherent and completely incompatible with the next, and the engineer trying to assemble a stack across categories has no translation table.&lt;/p&gt;

&lt;p&gt;The piece that pushed me over the edge was the novel-writing review, because it is a four-axis scorecard applied to five tools I had never seen compared under that framework, and the axes do not map onto anything the coding roundups use. 蛙蛙写作 ranks high on 网文场景适配度 but the piece never asks about IDE integration. Kimi ranks high on Chinese-language long-document processing, which is a column the picking roundups never built. 作家助手 ranks high on 码字场景效率 but 码字场景 is not a row in any coding scorecard I have read. To be fair I am taking the exact decimal scores with a grain of salt because the test corpus is never disclosed, but the structural tell is what has been rattling around in my head all morning. All seven frameworks have converged in the same search results page without anyone publishing the conversion table.&lt;/p&gt;

&lt;p&gt;The meta-pattern I want to call out is that the late-2025 roundups have not just forked into multiple formats, they have forked into format-category pairs where each format is glued to one category, and the cross-category stack work the engineer is doing has become a seven-framework translation job nobody signed up for. When I compare ChatGPT and Claude and Gemini for general-assistant work the broad-market piece gives me a vendor list and no rank. When I decide whether to add 蛙蛙写作 to my novel-writing pipeline alongside my Kimi habit the four-axis novel framework answers that question, but neither the coding scorecard nor the broad-market vendor list nor the dollar-per-month pricing column has a row for it. Honestly I am a little skeptical of any 2026 roundup workflow that pulls a single piece off the search results page and trusts it across categories, because each framework was written for one job and the cross-job integration is invisible.&lt;/p&gt;

&lt;p&gt;The practical takeaway I want to put down is that the late-2025 to early-2026 roundups are still useful for the within-framework anchor and not useful for the cross-framework re-rank most engineers are quietly trying to do this quarter. They name the per-category shortlist well, because the coding scorecard did name CodeBuddy and Cody and Ghostwriter and Codeium, the novel-writing piece did name 蛙蛙写作 and ChatGPT and DeepSeek and Kimi and 作家助手, and the broad-market piece did name ChatGPT and Claude and DeepSeek and 通义千问. They are not good at the cross-framework re-rank, because the engineer trying to decide whether to add a novel-writing tool to a coding stack on the strength of a 网文场景适配度 score has to discover that 网文场景适配度 has no published conversion to agent autonomy or IDE integration or cost, and that conversion is in none of the roundups. The seven frameworks are seven dialects of the same problem, and the dictionary between them does not exist yet.&lt;/p&gt;

&lt;p&gt;I will reassess in three months. The last time I said that I was mostly on Cursor and Claude Code for coding and ChatGPT for general assistant work, which is still roughly where I land, except that I now read every Juejin roundup as a category-specific dialect and treat any single post as one slice of a seven-slice workflow. What has changed is that I now check which framework a roundup is written in before I trust the verdict, and when the framework does not match the job I am actually trying to do I treat the rank as category trivia and reach for a different post. Give it six months and I expect either the roundups to publish an explicit axis-translation table at the top of the search results page, or the seven frameworks to harden into a permanent seven-dialect workflow the reader has to bridge by hand, and whichever moves first will tell me whether the format has finally noticed that the cross-category stack work is the engineering decision and the within-category rank is the supporting evidence.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cursor</category>
      <category>claude</category>
    </item>
  </channel>
</rss>
