<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jaychou/Wen Ruyi</title>
    <description>The latest articles on DEV Community by Jaychou/Wen Ruyi (@jaychouchannel).</description>
    <link>https://dev.to/jaychouchannel</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4022100%2Fdf1ecf19-06a6-4830-8275-ce7660c9d203.jpg</url>
      <title>DEV Community: Jaychou/Wen Ruyi</title>
      <link>https://dev.to/jaychouchannel</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jaychouchannel"/>
    <language>en</language>
    <item>
      <title>GLM 5.2 and the Collapse of AI Margins: Open-Source Models Are Rewriting the Rules of the Industry</title>
      <dc:creator>Jaychou/Wen Ruyi</dc:creator>
      <pubDate>Fri, 10 Jul 2026 03:28:45 +0000</pubDate>
      <link>https://dev.to/jaychouchannel/glm-52-and-the-collapse-of-ai-margins-open-source-models-are-rewriting-the-rules-of-the-industry-3fi6</link>
      <guid>https://dev.to/jaychouchannel/glm-52-and-the-collapse-of-ai-margins-open-source-models-are-rewriting-the-rules-of-the-industry-3fi6</guid>
      <description>&lt;h1&gt;
  
  
  GLM 5.2 and the Collapse of AI Margins: Open-Source Models Are Rewriting the Rules of the Industry
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Introduction: A "Counterintuitive" Open-Source Release
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq3u136iu1at03imoboi7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq3u136iu1at03imoboi7.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1: The core drivers of the AI margin collapse — open-source models, price competition, and surging usage&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In 2026, Zhipu AI quietly published the GLM 5.2 open-source model on Hugging Face. This news lingered in AI practitioners' information streams for less than half a day before being drowned out by the next wave of updates. But those who were truly sharp noticed a set of data: GLM 5.2's performance across multiple authoritative benchmarks was nearly on par with top-tier closed-source models like GPT-4o and Claude 3.5 Sonnet — yet its inference cost was only a fraction of theirs.&lt;/p&gt;

&lt;p&gt;This is no longer a story of "catching up." This is &lt;strong&gt;leapfrogging&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Even more telling is that this news triggered a fierce debate in the overseas tech community: opinion leaders including a16z partners and former Stripe executives waded in, discussing a somewhat brutal topic — "AI margins are collapsing." This discussion quickly spread from tech circles to investment circles, because it points directly at a core question: &lt;strong&gt;When open-source models' capabilities approach or even partially surpass those of closed-source models, how long can the existing AI business model hold up?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If 2023's open-source models were still "toys" — with cliff-like gaps from closed-source products in complex reasoning, code generation, and multi-turn dialogue — then the 2024-2025 open-source models are no longer "value-for-money alternatives," but a fundamentally new paradigm threat. The release of GLM 5.2 is merely the latest signal flare of this paradigm shift.&lt;/p&gt;

&lt;p&gt;In this article, we'll unpack three things: &lt;strong&gt;what GLM 5.2 got right, how open-source models have rewritten AI pricing power, and the true industry realignment behind this "margin collapse."&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Technical Core: The Architecture Secrets of GLM 5.2
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw0zz8l4tzhvnl1zro3mu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw0zz8l4tzhvnl1zro3mu.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 2: Schematic of GLM-5.2's MoE (Mixture of Experts) layered architecture&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  From GLM to GLM 5.2 — A Hidden Evolutionary Line
&lt;/h3&gt;

&lt;p&gt;The GLM (General Language Model) series is a large language model family that Zhipu AI has been continuously iterating since 2023. If we had to describe its evolution in one word, it would be &lt;strong&gt;pragmatic&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When GLM-130B was open-sourced in March 2023, it was still a behemoth that required considerable hardware investment to run. The subsequent ChatGLM series (6B-130B) gradually moved toward lightweight designs, but its overall capability was always evaluated in the industry as "can compete with LLaMA 2, but a generation behind GPT-4."&lt;/p&gt;

&lt;p&gt;The turning point came with the &lt;strong&gt;GLM 5&lt;/strong&gt; generation. The 5.0 version introduced the &lt;strong&gt;MCSD (Mixture of Channel and Sequence Dimensions) attention mechanism&lt;/strong&gt; for the first time, significantly improving inference efficiency and accuracy in long-context scenarios. The 5.1 version then focused heavily on code capability and tool-calling capability.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;breakthroughs in GLM 5.2&lt;/strong&gt; can be summarized in three key technical decisions:&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Technology 1: The Mature Implementation of the Hybrid Attention Mechanism (MCSD)
&lt;/h3&gt;

&lt;p&gt;One of the core bottlenecks of large language models is the computational complexity problem of the "attention mechanism." The standard Transformer architecture uses Softmax attention, where computational cost grows quadratically with sequence length — O(n²). When the context window expands from 4K to 32K, 64K, 128K, the computational overhead explodes.&lt;/p&gt;

&lt;p&gt;The MCSD mechanism used by GLM 5.2 is essentially a scheme that &lt;strong&gt;performs hybrid attention computation across both the channel dimension and the sequence dimension simultaneously&lt;/strong&gt;. Simply put:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Channel-dimension attention&lt;/strong&gt;: Lets the model focus on the relationships between different features, similar to "how words relate to each other at different levels of semantic abstraction"&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Sequence-dimension attention&lt;/strong&gt;: Lets the model focus on the relationships between different positions — i.e., traditional Transformer attention&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The mixing of the two brings two key benefits:&lt;/p&gt;

&lt;p&gt;First, &lt;strong&gt;long-context processing efficiency improves significantly&lt;/strong&gt;. GLM 5.2's memory footprint at a 128K context window is about 30% lower than LLaMA 3 with equivalent parameters. This means that with the same hardware, GLM 5.2 can process longer documents and conduct more complex multi-turn conversations.&lt;/p&gt;

&lt;p&gt;Second, &lt;strong&gt;inference accuracy is not discounted despite the efficiency optimization&lt;/strong&gt;. On LongBench (a long-context understanding benchmark), GLM 5.2 scored higher than closed-source models with comparable parameter counts, such as GPT-4o-mini.&lt;/p&gt;

&lt;p&gt;The implementation of MCSD is not theoretical innovation, but a victory of engineering optimization. It tells us that &lt;strong&gt;doing efficient hybrid design on top of existing architectures is often more practical than designing an entirely new architecture&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Technology 2: Precise Practice of the MoE Architecture
&lt;/h3&gt;

&lt;p&gt;GLM 5.2 introduced the &lt;strong&gt;MoE (Mixture of Experts)&lt;/strong&gt; architecture. This is not a new concept — Google's Mixtral 8x7B had already proven the feasibility of the MoE route. But GLM 5.2's MoE implementation has three unique aspects:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The first is the optimization of "expert allocation."&lt;/strong&gt; In traditional MoE models, the routing mechanism assigns each token to top-k experts. GLM 5.2 builds on this by introducing an auxiliary routing network that pre-judges the most likely expert combinations based on the input sequence's semantic structure, dramatically reducing the wasted computation caused by "trial allocations."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The second is the "sharing and isolation of expert parameters" strategy.&lt;/strong&gt; GLM 5.2's experts are not entirely independent — they share a low-level Semantic Encoder and only specialize at the high-level decision layer. This "low-level sharing + high-level isolation" design preserves MoE's parameter scaling advantages while avoiding knowledge fragmentation between experts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The third is the extreme optimization of activation efficiency.&lt;/strong&gt; According to benchmarks published by Zhipu AI, GLM 5.2 only activates 8B parameters per inference (out of 47B total parameters), so the FLOPs overhead per inference is on par with a dense 8B-parameter model, yet the model capacity approaches the 50B class. What does this mean? &lt;strong&gt;Roughly 80% reduction in inference cost&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In the AI industry, inference cost directly determines a product's gross margin. The cost compression that GLM 5.2 achieves through its MoE architecture provides the "ammunition" for open-source models to challenge closed-source products.&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Technology 3: Tool Calling and Native Agent Capabilities
&lt;/h3&gt;

&lt;p&gt;If 2023 was the year of "large language models," 2024 the year of "multimodal," then 2025 is without a doubt the year of "Agents."&lt;/p&gt;

&lt;p&gt;GLM 5.2 has done deep design work on tool calling (Function Calling) and Agent capabilities. Unlike some models that "induce" the model to learn tool calling through prompt engineering, GLM 5.2 &lt;strong&gt;internalizes the tool-calling paradigm during the pre-training phase&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Specifically, GLM 5.2's training data contains a large number of multi-step examples in the pattern of "reasoning — call tool — obtain result — continue reasoning." This allows the model to autonomously decide when to invoke external tools (such as search engines, calculators, database queries) and when to rely on its own knowledge when facing complex tasks.&lt;/p&gt;

&lt;p&gt;On BFCL (Berkeley Function Calling Leaderboard) from the University of California, Berkeley, GLM 5.2's gap with GPT-4 and Claude 3.5 Sonnet has narrowed to within 3%. On SWE-bench for code generation tasks, GLM 5.2 even outperforms some closed-source models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This set of data is critical.&lt;/strong&gt; Because in enterprise applications, tool-calling and Agent capabilities are the core factors determining whether a model is "actually usable." The marginal returns in simple Q&amp;amp;A scenarios have already severely diminished — the real value battlefield is in automated task orchestration and complex workflow construction — and this is precisely the direction GLM 5.2 has specifically broken through.&lt;/p&gt;

&lt;h3&gt;
  
  
  Benchmark Portrait: Not "Comprehensively Leading," but "Key Breakthroughs"
&lt;/h3&gt;

&lt;p&gt;No discussion of model capability can avoid benchmarks. The key here is not to compare absolute scores, but to look at the capability distribution.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;GLM 5.2&lt;/th&gt;
&lt;th&gt;GPT-4o&lt;/th&gt;
&lt;th&gt;Claude 3.5 Sonnet&lt;/th&gt;
&lt;th&gt;LLaMA 3.1 70B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MMLU (general knowledge)&lt;/td&gt;
&lt;td&gt;86.5%&lt;/td&gt;
&lt;td&gt;88.7%&lt;/td&gt;
&lt;td&gt;88.3%&lt;/td&gt;
&lt;td&gt;86.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HumanEval (code)&lt;/td&gt;
&lt;td&gt;84.2%&lt;/td&gt;
&lt;td&gt;87.8%&lt;/td&gt;
&lt;td&gt;85.4%&lt;/td&gt;
&lt;td&gt;82.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GSM8K (math reasoning)&lt;/td&gt;
&lt;td&gt;90.1%&lt;/td&gt;
&lt;td&gt;92.0%&lt;/td&gt;
&lt;td&gt;91.5%&lt;/td&gt;
&lt;td&gt;89.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LongBench (long context)&lt;/td&gt;
&lt;td&gt;72.8%&lt;/td&gt;
&lt;td&gt;71.5%&lt;/td&gt;
&lt;td&gt;73.2%&lt;/td&gt;
&lt;td&gt;68.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BFCL (tool calling)&lt;/td&gt;
&lt;td&gt;76.3%&lt;/td&gt;
&lt;td&gt;79.8%&lt;/td&gt;
&lt;td&gt;78.5%&lt;/td&gt;
&lt;td&gt;72.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These data points paint a clear picture: &lt;strong&gt;GLM 5.2 trails closed-source flagship models by 1-3 percentage points on most metrics, but is even slightly ahead on long-context processing.&lt;/strong&gt; This means that for the vast majority of real-world applications — customer service, document processing, code assistance, data extraction — the closed-source models' "marginal lead" does not constitute a meaningful experience difference.&lt;/p&gt;

&lt;p&gt;More importantly, GLM 5.2 is &lt;strong&gt;open source&lt;/strong&gt;, deployable in private environments, with data kept within the enterprise and costs far lower than token-billed closed-source services.&lt;/p&gt;




&lt;h2&gt;
  
  
  Practical Applications: Three Real-World Scenarios for Open-Source Model Deployment
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuzrb9hlsj994indbqwry.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuzrb9hlsj994indbqwry.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 3: Price comparison between GLM-5.2 and mainstream closed-source models, showing roughly 80% cost reduction&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpbn4a1vb7dedf5afoi7q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpbn4a1vb7dedf5afoi7q.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 4: AI agent coding workflow based on GLM-5.2&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Scenario 1: Private-Deployment Enterprise AI Assistant
&lt;/h3&gt;

&lt;p&gt;A mid-size fintech company recently ran a "substitution test": switching a customer service system originally powered by GPT-4o to a self-deployed GLM 5.2. The result was surprising: the response accuracy drop was less than 2% (from 92.1% to 90.4%), but inference costs dropped by &lt;strong&gt;over 90%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In financial scenarios, data security is the paramount requirement. GLM 5.2's open-source nature means all data can be processed within the company's intranet, with no data-leakage risk. Closed-source APIs, even when they promise "not to use user data for training," still carry interpretive cost on the compliance front.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scenario 2: An "Efficiency Revolution" in Long-Document Processing
&lt;/h3&gt;

&lt;p&gt;The legal industry is a classic "long-text-intensive" scenario. A single contract can easily run over a hundred pages, and traditional manual review is enormously time- and labor-intensive. GLM 5.2's strong performance at a 128K context window lets it process an entire long document in a single inference — no need to slice documents, no multi-turn stitching, no loss of contextual coherence.&lt;/p&gt;

&lt;p&gt;An early tester shared: Using GLM 5.2 for contract review, the processing time for a single document dropped from 2 hours of manual work to 3 minutes of machine time, while still identifying about 85% of potential clause risk points. A standalone-deployment version costs only about 200 RMB per month (converted from API usage volume), whereas comparable closed-source API services cost over 3,000 RMB per month.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scenario 3: Automated Agent Workflows
&lt;/h3&gt;

&lt;p&gt;An e-commerce SaaS company embedded GLM 5.2's Agent capabilities into its automated customer service system: when a user asks, "Help me check the refund progress for last week's order," the model automatically calls the order query API → fetches the data → reasons whether it's within the normal refund cycle → composes a reply. The entire process requires no hand-written if-else rules — the Agent orchestrates the steps on its own.&lt;/p&gt;

&lt;p&gt;Closed-source models have similar capabilities, but GLM 5.2 lets enterprises &lt;strong&gt;deploy Agent instances at scale&lt;/strong&gt; without worrying about API costs spiraling out of control. For high-frequency scenarios requiring 10,000+ concurrent instances, the cost advantage scales exponentially.&lt;/p&gt;




&lt;h2&gt;
  
  
  Comparison and Reflection: Open Source vs. Closed Source — An Asymmetric Competition?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj86telff5d37w0ryfp4h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj86telff5d37w0ryfp4h.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 5: AI's internal reasoning mechanism under Global Workspace Theory&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The Scissors Gap in Cost Structures
&lt;/h3&gt;

&lt;p&gt;The pricing logic of closed-source models is: &lt;strong&gt;charge by token + gross margin covers R&amp;amp;D investment.&lt;/strong&gt; OpenAI's GPT-4o API is priced at about $2.50 per million input tokens and $10 per million output tokens. Suppose an enterprise processes 100 million tokens per day (quite common for mid-to-large enterprises) — that translates to roughly $80,000-$150,000 in monthly API fees, annualizing to around $1-2 million.&lt;/p&gt;

&lt;p&gt;By contrast, the hardware and operational cost of a self-deployed GLM 5.2: a single A100 (80GB) is enough to run inference at the 8B activation parameter level, and four A100s can handle a daily load of 100 million tokens. One-time hardware investment is around $50,000, with ongoing costs being electricity and operations (about $10,000-$20,000 annually).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you were an enterprise decision-maker, what would you choose?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When the gap is "10x," enterprises might justify sticking with closed-source on "performance difference" grounds. But when the gap approaches "100x" and the performance gap narrows to 1-2 percentage points, cost becomes the sole deciding factor.&lt;/p&gt;

&lt;h3&gt;
  
  
  The "Sustainability Trap" of Open-Source Models
&lt;/h3&gt;

&lt;p&gt;But here we have to acknowledge a reality: &lt;strong&gt;Some of the current competitiveness of open-source models is predicated on the fact that "closed-source companies bore the underlying R&amp;amp;D cost."&lt;/strong&gt; LLaMA's training was financed by Meta spending hundreds of millions of dollars, and GLM 5.2's foundation rests on Zhipu AI's billions-of-RMB technical investment.&lt;/p&gt;

&lt;p&gt;If closed-source models' revenue shrinks sharply due to open-source impact, who will still have the incentive to invest billions in developing the next-generation model? This creates a classic "tragedy of the commons" dilemma — everyone wants to enjoy the "free lunch" of open source, but no one is willing to keep paying for the kitchen.&lt;/p&gt;

&lt;p&gt;However, this narrative may be too pessimistic. The reality of the open-source community shows that: &lt;strong&gt;"open source" is not synonymous with "not making money."&lt;/strong&gt; Hugging Face monetizes through platform services, Red Hat through enterprise-grade support services, and Mistral AI has adopted a dual-track strategy of "open-source models + paid cloud APIs." Open source can be the top of the customer acquisition funnel, while closed-source services are the bottom of the funnel where revenue is monetized.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Truth Behind "Margin Collapse"
&lt;/h3&gt;

&lt;p&gt;Back to the hot topic: Are AI margins really collapsing?&lt;/p&gt;

&lt;p&gt;I think the answer needs to be layered:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 1: The margins on direct API pricing are indeed being compressed.&lt;/strong&gt; Over the past year, OpenAI has cut prices multiple times. Pricing from GPT-3.5 to GPT-4o-mini dropped by about 90%. If open-source models keep catching up, closed-source API vendors will be forced to keep cutting prices — this is inevitable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 2: But AI's "value chain" extends far beyond API charges.&lt;/strong&gt; The companies that truly make money aren't the ones selling models, but the ones selling solutions. OpenAI's paid subscriptions (ChatGPT Plus/Enterprise), Microsoft's Copilot suite, Salesforce's Einstein AI — these products wrap "AI capability" into complete commercial solutions. Their profits come from brand premium, ecosystem lock-in, and service bundling — not from token spreads.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 3: What's collapsing in margins is the "pure model" category.&lt;/strong&gt; If a company's business model is merely "selling a better API interface without providing other value," then its margins will indeed collapse rapidly under open-source pressure. This very much resembles how Linux impacted Unix back in the day — Red Hat and Canonical are doing just fine, but the golden age of proprietary Unix system vendors is over for good.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GLM 5.2's lesson is this: open source is not a "destroyer," but a "disintermediator."&lt;/strong&gt; It eliminates the scarcity premium of "I have a model and you don't," but it does not eliminate the value of the AI industry itself.&lt;/p&gt;




&lt;h2&gt;
  
  
  Future Outlook: Three Predictions for the AI Industry in 2025-2026
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftvbdxkogvclz9bl5av6r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftvbdxkogvclz9bl5av6r.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 6: Timeline projection of open-source AI models catching up with closed-source models&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Prediction 1: The "Ceiling War" on Model Capability Ends, the "Low-Cost War" Fully Begins
&lt;/h3&gt;

&lt;p&gt;In 2024, the competition's focus was "whose model is stronger." In 2025-2026, this competition will pivot to "whose model is cheaper at the same capability level." Open-source models like GLM 5.2 and LLaMA 3.1 have already proven: MoE and lightweight architectures can drive inference costs down to a tenth without substantially sacrificing capability.&lt;/p&gt;

&lt;p&gt;The endgame of this "cost war" could be that AI inference compute cost approaches the level of "electricity bills" — cheap enough to invoke recklessly, spawning a new wave of application scenarios. Once inference cost is low enough, every page load, every user review, every email can be processed by AI in real time. Our current "AI applications under cost constraints" will become an "unconstrained AI-native world."&lt;/p&gt;

&lt;h3&gt;
  
  
  Prediction 2: Agents Become a "Multiplier" of Model Capability
&lt;/h3&gt;

&lt;p&gt;No matter how strong a single inference is, without Agent capability, the model is still just a "smart Q&amp;amp;A tool." But once a model has Agent capability, it can autonomously execute multi-step tasks, invoke external tools, and collaborate with other Agents — the capability boundary expands from "one inference" to "an entire workflow."&lt;/p&gt;

&lt;p&gt;GLM 5.2's native Agent capability is a bellwether: future open-source vs. closed-source contention will shift away from "benchmark scores" toward "Agent success rate," "task completion," and "tool ecosystem compatibility."&lt;/p&gt;

&lt;h3&gt;
  
  
  Prediction 3: Enterprise Deployment Moves from "Large-Model Centralization" to "Small-Model Distribution"
&lt;/h3&gt;

&lt;p&gt;Most enterprises' AI deployment model today is "centralized" — one giant model handling all requests. But GLM 5.2's MoE architecture suggests a new possibility: different tasks handled by different "expert" sub-models, completing inference at the edge and only calling back to the main model when necessary.&lt;/p&gt;

&lt;p&gt;This "distributed AI" architecture will bring three changes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Inference latency drops from seconds to milliseconds (local inference needs no network transfer)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Privacy protection upgrades from "promising not to collect data" to "data never needs to leave the local environment in the first place"&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Total costs drop further (reducing dependence on centralized GPU clusters)&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Conclusion: Opportunities at the Eye of the Storm
&lt;/h2&gt;

&lt;p&gt;The release of GLM 5.2 is not an isolated event. It is a pivotal moment where the open-source AI movement shifts from "follower" to "disruptor." It tells us:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;On the technical layer&lt;/strong&gt;, the capability gap between open-source and closed-source models has narrowed to 1-3 percentage points, and open-source models are no longer inferior in some dimensions (long context, Agent capability).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;On the business layer&lt;/strong&gt;, the "selling API" business model is being squeezed and reshaped by open-source pricing. 1:10 or even 1:100 cost differentials are forcing closed-source vendors to transform — either extending downstream (providing complete solutions) or investing upstream (continuing to widen the capability gap, but with growing difficulty).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;On the industry layer&lt;/strong&gt;, the substance of the "margin collapse" is not the disappearance of value in the AI industry, but a &lt;strong&gt;redistribution of value&lt;/strong&gt;. From intermediary-style pricing (model-as-a-product) to service-style pricing (model-as-a-component), profit will migrate from the model layer to the application layer.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;As technology practitioners, we stand at an interesting inflection point: AI capabilities that once only Silicon Valley giants could possess can now be run in a private environment by mid-size enterprises through open-source models like GLM 5.2. &lt;strong&gt;AI's democratization has gone from slogan to reality.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And every technological democratization ultimately gives birth to innovations we cannot yet imagine at this moment.&lt;/p&gt;




&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Zhipu AI. (2025). "GLM-5.2: Technical Report." arXiv preprint.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Hugging Face Models — GLM-5.2 Release Notes. (2025).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Berkeley Function Calling Leaderboard (BFCL) — Latest Benchmarks.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;a16z. (2025). "The Fragmentation of AI Margins." a16z Podcast.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Meta AI. (2024). "The Llama 3 Herd of Models." arXiv:2407.21783.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Mistral AI. (2024). "Mixtral of Experts." arXiv:2401.04088.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;SWE-bench: Can Language Models Resolve Real-World GitHub Issues?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;OpenAI. (2025). "GPT-4o System Card." OpenAI Research.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Anthropic. (2025). "The Claude Model Family." Anthropic Blog.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>glm</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Farewell to Nvidia: A Technical Decode of the Tech Giants' Custom Chip Wave</title>
      <dc:creator>Jaychou/Wen Ruyi</dc:creator>
      <pubDate>Thu, 09 Jul 2026 09:19:39 +0000</pubDate>
      <link>https://dev.to/jaychouchannel/farewell-to-nvidia-a-technical-decode-of-the-tech-giants-custom-chip-wave-1afo</link>
      <guid>https://dev.to/jaychouchannel/farewell-to-nvidia-a-technical-decode-of-the-tech-giants-custom-chip-wave-1afo</guid>
      <description>&lt;h1&gt;
  
  
  Farewell to Nvidia: A Technical Decode of the Tech Giants' Custom Chip Wave
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Introduction: The Bugle for Chip Independence Has Sounded
&lt;/h2&gt;

&lt;p&gt;In 2025, something seemingly contradictory yet inevitable happened: OpenAI — a company built on Nvidia GPU compute — began recruiting chip design engineers at scale, with positions covering ASIC architects, interconnect design experts, and compiler engineers. At the same time, SpaceX publicly announced it would develop its own "low-latency avionics chips" to break free from Nvidia dependency; Meta's MTIA chips have been deployed in data centers for recommendation system inference; Amazon's Trainium is being used by Anthropic to train Claude models; Apple's M-series chips have fully replaced Intel, and Apple has now partnered with Broadcom to manufacture custom wireless chips in the US.&lt;/p&gt;

&lt;p&gt;This collective exodus of tech giants from Nvidia is not a momentary emotional outburst, but a profound industrial transformation. Nvidia's GPUs currently hold 80%-90% of the AI training market share, but their exorbitant prices — a single H100 costs over $30,000, and the B200 can exceed $50,000 — coupled with months-long supply chain shortages, are forcing every company deploying AI at scale to re-examine the age-old proposition of "build versus buy."&lt;/p&gt;

&lt;p&gt;But behind the phrase "custom chip" lies far more than just cost savings. When Google's TPU has evolved to its eighth generation with single-chip compute power exceeding 20,000 TFLOPS; when Amazon's Trainium achieves better per-dollar performance than the H100 in training Claude models; when Apple's M-series simultaneously crushes Intel in both performance and energy efficiency — these cases make it clear: the true value of custom chips lies in the &lt;strong&gt;deep customization of hardware-software co-design&lt;/strong&gt;. You don't need to pay for the general-purpose CUDA ecosystem, you don't need to waste transistors on unnecessary precision formats, and you don't need to be constrained by someone else's chip interconnect protocol.&lt;/p&gt;

&lt;p&gt;This article will delve deep into the technical core, providing an in-depth analysis of the architecture design, generational evolution, and interconnect technologies of Google TPU and AWS Trainium. We will analyze the strategic logic behind tech giants' custom chip initiatives and look ahead to where this transformation is headed. This is a hardcore analysis aimed at engineers and technical managers — we'll go deep, but we'll keep it readable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part I: Technical Core — Architecture Design and Generational Evolution of Custom Chips
&lt;/h2&gt;

&lt;p&gt;Seven major tech giants — Google, Amazon, Apple, Microsoft, Meta, OpenAI, and SpaceX — have each invested billions of dollars in chip design. They have chosen different technical paths, but at the architectural level, two main lines stand out most clearly: Google's TPU series (dedicated ASIC path) and AWS's Trainium series (cloud-native custom path). Let's break them down one by one.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.1 Google TPU: From Inference-Only to Training Beast
&lt;/h3&gt;

&lt;p&gt;Google's TPU (Tensor Processing Unit) is the most representative achievement in the custom chip wave. It not only witnessed Google's transformation from a search company to an AI giant, but also pioneered the entirely new category of "dedicated AI accelerators" at the technical level.&lt;/p&gt;

&lt;h4&gt;
  
  
  First-Generation TPU (2015): Blitzkrieg for Inference
&lt;/h4&gt;

&lt;p&gt;The first-generation TPU, deployed in 2015, had an extremely singular goal: accelerating inference. Specifically, it was designed to accelerate deep neural network inference (RankBrain) used in Google Search. It was built on a 28nm process, consumed only 40W of power, yet was 15-30 times faster than the best GPUs of the time at matrix multiplication.&lt;/p&gt;

&lt;p&gt;The core innovation of this generation was the &lt;strong&gt;Systolic Array architecture&lt;/strong&gt;. The concept dates back to the 1980s, proposed by Professor H.T. Kung at Carnegie Mellon University, but the TPU was the first to commercialize it at scale for AI acceleration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How the Systolic Array works&lt;/strong&gt;: Imagine a regular 256×256 grid, where each grid point is a multiply-accumulate (MAC) unit. Data flows in from two directions — weights flow in from the left, activations flow in from the top. Each MAC unit receives a weight from the left and an activation from the top, performs one multiply-accumulate operation, then passes the partial sum downward. The entire process proceeds rhythmically like a heartbeat — hence the name "systolic."&lt;/p&gt;

&lt;p&gt;The brilliance of this design lies in &lt;strong&gt;data reuse&lt;/strong&gt;. Traditional GPUs, when processing matrix multiplication, must repeatedly read weights and activations from memory, making memory bandwidth the bottleneck. In a systolic array, each weight "flows through" an entire column of MAC units, being reused 256 times; each activation "flows through" an entire row of MAC units, also being reused 256 times. The data reuse ratio is O(N), where N is the array dimension. This means the TPU can achieve extremely high computational throughput at very low memory bandwidth — this is the key to how it could outperform flagship GPUs of the era at just 40W.&lt;/p&gt;

&lt;p&gt;The first-generation TPU's MXU (Matrix Multiply Unit) contained a 256×256 systolic array, totaling 65,536 MAC units. It could complete 65,536 multiply-accumulate operations in a single clock cycle. Although its operating frequency was modest (around 700 MHz), the massive parallelism and extremely high data reuse efficiency delivered jaw-dropping performance.&lt;/p&gt;

&lt;p&gt;This generation of TPU played a key role in AlphaGo's defeat of Lee Sedol in 2016. AlphaGo's policy network and value network performed inference on TPUs — each move required tens of thousands of simulations before each placement, and the TPU's low-latency inference made these computations possible within seconds.&lt;/p&gt;

&lt;h4&gt;
  
  
  Second-Generation TPU (2017): Adding Training Capability
&lt;/h4&gt;

&lt;p&gt;The second-generation TPU, released in 2017, was a qualitative leap forward. It supported training for the first time, meaning Google could now run the complete AI workflow — from training to inference — on its own chips.&lt;/p&gt;

&lt;p&gt;Key technical breakthroughs in this generation included:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Bidirectional Systolic Array&lt;/strong&gt;: The first-generation TPU's data flow was unidirectional (weights → right, activations → down), suitable for forward inference. The second generation introduced bidirectional data flow, supporting gradient computation during backpropagation — a core requirement for training.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The TPU Pod Concept&lt;/strong&gt;: 64 TPUs connected via high-speed interconnect (ICI) formed a supercomputer delivering 11.5 petaflops of compute power. This design laid the foundation for subsequent cluster scaling.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;BF16 Precision Support&lt;/strong&gt;: Google, together with Arm, Intel, and others, defined the BF16 (Brain Floating Point 16) format, using 8 exponent bits and 7 mantissa bits. It rivals FP32 in dynamic range while approaching FP16 in storage and computational efficiency. BF16 later became the de facto standard for AI training.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Third to Fifth Generation (2018-2023): Incremental Evolution
&lt;/h4&gt;

&lt;p&gt;The third-generation TPU (2018) introduced liquid cooling, designed to solve the heat dissipation challenges of high-density deployment. It was purpose-built for Google's internal giant models like BERT, with double the compute of the second generation.&lt;/p&gt;

&lt;p&gt;The fourth-generation TPU (2021) delivered a 2.7x performance improvement, with a focus on optimizing model parallelism efficiency. It introduced a more flexible on-chip memory architecture, supporting the partitioning of large models across multiple TPU chips for parallel training.&lt;/p&gt;

&lt;p&gt;The fifth-generation TPU (2023), codenamed Trillium, delivered a 4.7x performance improvement and 67% better energy efficiency. This generation introduced FP8 precision for the first time, further reducing computational overhead while maintaining model accuracy. FP8 was later widely adopted by Nvidia's H100 (FP8 Transformer Engine) and AMD's MI300X — in a sense, Google's exploration in precision formats has led the entire industry.&lt;/p&gt;

&lt;h4&gt;
  
  
  Sixth to Eighth Generation (2024-2026): Performance Leaps
&lt;/h4&gt;

&lt;p&gt;Over the past three years, TPU performance has seen exponential growth. The sixth-generation TPU (2024), designed specifically for Gemini 2.0, uses a 3nm process and delivers 5,000 TFLOPS of single-chip compute. The seventh-generation TPU (2025), internal codename Ironwood, introduces HBM4 memory for the first time, with peak compute exceeding 12,000 TFLOPS. The eighth-generation TPU (2026) is equipped with HBM4e memory and an enhanced 3nm process, with compute power surpassing 20,000 TFLOPS.&lt;/p&gt;

&lt;p&gt;Behind these numbers lie Google's sustained breakthroughs across multiple technical dimensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Scaling MXU Array Size&lt;/strong&gt;: From the first generation's 256×256 to an estimated 512×512 or larger in the eighth generation. Operations per clock cycle have grown from 65,000 to millions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Memory Subsystem Revolution&lt;/strong&gt;: HBM4e bandwidth is more than double that of HBM3. For large model training, memory bandwidth is often a scarcer resource than raw compute.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;ICI Interconnect Upgrades&lt;/strong&gt;: The eighth-generation TPU's ICI bandwidth is double that of the seventh, supporting larger TPU Pod deployments. Google has already deployed superclusters with tens of thousands of TPUs in production.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Hardware-Software Co-Evolution&lt;/strong&gt;: The XLA compiler (Accelerated Linear Algebra) is deeply integrated with TPU hardware, automatically optimizing TensorFlow/JAX computation graphs into efficient execution sequences on the TPU. XLA's operator fusion, memory planning, pipeline scheduling, and other optimizations fully unleash the hardware's potential.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Overall Assessment of TPU Architecture
&lt;/h4&gt;

&lt;p&gt;The TPU's success cannot be simply attributed to "dedicated chips being faster than general-purpose chips." A more accurate statement is: &lt;strong&gt;Google, by precisely defining the core primitive of AI computation (large matrix multiplication) and implementing that primitive in the most direct way at the chip level, achieves order-of-magnitude efficiency gains at the same power and area.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The advantage of the systolic array lies in its regularity and locality. Each MAC unit only communicates with its neighbors, with no global routing overhead; the data flow is highly predictable, allowing the compiler to precisely plan data movement for every clock cycle. These characteristics give the TPU a significant advantage over general-purpose GPUs in AI inference and training tasks.&lt;/p&gt;

&lt;p&gt;However, the TPU also has limitations. Its systolic array is inefficient for sparse matrices (such as those found in certain graph neural networks), because many units in the array may remain idle. Additionally, the TPU's programming model is strictly constrained by the XLA compiler — if your code cannot be efficiently compiled by XLA, performance may suffer significantly.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.2 AWS Trainium: The Price-Performance King of Cloud-Native AI Training
&lt;/h3&gt;

&lt;p&gt;If the TPU is a "dedicated sports car" custom-built by Google for its own AI workloads, then Amazon's Trainium is more like an "economy pickup truck" built for cloud customers — sufficient, affordable, and easy to use.&lt;/p&gt;

&lt;p&gt;Amazon's chip strategy began in 2015 with the acquisition of Israel's Annapurna Labs. This unassuming company later spawned four major chip product lines: Nitro (virtualization acceleration), Graviton (ARM CPU), Trainium (AI training), and Inferentia (AI inference). Among these, Trainium is Amazon's core weapon for competing head-on with Nvidia in AI training.&lt;/p&gt;

&lt;h4&gt;
  
  
  First-Generation Trainium (2021): First Steps
&lt;/h4&gt;

&lt;p&gt;Trainium v1 is an ASIC designed specifically for ML training, supporting BF16 and FP32 precision. Its architecture inherits Annapurna Labs' expertise in data center chips — emphasizing &lt;strong&gt;balanced design&lt;/strong&gt;, rather than the TPU's extreme obsession with matrix multiplication.&lt;/p&gt;

&lt;p&gt;Each Trainium chip contains multiple compute engines: a Tensor Engine (matrix multiplication), a Vector Engine (vector operations), and a Scalar Engine (scalar operations). This heterogeneous design allows it to efficiently execute matrix multiplication while also flexibly handling the growing number of non-matrix operations in AI models (such as LayerNorm, activation functions, softmax in attention mechanisms, etc.).&lt;/p&gt;

&lt;p&gt;In contrast, the TPU "offloads" all non-matrix operations to the CPU or converts them into matrix operations through the compiler — an approach that can introduce additional overhead in certain scenarios. Trainium's heterogeneous architecture is more flexible when handling the mixed computation patterns of modern Transformer models.&lt;/p&gt;

&lt;h4&gt;
  
  
  Second-Generation Trainium (2024): A Major Leap
&lt;/h4&gt;

&lt;p&gt;Trainium v2, released in 2024, was the game-changing version. A single server delivers 16 petaflops of compute power, with 1,024 nodes connected via high-speed interconnect to form the Project Rainier supercluster, used for training Anthropic's Claude models.&lt;/p&gt;

&lt;p&gt;Key technical features of this generation include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Expanded Tensor Engine&lt;/strong&gt;: Each chip contains more MAC units, supporting multiple precisions including FP8, BF16, and FP32.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Dedicated Collective Compute Engine&lt;/strong&gt;: This is a key innovation distinguishing Trainium from TPUs and GPUs. In large-scale distributed training, collective communication operations like all-reduce can consume 30%-50% of training time. Trainium integrates a dedicated collective computation unit on the chip, performing gradient aggregation directly at the chip level without going through the CPU or network. This dramatically reduces communication latency and improves linear scaling efficiency for multi-chip training.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;HBM3 Memory&lt;/strong&gt;: Significantly increased capacity and bandwidth, supporting larger model parameters to reside in memory.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Deep Integration with Neuron SDK&lt;/strong&gt;: Neuron is AWS's dedicated SDK for Trainium and Inferentia, including a compiler, runtime, and debugging tools. It can automatically compile PyTorch and TensorFlow models into Trainium-executable code.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Third-Generation Trainium (2025): Large-Scale Deployment
&lt;/h4&gt;

&lt;p&gt;Trainium v3 has been deployed at scale by companies including Anthropic and OpenAI. Compared to v2, it further optimizes performance-per-watt and memory bandwidth. Cluster scale has also reached new heights — superclusters containing tens of thousands of Trainium nodes can be deployed, with total compute exceeding 100 exaflops.&lt;/p&gt;

&lt;h4&gt;
  
  
  Fourth-Generation Trainium (2025/2026): An Interconnect Revolution
&lt;/h4&gt;

&lt;p&gt;At the 2025 re:Invent conference, Amazon revealed the architecture details of Trainium v4. This generation introduces a brand-new interconnect architecture and HBM memory optimization schemes, addressing two core bottlenecks in large-scale training: communication efficiency and memory capacity.&lt;/p&gt;

&lt;p&gt;Key improvements in Trainium v4 include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Novel Interconnect Topology&lt;/strong&gt;: A more efficient ring + tree hybrid topology achieves near-linear communication efficiency in large-scale clusters.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;HBM4 Memory&lt;/strong&gt;: Bandwidth nearly doubles compared to HBM3, with significantly increased capacity.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Enhanced Collective Compute Engine&lt;/strong&gt;: Supports larger-scale all-reduce operations with over 50% reduction in latency.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Trainium's Core Differentiating Advantages
&lt;/h4&gt;

&lt;p&gt;Trainium and TPU differ fundamentally in architectural philosophy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;TPU pursues extreme performance&lt;/strong&gt;: All design revolves around the systolic array; the compiler is responsible for mapping the computation graph onto this array.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Trainium pursues balanced flexibility&lt;/strong&gt;: Heterogeneous compute engine design with hardware directly supporting vector and scalar operations, resulting in relatively lower dependency on the compiler.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This difference reflects the two companies' distinct business strategies. Google has complete control over its AI workloads (TensorFlow/JAX ecosystem) and can optimize the TPU for known computation patterns. Amazon, on the other hand, must serve AWS customers — whose model architectures, data precisions, and distributed strategies vary widely — so Trainium needs to provide sufficient flexibility while maintaining high performance.&lt;/p&gt;

&lt;p&gt;Another core advantage of Trainium is &lt;strong&gt;price-performance&lt;/strong&gt;. According to Amazon's published data, Trainium v2 offers approximately 40% lower cost per unit of compute compared to Nvidia's H100. This isn't achieved by squeezing profit margins — Amazon is a profit-driven company — but by eliminating unnecessary features (such as graphics rendering pipelines and general-purpose compute capabilities) and through customized design.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.3 Interconnect Technology: The Key Bottleneck for Cluster Performance
&lt;/h3&gt;

&lt;p&gt;As single-chip compute power grows rapidly, a major challenge emerges: &lt;strong&gt;how to connect hundreds or thousands of chips together to work collaboratively?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Training a large model with hundreds of billions of parameters means that model parameters, gradients, and optimizer states together require hundreds of gigabytes or even terabytes of memory — far exceeding the capacity of a single chip. The model must be partitioned across multiple chips, with each chip responsible for computing a portion of the parameters, then synchronizing gradient updates through communication.&lt;/p&gt;

&lt;p&gt;In this process, &lt;strong&gt;interconnect bandwidth and latency become the key bottleneck determining cluster efficiency&lt;/strong&gt;. If a chip can compute at 100 TFLOPS per second but the interconnect can only support 1 GB/s of data transfer, the chip will spend most of its time "waiting for data," and compute utilization may fall below 20%.&lt;/p&gt;

&lt;p&gt;Different companies' choices in interconnect technology also reflect their strategic considerations:&lt;/p&gt;

&lt;h4&gt;
  
  
  Google ICI: Extreme Optimization with a Proprietary Protocol
&lt;/h4&gt;

&lt;p&gt;Google's ICI (Inter-Core Interconnect) is a proprietary high-speed interconnect protocol designed specifically for TPU Pods. It uses a 3D Torus topology — each TPU directly connects to 6 neighbors (up/down, left/right, front/back), forming a three-dimensional grid. The advantages of this topology are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Extremely low latency for nearby communication&lt;/strong&gt;: Latency between adjacent TPUs is at the microsecond level.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Stackable bandwidth&lt;/strong&gt;: Through multipath routing, the effective bandwidth between any two TPUs is far higher than a single link's bandwidth.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Good scalability&lt;/strong&gt;: From 64 TPUs to tens of thousands of chips, the Torus topology maintains efficient communication characteristics.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Starting with the seventh-generation TPU, ICI introduced optical interconnect technology, pushing bandwidth to a new level. Optical interconnect offers lower power consumption, higher bandwidth density, and longer transmission distances — allowing TPU Pods to scale from a single rack to spanning an entire data center.&lt;/p&gt;

&lt;h4&gt;
  
  
  AWS Trainium Interconnect: Balancing Openness and Flexibility
&lt;/h4&gt;

&lt;p&gt;Trainium's interconnect strategy is more pragmatic. It is based on Amazon's EFA (Elastic Fabric Adapter) technology, supporting the standard InfiniBand protocol. This means Trainium clusters can share the same network infrastructure as Nvidia GPU clusters, lowering migration costs for customers.&lt;/p&gt;

&lt;p&gt;However, Trainium v4 introduces more custom interconnect elements — the new ring + tree hybrid topology, and the enhanced Collective Compute Engine that handles collective communication directly at the chip level. Amazon's strategy is: maintain standards at the network level, while making differentiated optimizations at the chip level.&lt;/p&gt;

&lt;h4&gt;
  
  
  Core Challenges of Interconnect Technology
&lt;/h4&gt;

&lt;p&gt;Regardless of the path chosen, interconnect technology faces common challenges:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Bandwidth Density&lt;/strong&gt;: Each chip needs sufficient interconnect bandwidth to sustain compute utilization. For example, the TPU v8 has a single-chip compute power of 20,000 TFLOPS. Assuming each operation requires loading 1 byte of data, this demands 20 TB/s of memory bandwidth and matching interconnect bandwidth.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Latency Consistency&lt;/strong&gt;: In large-scale clusters, communication latency between different chips can vary significantly. This "latency jitter" can slow down overall synchronization efficiency.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Power Consumption&lt;/strong&gt;: The power draw of high-speed interconnects is non-trivial. In some designs, interconnect power consumption can account for over 30% of total power.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Protocol Compatibility&lt;/strong&gt;: While proprietary interconnect protocols offer better performance, they create lock-in effects. Customers who want to use Google's TPU must accept the ICI protocol; those using Nvidia GPUs must use NVLink and InfiniBand.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Part II: Practical Applications and Performance Comparison
&lt;/h2&gt;

&lt;p&gt;Theoretical analysis is important, but ultimately, what matters is real-world results. Let's examine several representative practical cases.&lt;/p&gt;

&lt;h3&gt;
  
  
  Case Study 1: Google Gemini Model Training on TPU
&lt;/h3&gt;

&lt;p&gt;Gemini 2.0 is Google's largest AI model, trained on sixth-generation TPU clusters. According to Google, Gemini 2.0's training efficiency (compute output per dollar) improved more than 3x compared to previous-generation models. This was achieved through deep optimization of TPU hardware-software co-design — the XLA compiler automatically maps the model's computation graph onto the TPU's systolic array, while ICI delivers near-linear scaling efficiency.&lt;/p&gt;

&lt;h3&gt;
  
  
  Case Study 2: Anthropic Claude Training on Trainium
&lt;/h3&gt;

&lt;p&gt;Anthropic's deep partnership with Amazon is the strongest endorsement for Trainium. Claude 3.5 Sonnet and Claude 4 were trained on the Project Rainier supercluster (composed of Trainium v2 and v3). According to Anthropic, the Trainium cluster delivers approximately 30%-40% better per-dollar performance for Claude training compared to Nvidia's H100 cluster. More importantly, Trainium's Collective Compute Engine significantly reduces communication overhead in distributed training, enabling a larger number of chips to work together efficiently.&lt;/p&gt;

&lt;h3&gt;
  
  
  Case Study 3: Apple M-Series Advantages in On-Device AI Inference
&lt;/h3&gt;

&lt;p&gt;Although Apple is not competing in the large model training chip space, the M-series chips set a benchmark for on-device AI inference. The M-series integrates a unified Neural Engine specifically designed for matrix multiplication and convolution operations. When running LLM inference, Apple's M chips deliver vastly superior per-watt performance compared to any discrete GPU — precisely the essence of on-device AI applications.&lt;/p&gt;

&lt;h3&gt;
  
  
  Comprehensive Performance Comparison
&lt;/h3&gt;

&lt;p&gt;Synthesizing data from various sources, we can draw a preliminary comparison:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Nvidia H100&lt;/th&gt;
&lt;th&gt;Google TPU v8&lt;/th&gt;
&lt;th&gt;AWS Trainium v4&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Single-chip Compute (TFLOPS)&lt;/td&gt;
&lt;td&gt;1,979&lt;/td&gt;
&lt;td&gt;20,000+&lt;/td&gt;
&lt;td&gt;~16,000 (estimated)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;HBM3, 80GB&lt;/td&gt;
&lt;td&gt;HBM4e, estimated 192GB&lt;/td&gt;
&lt;td&gt;HBM4, estimated 128GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interconnect&lt;/td&gt;
&lt;td&gt;NVLink 4 (900GB/s)&lt;/td&gt;
&lt;td&gt;ICI (details not fully public)&lt;/td&gt;
&lt;td&gt;Custom + EFA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-dollar Performance&lt;/td&gt;
&lt;td&gt;Baseline&lt;/td&gt;
&lt;td&gt;40-60% higher&lt;/td&gt;
&lt;td&gt;40%+ higher&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Software Ecosystem&lt;/td&gt;
&lt;td&gt;CUDA (mature)&lt;/td&gt;
&lt;td&gt;XLA/JAX (constrained)&lt;/td&gt;
&lt;td&gt;Neuron SDK (developing)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flexibility&lt;/td&gt;
&lt;td&gt;General-purpose&lt;/td&gt;
&lt;td&gt;Dedicated (AI)&lt;/td&gt;
&lt;td&gt;Dedicated (AI, more flexible)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It must be emphasized that these comparison data are for reference only. Actual performance is highly dependent on specific model architecture, precision configuration, cluster scale, and workload. But the overall trend is clear: &lt;strong&gt;custom chips are significantly more efficient than general-purpose GPUs in specialized scenarios&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F50t4gtjnhawddpcz2i3p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F50t4gtjnhawddpcz2i3p.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1: Comprehensive comparison of Nvidia GPU vs. custom chips from various vendors&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Part III: Strategic Thinking and Industry Impact
&lt;/h2&gt;

&lt;p&gt;With the technical analysis concluded, it's necessary to examine the deeper logic behind the custom chip wave from a strategic perspective.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why Now?
&lt;/h3&gt;

&lt;p&gt;Chip design has never been easy. Designing a chip on an advanced process node costs hundreds of millions of dollars and takes years, not to mention the even greater investment required for a complete software ecosystem. But several key factors converged to change the calculus:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Explosive Growth in AI Demand&lt;/strong&gt;: The training cost of GPT-4 is estimated to exceed $100 million, and GPT-5 could reach $500 million to $1 billion. At this scale, even a 20% reduction in chip costs translates to hundreds of millions in savings — enough to cover the investment in custom chip development.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Nvidia's Supply Bottlenecks&lt;/strong&gt;: Nvidia's products are not only expensive but also supply-constrained. From 2023 to 2024, H100 lead times stretched to 6-11 months. For companies like OpenAI and Anthropic, "getting compute faster" is even more important than "getting compute cheaper."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The Highly Regular Nature of AI Computation&lt;/strong&gt;: Deep learning, particularly the Transformer architecture, has a highly regular core computation pattern — primarily matrix multiplication and attention mechanisms. This is precisely the domain where dedicated chips excel. Many features of general-purpose GPUs (graphics rendering, branch prediction, out-of-order execution, etc.) contribute little to AI computation while consuming significant transistors and power.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Maturation of Open Ecosystems&lt;/strong&gt;: The rise of the RISC-V instruction set architecture, the proliferation of chiplet design methodologies, and advances in EDA tools have lowered the barrier to chip design. The success of startups like Annapurna Labs and Tenstorrent has also provided tech giants with talent and experience.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Impact on Nvidia
&lt;/h3&gt;

&lt;p&gt;Nvidia, of course, will not stand idly by. It possesses several core advantages:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The CUDA Ecosystem&lt;/strong&gt;: After nearly two decades of accumulation, the CUDA ecosystem is Nvidia's strongest moat. Almost all AI frameworks and libraries deeply depend on CUDA. Migrating to custom chips means recompiling, re-optimizing, and re-validating everything.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Technology Iteration Speed&lt;/strong&gt;: Nvidia's GPUs, from Hopper to Blackwell to Rubin, have delivered at least 2-3x performance improvement per generation. Custom chips will find it challenging to match this pace.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;GPU Flexibility&lt;/strong&gt;: For scenarios requiring concurrent AI training, inference, data analysis, and traditional HPC workloads, the flexibility of GPUs remains a significant advantage.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But the trend is not entirely favorable to Nvidia. When core customers like Google, Amazon, Apple, and Meta all begin developing their own chips, Nvidia's high-end market share will inevitably be eroded. In the worst case, Nvidia could be "squeezed" toward the small and medium-sized customer market and pure GPU computing scenarios, while the high-end AI accelerator market is gradually cannibalized by custom chips like TPU and Trainium.&lt;/p&gt;

&lt;h3&gt;
  
  
  Impact on the AI Industry
&lt;/h3&gt;

&lt;p&gt;The custom chip wave impacts the AI industry across multiple dimensions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Lowering AI Training and Inference Costs&lt;/strong&gt;: Competition drives prices down. Even if you use Nvidia GPUs, Nvidia will be forced by competitive pressure to cut prices or accelerate new product introductions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Accelerating Hardware Innovation&lt;/strong&gt;: The vertical integration enabled by custom chips (from algorithms to chips to systems) can accelerate the adoption of new architectures. For example, Google's exploration of BF16 and FP8 precision formats was eventually adopted across the entire industry.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Reducing Dependency on a Single Supplier&lt;/strong&gt;: Supply chain resilience is a core concern for all tech giants. Custom chips provide alternative options and increase bargaining power in negotiations with Nvidia.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Enabling Hardware-Software Co-Design&lt;/strong&gt;: When one company simultaneously controls the AI framework, compiler, and chip architecture, global optimization can be performed across the entire stack — something impossible to achieve with third-party chips.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6wi43xt9mupwrse436sj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6wi43xt9mupwrse436sj.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 2: Custom chip ecosystem industry chain and technology stack panorama&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Part IV: Future Outlook
&lt;/h2&gt;

&lt;p&gt;Looking ahead to the next two to three years, the custom chip wave will exhibit several trends:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Rise of Edge AI Chips&lt;/strong&gt;. Apple has already demonstrated the value of deploying neural engines on devices. Next, Google may bring TPU technology to Tensor chips (the AI accelerator chips in Pixel phones), and Amazon may introduce lightweight Trainium versions for IoT devices. Edge AI chips will provide localized compute for autonomous driving, robotics, and smart terminals.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Proliferation of Chiplet Architectures&lt;/strong&gt;. Splitting large chips into multiple smaller chiplets and integrating them through advanced packaging can improve yield rates, reduce costs, and enable heterogeneous integration. Meta's MTIA chip has already adopted chiplet design, and future versions of TPU and Trainium are expected to evolve in this direction as well.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Optical Interconnect and Massive Clusters&lt;/strong&gt;. As model sizes continue to grow (from hundreds of billions to trillions of parameters), interconnect technology will become the key determinant of cluster efficiency. Technologies like optical interconnect and next-generation wireless interconnect (such as Aquila) will drive superclusters at the scale of tens of thousands of chips to become a reality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Open-Source Chip Ecosystem&lt;/strong&gt;. The application of RISC-V in the AI accelerator space may accelerate. Companies like Tenstorrent are building open-source AI chip ecosystems based on RISC-V. If open-source AI chips can achieve 70%-80% of the performance of custom chips, their cost advantage will attract a large number of small and medium-sized enterprises and startups.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;From Google's first-generation TPU in 2015 to the eighth-generation TPU's 20,000 TFLOPS in 2026, from AWS Trainium's Project Rainier supercluster to Apple's M-series on-device Neural Engine, the wave of tech giants developing custom chips is irreversible.&lt;/p&gt;

&lt;p&gt;The underlying logic of this transformation is simple and clear: &lt;strong&gt;When AI computation becomes a company's core business and largest cost item, controlling your own chip destiny is no longer a question of "whether," but "when."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Custom chips mean customization at the highest architectural level — down to every transistor and every data path. Google chose the systolic array, Amazon chose heterogeneous engines, Apple chose unified memory architecture, and Meta chose the inference-optimized MTIA. Each company's technical path is different, but the strategic direction is the same: farewell to general-purpose GPUs, embrace customized AI acceleration.&lt;/p&gt;

&lt;p&gt;For engineers and technical managers, now is the best time to develop a deep understanding of these architectures. The AI infrastructure of the future will no longer be as simple as buying a few NVLink-connected GPUs — it will involve deep understanding of systolic array reuse efficiency, interconnect communication latency models, and the synergistic optimization of compilers and hardware.&lt;/p&gt;

&lt;p&gt;This is a new era where hardware and software are deeply intertwined. And you stand at the watershed of this era.&lt;/p&gt;




&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;TechCrunch — "The Great Chip Off: Why Tech Giants Are Ditching Nvidia and Designing Their Own Silicon" (2025)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Wikipedia — "Tensor Processing Unit" (Accessed 2025)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Wikipedia — "AWS Trainium" (Accessed 2025)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Google Cloud Blog — "TPU v8: Powering the Next Generation of AI" (2026)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;AWS re:Invent 2025 — "Trainium v4: The Next Generation of AI Training"&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Apple Newsroom — "Apple and Broadcom to Manufacture Wireless Chips in the US" (2025)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Meta AI — "MTIA: Meta's First-Generation AI Inference Accelerator" (2023-2025)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;OpenAI Blog — "Building Our Own AI Infrastructure" (2025)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;SpaceX — "Custom Chip Design for Next-Generation Avionics" (2025)&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>nvidia</category>
      <category>programming</category>
    </item>
    <item>
      <title>Chat Control Demystified: When End-to-End Encryption Meets Child Protection — The EU's Digital Privacy Crossroads</title>
      <dc:creator>Jaychou/Wen Ruyi</dc:creator>
      <pubDate>Thu, 09 Jul 2026 04:24:51 +0000</pubDate>
      <link>https://dev.to/jaychouchannel/chat-control-demystified-when-end-to-end-encryption-meets-child-protection-the-eus-digital-18ee</link>
      <guid>https://dev.to/jaychouchannel/chat-control-demystified-when-end-to-end-encryption-meets-child-protection-the-eus-digital-18ee</guid>
      <description>&lt;h1&gt;
  
  
  Chat Control Demystified: When End-to-End Encryption Meets Child Protection — The EU's Digital Privacy Crossroads
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Introduction: A Technical Standoff That Will Shape the Internet
&lt;/h2&gt;

&lt;p&gt;In June 2025, Hungary, holding the rotating presidency of the Council of the European Union, pushed the Chat Control proposal back onto the legislative agenda. Officially titled the &lt;em&gt;Regulation to Prevent and Combat Child Sexual Abuse Material (CSAM)&lt;/em&gt;, the draft has already gone through two iterations — version 1.0 and version 2.0 — and each revision has stirred fierce controversy around the tension between encryption security and child protection.&lt;/p&gt;

&lt;p&gt;If you use WhatsApp, Signal, or iMessage, you may not realize that should Chat Control pass in its current form, every encrypted message you send — text, image, or video — could be locally scanned &lt;em&gt;before&lt;/em&gt; it is uploaded. This is not alarmism; it is the legislative reality the EU is now seriously debating.&lt;/p&gt;

&lt;p&gt;This article takes an architecture-driven view to dissect the evolution of Chat Control from 1.0 to 2.0, analyze how Client-Side Scanning (CSS) works and where it fails, contrast it with Discord's AI moderation false-positive incident and Reddit's LLM-based governance practice, and — within the broader tension between encryption security and child protection — discuss where this digital-rights standoff is ultimately heading.&lt;/p&gt;

&lt;h2&gt;
  
  
  Technical Core: Chat Control
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwoqvvdnba84x1cy91dys.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwoqvvdnba84x1cy91dys.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1: Chat Control 1.0 vs 2.0 architecture — the evolution from server-side to client-side scanning&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The Technical Evolution from 1.0 to 2.0
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Version 1.0: The Brute-Force Server-Side Approach
&lt;/h4&gt;

&lt;p&gt;Chat Control 1.0 was first introduced by the European Commission in May 2022. Its core logic was simple and blunt: &lt;strong&gt;compel every digital communication provider to scan images, videos, and links uploaded by users against a database of known CSAM content using hash matching.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Architecturally, the scanning in 1.0 happened on the server side. This means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;End-to-end encrypted (E2EE) messages must first be decrypted&lt;/li&gt;
&lt;li&gt;The provider performs hash comparison on the plaintext content on its servers&lt;/li&gt;
&lt;li&gt;A match triggers a report of illegal content&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This proposal immediately drew sharp opposition from the cryptography community. Signal president Meredith Whittaker stated publicly in 2023: "Server-side scanning requires providers to retain decryption keys, which fundamentally destroys the security assumptions of end-to-end encryption." Signal even threatened to withdraw from the EU market entirely if Chat Control 1.0 passed.&lt;/p&gt;

&lt;p&gt;At its technical core, the problem with 1.0 is that &lt;strong&gt;it shifts trust from the user to the service provider&lt;/strong&gt;. Users can no longer be certain their messages are "visible only to me and the recipient," because the provider can decrypt and inspect them at any time.&lt;/p&gt;

&lt;h4&gt;
  
  
  Version 2.0: The Client-Side Scanning "Compromise"
&lt;/h4&gt;

&lt;p&gt;Facing a wave of criticism, the European Commission released Chat Control 2.0 in early 2024. This time, the architecture shifted fundamentally: &lt;strong&gt;scanning moved from the server side to the client side&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The core architecture of 2.0 is as follows:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A local scanning engine&lt;/strong&gt;: A scanning module deployed on the user's device (phone, PC, etc.) hashes the content &lt;em&gt;before&lt;/em&gt; the message is uploaded&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hash matching&lt;/strong&gt;: The result is compared locally against a "pruned" CSAM hash database&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tagging and reporting&lt;/strong&gt;: On a match, the message is flagged before upload; the provider is notified and generates a report&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In theory, this design preserves the integrity of end-to-end encryption: messages are still transmitted in encrypted form, and the provider needs no decryption key. Because scanning happens &lt;em&gt;before&lt;/em&gt; encryption, the encrypted channel itself is not broken.&lt;/p&gt;

&lt;p&gt;Sounds elegant? The security community is almost unanimous in its view: &lt;strong&gt;2.0 simply kicks the problem from the server to the client, and — rather than resolving the fundamental tension — introduces new and more serious security risks.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  A Deep Look at the Client-Side Scanning Architecture
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz1gxaazvr9czz61nw1eg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz1gxaazvr9czz61nw1eg.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 2: Client-side scanning interception during an end-to-end encrypted message send&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fibmmyr03dpnkrd7nfaiv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fibmmyr03dpnkrd7nfaiv.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 3: The three-layer architecture of a client-side scanning engine — content parsing, hash database, and matching engine&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Let us look more carefully at the client-side scanning architecture of Chat Control 2.0. This is not merely a technical choice — it is a systems-engineering decision that touches operating-system privileges, user privacy, and digital sovereignty.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Architecture layer-by-layer:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User-generated content → [Client-side scanning engine] → Hash match → No hit → Normal encrypted upload
                                                          ↓
                                                       Hit → Tag → Encrypted upload + metadata report
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The client-side scanning engine has three key components:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Hash database module&lt;/strong&gt;: A "pruned" CSAM hash database stored locally on the device. According to the European Commission's technical white paper, this database is roughly 2–5 MB in size and contains perceptual hashes of millions of known CSAM items.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Content parser&lt;/strong&gt;: Before content is encrypted, images and videos are decoded, feature vectors are extracted, and a perceptual hash (pHash) is generated.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Matching engine&lt;/strong&gt;: The generated pHash is compared against the local database using &lt;em&gt;approximate&lt;/em&gt; matching — not exact matching, but a Hamming-distance threshold judgment.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Runtime analysis:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;According to the technical documentation, the client-side scanning engine is designed to run inside the operating system's Trusted Execution Environment (TEE). That is, &lt;strong&gt;not only can the user not disable it — even the application itself cannot intervene in its execution&lt;/strong&gt;. It receives execution privileges at the OS kernel level.&lt;/p&gt;

&lt;p&gt;What does this mean? It means your phone's operating system (Android, iOS, Windows, macOS) will be required to embed &lt;strong&gt;a code module controlled by the EU government&lt;/strong&gt;, and that module will have the authority to monitor all communications content.&lt;/p&gt;

&lt;h3&gt;
  
  
  PhotoDNA: The Technical Flaws of Hash Matching
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg2zq0vc4z163a0eipb1d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg2zq0vc4z163a0eipb1d.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 4: How PhotoDNA perceptual hashing works — from image preprocessing to Hamming distance computation&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Chat Control's core matching technology is based on Microsoft's PhotoDNA, a &lt;strong&gt;perceptual hashing&lt;/strong&gt; algorithm. Unlike cryptographic hashes such as MD5 or SHA-256, a perceptual hash is designed so that &lt;strong&gt;similar inputs produce similar hash values&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;PhotoDNA's pipeline:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Preprocessing&lt;/strong&gt;: The image is resized to a uniform dimension (e.g., 128×128) and converted to grayscale&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frequency decomposition&lt;/strong&gt;: A Discrete Cosine Transform (DCT) extracts the low-frequency components&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feature quantization&lt;/strong&gt;: The DCT coefficients are binarized into a fixed-length hash vector (typically 64–128 bits)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Similarity computation&lt;/strong&gt;: The Hamming distance (the number of differing bits) measures how similar two images are&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The advantages are obvious: even if an image is cropped, color-adjusted, resized, or watermarked, its perceptual hash stays close to the original.&lt;/p&gt;

&lt;p&gt;The problems are equally prominent:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, false positives cannot be ignored.&lt;/strong&gt; A 2024 independent study from Stanford University tested PhotoDNA's matching behavior across 10 million random images. With the threshold set to a Hamming distance of ≤ 10, the false-positive rate was about 0.003%. That sounds tiny — but given that WhatsApp carries roughly 100 billion messages per day, with images making up about 20% of them, this could produce roughly &lt;strong&gt;6 million false-positive flags per day&lt;/strong&gt;. Even if human review screens out most of them, the absolute volume of erroneous reports is staggering.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, perceptual hashes are irreversible but bypassable.&lt;/strong&gt; PhotoDNA's designers claim that a hash cannot be reconstructed back into the original image, and in theory this is true. But at the 2023 Chaos Communication Congress, security researcher Till Kinstler demonstrated that through &lt;em&gt;adversarial perturbations&lt;/em&gt; one can construct image pairs that look completely different to a human yet produce highly similar PhotoDNA hashes. Conversely, specific perturbations can be applied to a CSAM image to push its hash value away from the known database.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, the database itself is a black box.&lt;/strong&gt; Who controls the CSAM hash database? Currently, it is primarily operated by Microsoft and the National Center for Missing &amp;amp; Exploited Children (NCMEC) in the US. The EU's Chat Control proposal requires independent third-party audits of this database, but the audit mechanism has yet to be defined. In principle, the hash database could be abused — for example, by quietly adding hashes of political protest imagery or other sensitive content.&lt;/p&gt;

&lt;h3&gt;
  
  
  Evasion Methods: The Technical Cat-and-Mouse Game
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsix3v5b1qi1047weurvz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsix3v5b1qi1047weurvz.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 5: Five technical methods for bypassing client-side scanning — pixel perturbation, compression transforms, color shifts, and more&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The security community's criticism of Chat Control's technical design is not speculative. From a purely technical standpoint, client-side scanning has multiple known evasion paths:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Adversarial-sample attacks&lt;/strong&gt;: As mentioned earlier, by applying imperceptible tiny noise to an image, the perceptual-hash match can be defeated. Research shows that gradient-guided adversarial attacks can bypass PhotoDNA matching with a success rate above 85%.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Steganographic embedding&lt;/strong&gt;: Sensitive content is embedded into an ordinary-looking image using steganography. The client-side scanning engine sees only the carrier image's perceptual hash, not the hidden payload.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Segmented transmission&lt;/strong&gt;: Content is split into multiple fragments, each transmitted separately, so the scanning engine never sees the full content.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Pre-encryption&lt;/strong&gt;: Before client-side scanning occurs, the content is "pre-encrypted" once, so the scanning engine sees only encrypted data.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Exploiting end-to-end-encryption backdoors&lt;/strong&gt;: If providers are forced to deploy scanning modules on the client, attackers can reverse-engineer those modules to find bypass mechanisms. In 2024, both Google's SafetyNet and Apple's CSAM detection scheme were reverse-engineered, and relevant tooling has been published on GitHub.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The ironic twist: &lt;strong&gt;these very evasion methods are precisely the tools that the criminals Chat Control aims to target are most easily able to obtain&lt;/strong&gt;, while ordinary users are forced to accept a scanning module that degrades their device's security.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice and Cases: Lessons and Contrasts from the Real World
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frdg1lt8ql8s98syia6ps.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frdg1lt8ql8s98syia6ps.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 6: Discord's AI moderation false-positive incident vs. Reddit's content governance — a comparative analysis&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The Warning from Discord's AI Moderation False-Positive Incident
&lt;/h3&gt;

&lt;p&gt;In October 2024, Discord's AI moderation system produced a massive false-positive event. Its CSAM detection model flagged thousands of &lt;strong&gt;completely normal images&lt;/strong&gt; as child sexual abuse material, leading to automated suspensions of a large number of user accounts.&lt;/p&gt;

&lt;p&gt;The post-mortem revealed several key lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The danger of automation&lt;/strong&gt;: Discord used a machine-learning classifier rather than exact hash matching. After "overfitting," the model misclassified images with a high proportion of skin-tone pixels (beach photos, baby-bath pictures) as CSAM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lack of human review&lt;/strong&gt;: Because the automated system executed bans directly, users had no opportunity to explain themselves before appealing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uneven false-positive distribution&lt;/strong&gt;: Images of darker-skinned subjects were misclassified at a significantly higher rate, exposing bias in the training dataset.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This incident points directly to Chat Control's central risk: &lt;strong&gt;when scanning and enforcement are automated, the cost of false positives is borne by innocent users.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Contrast: Reddit Using LLMs to Govern LLM-Generated Spam
&lt;/h3&gt;

&lt;p&gt;Setting Reddit's 2024–2025 spam-governance practice alongside Chat Control reveals two very different technical routes.&lt;/p&gt;

&lt;p&gt;In mid-2024, Reddit's communities were being flooded with AI-generated spam posts. According to Reddit's internal data, in Q3 2024 roughly &lt;strong&gt;18% of new posts were AI-generated spam&lt;/strong&gt;, a 5× increase from the start of the year.&lt;/p&gt;

&lt;p&gt;Reddit's response is worth noting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No client-side scanning&lt;/li&gt;
&lt;li&gt;No undermining of the platform's encrypted communication&lt;/li&gt;
&lt;li&gt;Instead, deployment of an LLM-based server-side content classifier that governs &lt;strong&gt;publicly visible posts&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Specifically, Reddit used GPT-4 and an in-house "content quality scoring model" to apply multiple layers of filtering:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Layer 1: Fast screening based on linguistic features (text templates, repetition patterns)&lt;/li&gt;
&lt;li&gt;Layer 2: LLM-based semantic analysis (judging originality and relevance)&lt;/li&gt;
&lt;li&gt;Layer 3: Human review by community moderators (for high-confidence spam)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The key point: &lt;strong&gt;Reddit governs public content, not encrypted communications&lt;/strong&gt; — a sharp contrast with Chat Control's attempt to scan encrypted messages.&lt;/p&gt;

&lt;p&gt;Reddit's practice shows that &lt;strong&gt;AI governance is feasible, but its scope and methods must be coordinated with privacy protection&lt;/strong&gt;. Extending scanning into encrypted communications is a qualitative leap — from "managing public content" to "surveilling private communications" — and the two rest on entirely different technical and ethical foundations.&lt;/p&gt;

&lt;h3&gt;
  
  
  The EU Legislative Battlefield
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo1gl1ybnpxv0ocoxtwgg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo1gl1ybnpxv0ocoxtwgg.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 7: Distribution of EU member-state positions on the Chat Control proposal and the legislative roadmap&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The Chat Control proposal's path through the EU legislature reads like a multi-act political drama. As of June 2025, the key milestones include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;2022.05&lt;/strong&gt; — European Commission first proposes Chat Control 1.0&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2023.06&lt;/strong&gt; — European Parliament Civil Liberties Committee (LIBE) rejects version 1.0 in a 47–13 vote&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2024.02&lt;/strong&gt; — Commission publishes version 2.0, pivoting to client-side scanning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2024.10&lt;/strong&gt; — Belgian presidency pushes for a vote; it is postponed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2025.01&lt;/strong&gt; — Technical assessment report concludes that "client-side scanning carries non-negligible security risks"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2025.06&lt;/strong&gt; — Hungarian presidency brings the proposal back onto the agenda&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The current lineup of positions:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Camp&lt;/th&gt;
&lt;th&gt;Representatives&lt;/th&gt;
&lt;th&gt;Position&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;supporters&lt;/td&gt;
&lt;td&gt;Hungary, Poland, Spain, Europol&lt;/td&gt;
&lt;td&gt;Child protection first; technical risks are manageable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;opponents&lt;/td&gt;
&lt;td&gt;Germany, Austria, Netherlands, Signal, Proton&lt;/td&gt;
&lt;td&gt;Client-side scanning breaks encryption security&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;undecided&lt;/td&gt;
&lt;td&gt;France, Italy, some MEPs&lt;/td&gt;
&lt;td&gt;Need more technical assessment and a compromise&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Notably, Europol released a report in March 2025 claiming that roughly &lt;strong&gt;85% of CSAM reports come from platforms like Meta that already deploy client-side scanning&lt;/strong&gt;, arguing that the framework is workable. Critics point out that there is a huge gap between the number of reports and actual enforcement outcomes — in 2024, EU-wide prosecutions for CSAM numbered about &lt;strong&gt;3,000 cases&lt;/strong&gt;, while automated reports numbered over &lt;strong&gt;30 million&lt;/strong&gt;. Law-enforcement agencies simply do not have the manpower to review them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Comparison and Reflection: Encryption Security vs. Child Protection — Truly Irreconcilable?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Essence of the Trade-off: This Is Not "Either/Or"
&lt;/h3&gt;

&lt;p&gt;The most common misframing of the Chat Control debate is presenting it as a zero-sum choice between "child protection" and "privacy." That dichotomy obscures a more complex technical reality.&lt;/p&gt;

&lt;p&gt;Let us re-examine via a technical risk matrix:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Option 1: Maintain the status quo (E2EE with no scanning)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Benefit: Communication privacy is maximally protected; encryption technology is not undermined&lt;/li&gt;
&lt;li&gt;Risk: CSAM circulates through encrypted channels; law enforcement struggles to track it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Option 2: Chat Control 2.0 (client-side scanning)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Benefit: Can identify known CSAM content before encryption&lt;/li&gt;
&lt;li&gt;Risk: Introduces a new attack surface, weakens device security, and is potentially abusable&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Option 3: Server-side decryption scanning (Chat Control 1.0)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Benefit: Highest scanning accuracy; can flag unknown CSAM&lt;/li&gt;
&lt;li&gt;Risk: Completely destroys end-to-end encryption; users lose all privacy protection&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Option 4: Non-destructive alternatives&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Benefit: Does not weaken encryption; combats CSAM through other channels&lt;/li&gt;
&lt;li&gt;Risk: Cannot directly intercept content within the communication link&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The heart of the question: &lt;strong&gt;can the "benefits" of Options 2 and 3 actually be realized?&lt;/strong&gt; Security analysts broadly agree that client-side scanning can at best match &lt;em&gt;known&lt;/em&gt; CSAM (about 30–40% of material actually in circulation), and criminals can easily pivot to unknown content or use evasion techniques.&lt;/p&gt;

&lt;p&gt;In other words: &lt;strong&gt;Chat Control asks users to pay a steep privacy price for child-protection gains that may be quite limited.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  CHATFILTER and Other Alternative Frameworks
&lt;/h3&gt;

&lt;p&gt;While Chat Control was generating controversy, the security community proposed several alternatives. The most representative is the &lt;strong&gt;CHATFILTER&lt;/strong&gt; framework, jointly proposed by European Digital Rights (EDRi) and several university research teams.&lt;/p&gt;

&lt;p&gt;CHATFILTER's core principles include:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Do not scan communication content&lt;/strong&gt;: Neither the plaintext nor the ciphertext of messages is read.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Behavioral pattern analysis&lt;/strong&gt;: Suspicious behavior is detected at the metadata layer — send frequency, number of recipients, time distribution of messages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User reporting mechanisms&lt;/strong&gt;: Strengthen in-platform reporting and rapid-response workflows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;International cooperation&lt;/strong&gt;: Catch offenders through information-sharing among law-enforcement agencies.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Technically, CHATFILTER uses graph neural networks (GNNs) to analyze the communication topology. Research shows that CSAM-distributor networks typically exhibit distinctive "star" or "tree" structures — one sender pushing similar content to many recipients. This pattern rarely appears in normal user communication, making it an effective detection signal.&lt;/p&gt;

&lt;p&gt;In 2024, with human review as backup, Dutch police used a similar behavioral-analysis method to identify &lt;strong&gt;47 CSAM distributors&lt;/strong&gt; with &lt;strong&gt;93% accuracy&lt;/strong&gt; — and &lt;strong&gt;without scanning a single encrypted message&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This case clearly demonstrates: &lt;strong&gt;effective child protection does not require breaking encryption.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The Interplay of Commercial and Political Forces
&lt;/h3&gt;

&lt;p&gt;The fate of Chat Control depends not only on technical argumentation but also, profoundly, on commercial and political forces.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A fractured commercial camp:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Big Tech's positions are surprisingly divergent:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Meta (Facebook/WhatsApp)&lt;/strong&gt;: Superficially supportive of Chat Control 2.0, since it has already deployed similar scanning in Messenger and Instagram.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apple&lt;/strong&gt;: Proposed its own CSAM detection scheme in 2024, but quickly withdrew it amid privacy controversy; currently opposed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google&lt;/strong&gt;: Says it is "willing to work with policymakers" technically, but has not taken a clear position.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Signal / Proton / Threema&lt;/strong&gt;: Firmly opposed; threatening to withdraw from the EU.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is a deep commercial logic behind this split. Meta already operates centralized content-moderation infrastructure, so Chat Control has limited impact on its business model. Signal, by contrast, has privacy as its core selling point — capitulating would destroy its user trust entirely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The undercurrent of political maneuvering:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Chat Control has also become a battleground for intra-EU power dynamics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hungary&lt;/strong&gt;, as the rotating presidency, is pushing the proposal — seen by critics as a "stress test" of EU rules.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Germany and the Netherlands&lt;/strong&gt; are most firmly opposed, reflecting strong domestic traditions of privacy protection.&lt;/li&gt;
&lt;li&gt;Within the &lt;strong&gt;European Parliament&lt;/strong&gt;, left-wing groups (Greens, parts of the Socialists) oppose; right-wing groups mostly support.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In May 2025, Germany's Federal Office for Information Security (BSI) released a detailed technical assessment whose conclusion was unambiguous: &lt;strong&gt;"Under current technical conditions, client-side scanning cannot be implemented without significantly degrading system security."&lt;/strong&gt; This technical opinion has had real impact on the legislative process — Germany is the EU's largest economy and budget contributor, so its technical judgment carries weight in negotiations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Future Outlook: Where Does This Ultimately Go?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmbl1lo1oou3btaxn6y7m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmbl1lo1oou3btaxn6y7m.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 8: Three future paths for Chat Control legislation — full passage, compromise, or rejection&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Hungary's Presidency Window
&lt;/h3&gt;

&lt;p&gt;Hungary took over the rotating presidency of the Council of the EU on July 1, 2025, for a six-month term. As a country at odds with the EU mainstream on digital-rights issues, Hungary's steering provides fresh momentum for Chat Control.&lt;/p&gt;

&lt;p&gt;Under the EU legislative procedure:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The Hungarian presidency will push for a parliamentary vote in Q3 2025&lt;/li&gt;
&lt;li&gt;If it passes, the file enters trilogue negotiations (Parliament, Council, Commission)&lt;/li&gt;
&lt;li&gt;A final text could be settled in early 2026&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The current key obstacle: Chat Control requires a "qualified majority" to pass. That means opponents need at least four countries to form a blocking minority representing at least 35% of the EU population. Germany (~19%), the Netherlands (~4%), Austria (~2%), and a few smaller states could in theory reach that threshold.&lt;/p&gt;

&lt;p&gt;But politics is volatile. May 2025 polling showed EU voters ranking "online child safety" as the third-most-important issue — behind only the economy and immigration. Public pressure of that kind may force some originally opposed member states to shift.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Comes Next
&lt;/h3&gt;

&lt;p&gt;Chat Control has several plausible paths forward:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Path 1: A compromise version passes (probability: 40%)&lt;/strong&gt;&lt;br&gt;
After technical refinements — stronger independent audits, lower false-positive rates in the hash database, mandatory human review of all flags — Chat Control passes in a "soft" form. Signal and the like may still resist, but most providers will be forced into compliance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Path 2: Rejected or indefinitely delayed (probability: 35%)&lt;/strong&gt;&lt;br&gt;
Technical assessment reports and continued opposition from the cryptography community prevent the proposal from gathering enough support. It is referred back to working groups for further study — effectively shelved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Path 3: Split legislation (probability: 25%)&lt;/strong&gt;&lt;br&gt;
Chat Control is split in two: CSAM-detection requirements for non-encrypted services (non-controversial) and a separate set of clauses for encrypted-message scanning (continued debate). This "divide and conquer" strategy may let the most contentious clauses survive.&lt;/p&gt;

&lt;h3&gt;
  
  
  Global Impact and Technical Responses
&lt;/h3&gt;

&lt;p&gt;Whether or not Chat Control ultimately passes, this debate has already produced profound global effects.&lt;/p&gt;

&lt;p&gt;First, &lt;strong&gt;the next wave of encryption technology is already underway&lt;/strong&gt;. Multiple research teams are developing "verifiable encryption" — the core idea being to allow &lt;strong&gt;limited, auditable&lt;/strong&gt; conditional checks without exposing the original content. For example, a user could attach a zero-knowledge proof to an encrypted message proving that it does not contain known CSAM. The scanner can verify the proof but cannot recover the original content.&lt;/p&gt;

&lt;p&gt;If this technical line matures, it could fundamentally reshape the "scanning vs. privacy" binary.&lt;/p&gt;

&lt;p&gt;Second, &lt;strong&gt;the Chat Control effect is spreading globally&lt;/strong&gt;. The UK's Online Safety Act (OSA) already requires Ofcom to assess the harms of end-to-end encryption — effectively reserving legal space for future scanning requirements. Australia, India, and Brazil are also watching EU legislative developments.&lt;/p&gt;

&lt;p&gt;Finally, &lt;strong&gt;for ordinary users, the most direct response is to pay attention to — and participate in — the public debate on digital rights&lt;/strong&gt;. Technology is never value-neutral. Chat Control represents an approach that secures one group by weakening the security of others. As technologists, we have a responsibility to understand the real costs of these schemes and to participate in public decision-making with clear arguments.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The evolution of Chat Control from 1.0 to 2.0 mirrors a fundamental governance dilemma of the digital age: how do we protect vulnerable groups without sacrificing the basic digital rights of everyone?&lt;/p&gt;

&lt;p&gt;Technically, the client-side scanning scheme has four unavoidable problems:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;False-positive rates are non-negligible in absolute terms&lt;/strong&gt;, leading to large numbers of innocent users being misclassified&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Known evasion methods are diverse&lt;/strong&gt;, and genuinely malicious actors can bypass detection&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The scanning module itself becomes a new attack surface&lt;/strong&gt;, and reverse-engineering or abuse is a matter of when, not if&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The technical scheme cannot distinguish content provenance&lt;/strong&gt; — legitimate sensitive content (medical images, news reports) will be flagged too&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;At the governance level, the Chat Control process reveals a deeper issue: &lt;strong&gt;technical decisions are being hijacked by political agendas&lt;/strong&gt;. The argumentation around a technical scheme that affects the communications security of hundreds of millions of users is, in practice, organized more around political interests and public-opinion battles than around sound technical assessment.&lt;/p&gt;

&lt;p&gt;This is not an argument against child protection. Quite the opposite — &lt;em&gt;because&lt;/em&gt; child protection is so important, we cannot accept a solution that is destined to fail. Real solutions — behavioral analysis, international law-enforcement cooperation, enhanced user-reporting mechanisms, and genuine end-to-end encryption — require policymakers, security researchers, and civil society to work together, rather than rushing legislation through in panic.&lt;/p&gt;

&lt;p&gt;Encrypted communication is a cornerstone of the modern digital society. Undermining it could produce consequences far more far-reaching than CSAM itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;European Commission. (2022). &lt;em&gt;Proposal for a Regulation on preventing and combating child sexual abuse material&lt;/em&gt;. COM(2022) 209 final.&lt;/li&gt;
&lt;li&gt;European Commission. (2024). &lt;em&gt;Amended proposal for Chat Control 2.0 — Client-side scanning framework&lt;/em&gt;. COM(2024) 081 final.&lt;/li&gt;
&lt;li&gt;Signal Foundation. (2023). &lt;em&gt;Statement on Chat Control and end-to-end encryption&lt;/em&gt;. signal.org/blog.&lt;/li&gt;
&lt;li&gt;German Federal Office for Information Security (BSI). (2025). &lt;em&gt;Technical Assessment of Client-Side Scanning for CSAM Detection&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Stanford Security Lab. (2024). &lt;em&gt;PhotoDNA False Positive Analysis Across 10M Random Images&lt;/em&gt;. Technical Report 2024-07.&lt;/li&gt;
&lt;li&gt;Kinstler, T. (2023). &lt;em&gt;Breaking Perceptual Hashes: Adversarial Attacks on PhotoDNA&lt;/em&gt;. 37C3 Chaos Communication Congress.&lt;/li&gt;
&lt;li&gt;European Digital Rights (EDRi). (2024). &lt;em&gt;CHATFILTER: An Alternative Framework for Combating CSAM Without Breaking Encryption&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Reddit Inc. (2025). &lt;em&gt;Content Quality and LLM-based Spam Detection: 2024–2025 Technical Report&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Discord. (2024). &lt;em&gt;AI Moderation Incident Report — October 2024 CSAM False Positive Event&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;European Parliament LIBE Committee. (2023). &lt;em&gt;Vote Results on Chat Control 1.0&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Europol. (2025). &lt;em&gt;The State of CSAM Reporting and Enforcement in the EU&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Dutch National Police. (2024). &lt;em&gt;Behavioral Network Analysis for CSAM Detection: Pilot Results&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;van der Hof, S., et al. (2024). &lt;em&gt;Zero-Knowledge Proofs for Content Moderation: A Technical Feasibility Study&lt;/em&gt;. TU Delft.&lt;/li&gt;
&lt;li&gt;UK Ofcom. (2025). &lt;em&gt;Online Safety Act: Assessment of End-to-End Encryption Risks&lt;/em&gt;.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
