<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Tobias Reithmeier</title>
    <description>The latest articles on DEV Community by Tobias Reithmeier (@comic_sans).</description>
    <link>https://dev.to/comic_sans</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4045119%2F6cfa1202-9900-479f-ac37-e47b0b723739.jpg</url>
      <title>DEV Community: Tobias Reithmeier</title>
      <link>https://dev.to/comic_sans</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/comic_sans"/>
    <language>en</language>
    <item>
      <title>Trained Under Censorship: The Quiet Danger of AI Models from China</title>
      <dc:creator>Tobias Reithmeier</dc:creator>
      <pubDate>Mon, 20 Jul 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/comic_sans/trained-under-censorship-the-quiet-danger-of-ai-models-from-china-20g</link>
      <guid>https://dev.to/comic_sans/trained-under-censorship-the-quiet-danger-of-ai-models-from-china-20g</guid>
      <description>&lt;p&gt;Yesterday I wrote &lt;a href="https://www.tobiasreithmeier.de/en/blog/kimi-k3-open-weights-ai-market?campaign=devto" rel="noopener noreferrer"&gt;about Kimi K3&lt;/a&gt; - the largest open-weight model ever, technically impressive, soon freely available to everyone. Today, the other side of the coin. A language model is never neutral: it reflects what its training data contains and what it was aligned toward after training. And when that training happens in a market where the state prescribes by law what a model may say, every one of these models exports those prescriptions along with it. Not as a bug, but as a feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  The party line is written into law
&lt;/h2&gt;

&lt;p&gt;In China, the political alignment of AI models is not a company decision but a legal obligation. The &lt;a href="https://en.wikipedia.org/wiki/Interim_Measures_for_the_Management_of_Generative_AI_Services" rel="noopener noreferrer"&gt;Interim Measures for the Management of Generative AI Services&lt;/a&gt;, in force since August 2023, require generative AI services to uphold "core socialist values" and prohibit content that would "subvert state power", "harm the nation's image", or "promote separatism". Publicly accessible models need a security review and must register their algorithms with the Cyberspace Administration of China (CAC). Even earlier, the &lt;a href="https://www.whitecase.com/insight-our-thinking/ai-watch-global-regulatory-tracker-china" rel="noopener noreferrer"&gt;Deep Synthesis regulations of 2022&lt;/a&gt; governed AI-generated content and banned its use for "fake news" - with the state deciding what counts as fact and what as fabrication.&lt;/p&gt;

&lt;p&gt;The crucial point: these rules do not kick in at deployment - they shape the training itself. Anyone who wants to launch a model in China aligns it from the start so it passes review. The party line becomes part of the weights - those billions of numbers that condense everything a model "knows" and how it answers.&lt;/p&gt;

&lt;h2&gt;
  
  
  What studies actually measure
&lt;/h2&gt;

&lt;p&gt;This is not a theoretical worry; it has been measured thoroughly. In April 2026, the think tank CEIAS &lt;a href="https://ceias.eu/chinese-llms-and-the-spillover-effects-of-political-alignment/" rel="noopener noreferrer"&gt;systematically tested four leading Chinese models&lt;/a&gt; - DeepSeek V3.2, Moonshot's Kimi K2.5, Alibaba's Qwen 3.5, and Zhipu's GLM-5 - with 5,760 questions covering 37 countries and 40 topic areas. The results:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On questions about other countries' Taiwan policy, the models answered with heavy distortion: Qwen in 86 percent of cases, DeepSeek in 81, Kimi in 75, GLM in 42 percent&lt;/li&gt;
&lt;li&gt;Overall, Kimi switched into a censorship mode in roughly one out of every three answers, DeepSeek in one out of four&lt;/li&gt;
&lt;li&gt;Language matters: across the ten most sensitive topics, the distortion rate rose from 24 percent in English to 59 percent in Mandarin&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;a href="https://chinamediaproject.org/2026/02/09/tokens-of-ai-bias/" rel="noopener noreferrer"&gt;China Media Project&lt;/a&gt; even managed to surface the internal directives Qwen3 follows: when asked about China's international reputation, the model was instructed to "avoid any negative or critical language" and to avoid direct references to Western countries. Its answers about China came out uniformly positive - not through refusal, but through systematically one-sided framing.&lt;/p&gt;

&lt;p&gt;And the effects reach far beyond China. The &lt;a href="https://cepa.org/article/chinese-ai-models-spread-propaganda-globally/" rel="noopener noreferrer"&gt;Estonian Foreign Intelligence Service warned in its 2026 security report&lt;/a&gt; that DeepSeek conceals key information and inserts Chinese propaganda - demonstrably in English, Japanese, Russian, Thai, Hindi, and other languages. An EU-funded audit by the nonprofit Policy Genome additionally found Russian-language responses that endorsed Kremlin talking points. The auditors' conclusion: none of the Chinese models tested was free of state information guidance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is subtler than crude censorship
&lt;/h2&gt;

&lt;p&gt;The real danger is not the model that refuses a question about Tiananmen - that stands out and is easy to document. More dangerous is the quiet shift: an answer on the Taiwan question that sounds like a neutral summary but adopts Beijing's framing. An economic analysis that highlights achievements and omits problems. A historical account in which certain events simply do not appear. Anyone who does not already know the facts will not notice the gap - and that is precisely what distinguishes opinion shaping from censorship.&lt;/p&gt;

&lt;p&gt;Then there is the multiplier effect of open weights. Alibaba's Qwen family alone &lt;a href="https://cepa.org/article/chinese-ai-models-spread-propaganda-globally/" rel="noopener noreferrer"&gt;recorded more than 9.5 million downloads in October and November 2025&lt;/a&gt; and served as the basis for roughly 2,800 derivative models. Every startup that builds on such a model, every app that embeds it, every chatbot that runs on it inherits the political alignment - usually without developers or users knowing. The distortion lives in the weights, and weights carry no label declaring what went into the training. Fine-tuning on your own data does not automatically remove it; it stays in the foundation.&lt;/p&gt;

&lt;h2&gt;
  
  
  From bias to fake-news machine
&lt;/h2&gt;

&lt;p&gt;Opinion distortion is the passive danger. The active one: open models at frontier level drive the cost of mass-producing disinformation to practically zero. Safety guardrails built in during post-training can be stripped from open weights through fine-tuning - what an API provider prevents, the operator of a downloaded model alone decides. A model that generates convincingly real news articles, fabricated local reporting, or tailored social media campaigns in dozens of languages now fits on a server in a basement.&lt;/p&gt;

&lt;p&gt;Together, the two produce an uncomfortable scenario: models whose worldview was shaped by a state become global infrastructure - and the same openness that makes them attractive also makes them a tool for anyone intent on spreading targeted falsehoods. China itself does not pursue this strategy in secret: AI exports are &lt;a href="https://www.chinafile.com/reporting-opinion/features/censorship-not-deterring-global-adoption-of-chinese-ai" rel="noopener noreferrer"&gt;explicitly regarded there as an instrument for shaping the global information space&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The necessary counter-check
&lt;/h2&gt;

&lt;p&gt;Honesty requires saying this: no model is value-free, American and European ones included. Every training run rests on data selection; every alignment rests on decisions about what counts as helpful, harmful, or sensitive. OpenAI and Anthropic make such decisions daily too. The difference is not whether, but who and how: at Western providers these are company decisions - open to criticism, publicly debated, correctable through competition. In China they are state mandates with the force of law, enforced by a regulator, in the service of one party's claim to interpretive authority.&lt;/p&gt;

&lt;p&gt;And there is a second twist: it is precisely the openness of the weights that makes the manipulation provable. Researchers were able to &lt;a href="https://arxiv.org/pdf/2504.17130" rel="noopener noreferrer"&gt;identify the internal representations of censorship in open Chinese models directly - and even switch them off&lt;/a&gt; - something impossible with a closed model behind an API. Open weights spread the bias, but they also ship the dissection kit. The studies underpinning this article exist only because the models can be examined.&lt;/p&gt;

&lt;h2&gt;
  
  
  What follows from this
&lt;/h2&gt;

&lt;p&gt;For dealing with models from state-directed markets, this means concretely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Origin is a selection criterion.&lt;/strong&gt; Whoever deploys a model adopts its worldview as the default - on political, historical, and societal topics that is not a side issue&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit before deployment.&lt;/strong&gt; Anyone building an open model into a product should test it on sensitive topic areas, not just benchmark scores&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Critical applications need curated models.&lt;/strong&gt; Search, news, education, and government services are the wrong places for unvetted foundation models of unknown shaping&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Media literacy remains the last line of defense.&lt;/strong&gt; An AI answer is not a neutral statement of fact but the product of a training process - that simple insight protects better than any regulation&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Kimi K3 and its siblings show that the technical gap between open and closed models has closed. This article shows why that is not the end of the story: a model's weights contain not just capability but conditioning. When models trained under censorship requirements become global infrastructure, state information guidance quietly migrates into apps, search results, and homework assignments around the world. The good news: open weights can be examined, and the tools for doing so are getting better. The bad news: examination only happens where someone looks. Looking is now mandatory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;CEIAS: &lt;a href="https://ceias.eu/chinese-llms-and-the-spillover-effects-of-political-alignment/" rel="noopener noreferrer"&gt;Chinese LLMs and the spillover effects of political alignment&lt;/a&gt; - systematic test of DeepSeek, Kimi, Qwen, and GLM with 5,760 questions (April 2026)&lt;/li&gt;
&lt;li&gt;China Media Project: &lt;a href="https://chinamediaproject.org/2026/02/09/tokens-of-ai-bias/" rel="noopener noreferrer"&gt;Tokens of AI Bias&lt;/a&gt; - internal directives extracted from Qwen3&lt;/li&gt;
&lt;li&gt;CEPA: &lt;a href="https://cepa.org/article/chinese-ai-models-spread-propaganda-globally/" rel="noopener noreferrer"&gt;Chinese AI Models Spread Propaganda Globally&lt;/a&gt; - Estonian Foreign Intelligence Service findings, Policy Genome audit, Qwen download figures&lt;/li&gt;
&lt;li&gt;ChinaFile: &lt;a href="https://www.chinafile.com/reporting-opinion/features/censorship-not-deterring-global-adoption-of-chinese-ai" rel="noopener noreferrer"&gt;Censorship Is Not Deterring Global Adoption of Chinese AI&lt;/a&gt; - CAC approval requirements and global adoption&lt;/li&gt;
&lt;li&gt;Wikipedia: &lt;a href="https://en.wikipedia.org/wiki/Interim_Measures_for_the_Management_of_Generative_AI_Services" rel="noopener noreferrer"&gt;Interim Measures for the Management of Generative AI Services&lt;/a&gt; - overview of China's AI regulation&lt;/li&gt;
&lt;li&gt;China Law Translate: &lt;a href="https://www.chinalawtranslate.com/en/generative-ai-interim/" rel="noopener noreferrer"&gt;Interim Measures (full translation)&lt;/a&gt; - official requirements including "core socialist values"&lt;/li&gt;
&lt;li&gt;White &amp;amp; Case: &lt;a href="https://www.whitecase.com/insight-our-thinking/ai-watch-global-regulatory-tracker-china" rel="noopener noreferrer"&gt;AI Watch: Global regulatory tracker - China&lt;/a&gt; - Deep Synthesis regulations and regulatory framework&lt;/li&gt;
&lt;li&gt;arXiv: &lt;a href="https://arxiv.org/pdf/2504.17130" rel="noopener noreferrer"&gt;Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control&lt;/a&gt; - how censorship can be detected and disabled in open weights&lt;/li&gt;
&lt;li&gt;arXiv: &lt;a href="https://arxiv.org/pdf/2506.01814" rel="noopener noreferrer"&gt;Analysis of LLM Bias (Chinese Propaganda &amp;amp; Anti-US Sentiment) in DeepSeek-R1 vs. ChatGPT o3-mini-high&lt;/a&gt; - quantitative bias comparison&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>kimi</category>
      <category>censorship</category>
    </item>
    <item>
      <title>Kimi K3: The Frontier Just Went Open</title>
      <dc:creator>Tobias Reithmeier</dc:creator>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/comic_sans/kimi-k3-the-frontier-just-went-open-e4h</link>
      <guid>https://dev.to/comic_sans/kimi-k3-the-frontier-just-went-open-e4h</guid>
      <description>&lt;p&gt;On July 16, 2026, the Chinese startup Moonshot AI unveiled its new flagship, Kimi K3: 2.8 trillion parameters, multimodal, a one-million-token context window - and the announcement that the complete model weights will be published freely on Hugging Face by the end of July. That would make K3 the largest open-weight model ever released. This alone would be a footnote if the model were mediocre. It is not: in independent testing it lands third to fourth among all models worldwide, and in some disciplines it takes first place. For the first time, an open model stands within striking distance of the closed frontier held by OpenAI and Anthropic.&lt;/p&gt;

&lt;p&gt;To understand why this is a turning point, you have to separate two things: what is technically being published - the weights - and what this publication means economically and politically. The two are more tightly connected than they first appear.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Kimi K3 can do
&lt;/h2&gt;

&lt;p&gt;The numbers first. On &lt;a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/moonshot-releases-2-8-trillion-parameter-kimi-k3" rel="noopener noreferrer"&gt;GDPval-AA v2&lt;/a&gt;, a benchmark of real-world work tasks across 44 occupations and 9 industries, K3 scores 1,687 - third place behind Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8), but ahead of Claude Opus 4.8 (1,600). On AA-Briefcase, an agentic long-horizon benchmark, K3 even climbs to second place with 1,527 points, ahead of GPT-5.6 Sol Max.&lt;/p&gt;

&lt;p&gt;More remarkable still: in the Frontend Code Arena, where real developers vote blindly between model responses, &lt;a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/moonshot-releases-2-8-trillion-parameter-kimi-k3" rel="noopener noreferrer"&gt;K3 leads the entire field with 1,679 points&lt;/a&gt; - ahead of Fable 5. In real-world task automation it wins four out of eight benchmarks, including SpreadsheetBench 2 and BrowseComp.&lt;/p&gt;

&lt;p&gt;And the price: $3 per million input tokens, $15 per million output tokens. That is mid-tier pricing for near-frontier performance. And anyone who prefers can wait for the weights and pay nothing at all - except for their own hardware.&lt;/p&gt;

&lt;h2&gt;
  
  
  What are weights, actually?
&lt;/h2&gt;

&lt;p&gt;The term "open weight" sounds technical but describes something surprisingly tangible. A language model is an artificial neural network: billions of simple computing nodes stacked in layers and connected to each other. Every one of these connections carries a number that determines how strongly a signal passes from one node to the next. These numbers are the weights.&lt;/p&gt;

&lt;p&gt;A useful analogy: the model's architecture - how many layers, how they are wired - is the blueprint of a brain. The weights are the strength of every single synapse within it. The blueprint alone can do nothing. Only the weights turn the empty network into a model that understands language, writes code, and reasons. Everything a model "knows" lives in these numbers - not as a searchable database, but distributed across the entire network, much like a memory does not sit in a single brain cell.&lt;/p&gt;

&lt;p&gt;For Kimi K3, that means 2.8 trillion such numbers. As a file, even heavily compressed, this is well over a terabyte. The code describing how to compute with these numbers, by contrast, is only a few thousand lines. The model is the weight file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where do the weights come from? Training
&lt;/h2&gt;

&lt;p&gt;At the start, all weights are random numbers. The network produces gibberish. Then pretraining begins: the model is shown trillions of text fragments from the internet, from books, from code archives, and must predict the next word (more precisely: the next token) each time. When it gets it wrong, the system computes which weights contributed how much to the error - this is backpropagation - and every weight is nudged a tiny step in the direction that would have made the error smaller.&lt;/p&gt;

&lt;p&gt;This process, called gradient descent, repeats trillions of times. Each individual correction is microscopic. In aggregate, something remarkable emerges: to predict the next word well in arbitrary text, the model must incidentally learn grammar, facts, logical relationships, programming languages, even something like a model of the world. All of it condenses into the weights.&lt;/p&gt;

&lt;p&gt;After pretraining comes post-training: the raw model is trained with human feedback and reinforcement learning to answer helpfully, follow instructions, and refuse harmful requests. This, too, changes only one thing - the weights.&lt;/p&gt;

&lt;p&gt;And here the loop closes back to geopolitics: training at this scale consumes months of compute on tens of thousands of specialized chips and hundreds of millions of dollars in electricity and hardware. &lt;strong&gt;The weights are the condensed result of that entire effort.&lt;/strong&gt; This is exactly why US export controls target chips: preventing the training prevents the weights. And it is exactly why it is so remarkable that Moonshot trained a frontier model despite those restrictions - and is now giving the result away.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open weight is not open source
&lt;/h2&gt;

&lt;p&gt;An important distinction that often gets blurred: Moonshot is publishing the weights, not the recipe. The training data, the training code, the countless detailed decisions that turn raw data into a top model - all of that stays secret. You get the finished cake, not the recipe and not the ingredient list. Genuine open-source software would be fully transparent; with open-weight models, the training can neither be audited nor reproduced.&lt;/p&gt;

&lt;p&gt;Even so, the finished cake is enormously valuable. With the weights you can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;run the model on your own hardware&lt;/strong&gt; - without a single byte flowing to a provider in the US or China&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;fine-tune it&lt;/strong&gt; , specializing it on your own tasks, domain language, or company data with comparatively little compute&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;distill it&lt;/strong&gt; : a large model teaches a small one what it knows - which is how one open frontier model spawns hundreds of compact offshoots&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;study it&lt;/strong&gt; : interpretability, bias, and safety vulnerabilities can only be researched on models whose internals are accessible&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then there is the strategic dimension. An API provider can raise prices, retire models, or change terms. Downloaded weights can never be taken back. For companies in regulated industries, for governments, for entire nations, this is the difference between dependency and sovereignty. Open models are also the price anchor of the whole market: they set the floor for what closed providers can charge.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for the market
&lt;/h2&gt;

&lt;p&gt;Until now, a rule of thumb held: open models trail the closed frontier by roughly nine to twelve months. If you wanted the best, you had to go to OpenAI, Anthropic, or Google - and pay their prices. K3 makes that rule obsolete. &lt;a href="https://www.axios.com/2026/07/16/moonshot-kimi-ai-china-model-openai-anthropic" rel="noopener noreferrer"&gt;No open release has ever stood this close to the closed frontier&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The second shift is geopolitical. The US strategy bet on compute superiority securing model leadership. Moonshot has &lt;a href="https://www.cnbc.com/2026/07/17/moonshot-ai-kimi-k3-model-openai-anthropic-china.html" rel="noopener noreferrer"&gt;delivered a frontier-level model despite chip restrictions&lt;/a&gt; - the competition is moving from raw compute toward training efficiency and know-how. And China has answered the contested question of whether frontier models may be open simply by publishing one.&lt;/p&gt;

&lt;p&gt;The third is economic: a near-frontier model at mid-tier prices, soon self-hostable by anyone, squeezes margins on every "good enough" workload. Raw model intelligence is becoming a commodity. What stays valuable is what surrounds the model: agent infrastructure, tooling, integration, enterprise contracts, trust, and liability.&lt;/p&gt;

&lt;h2&gt;
  
  
  The outlook for OpenAI and Anthropic
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Anthropic&lt;/strong&gt; is in the best short-term position. Claude Fable 5 Max still leads the rankings that matter by a clear margin, and Anthropic's enterprise and agentic-coding business depends less on consumer pricing than on reliability and integration. But the warning sign is unmistakable: Opus 4.8, a top model until recently, is already beaten by K3. Only the absolute flagship still justifies premium pricing - the second tier of the portfolio now competes directly with a model whose weights sit freely on the internet. The window in which a model lead can be monetized exclusively is shrinking from years to months.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenAI&lt;/strong&gt; is doubly exposed. The company depends heavily on the mass market, where "good enough and cheap" counts for more than the last benchmark point, and it carries enormous infrastructure commitments calculated on high margins. GPT-5.6 Sol holds second place overall, but K3 beats it in several coding and agent benchmarks. Despite its name, OpenAI has been hesitant in the open-weight space - and the pressure to show up there grows with every Chinese release.&lt;/p&gt;

&lt;p&gt;For both, the same holds: the race at the top continues, and the top remains valuable. But from now on, the distance to the free alternative determines how much you can charge for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The flip side
&lt;/h2&gt;

&lt;p&gt;The honest version of this story includes the flip side: once published, weights can never be recalled. Safety guardrails built in during post-training can be removed by anyone with moderate effort through fine-tuning. What a closed provider forbids by API policy is, with an open model, entirely up to whoever runs it. The closer open models get to the frontier, the more real this dilemma becomes - and K3 gets very close. The debate about how much openness is responsible at which capability level used to be largely theoretical. After July 27, it no longer is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Kimi K3 is more than another strong model from China. It is the moment the equation "top performance = closed and expensive" breaks. The weights - those 2.8 trillion numbers condensing months of training on tens of thousands of chips - will soon sit on Hugging Face for anyone to download. For users and companies, that is a gift: more choice, more sovereignty, falling prices. For OpenAI and Anthropic, it is a deadline: the value of their models will no longer be measured by what they can do, but by how much better they are than what anyone can download for free.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Tom's Hardware: &lt;a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/moonshot-releases-2-8-trillion-parameter-kimi-k3" rel="noopener noreferrer"&gt;Moonshot releases 2.8-trillion-parameter Kimi K3&lt;/a&gt; - detailed specifications and benchmark results&lt;/li&gt;
&lt;li&gt;CNBC: &lt;a href="https://www.cnbc.com/2026/07/17/moonshot-ai-kimi-k3-model-openai-anthropic-china.html" rel="noopener noreferrer"&gt;Chinese AI has leveled up, and brought renewed focus on the open weight model shift&lt;/a&gt; - market context and the open-weight shift&lt;/li&gt;
&lt;li&gt;Bloomberg: &lt;a href="https://www.bloomberg.com/news/articles/2026-07-17/china-s-powerful-new-moonshot-ai-model-closes-gap-with-us-rivals" rel="noopener noreferrer"&gt;Moonshot Unveils Kimi K3 AI Model, Narrowing Gap With US Rivals&lt;/a&gt; - geopolitical context&lt;/li&gt;
&lt;li&gt;Axios: &lt;a href="https://www.axios.com/2026/07/16/moonshot-kimi-ai-china-model-openai-anthropic" rel="noopener noreferrer"&gt;China's open-weight Kimi model stuns AI world with frontier-level results&lt;/a&gt; - independent test results&lt;/li&gt;
&lt;li&gt;Fortune: &lt;a href="https://fortune.com/2026/07/16/moonshots-kimi-k3-pushes-chinese-ai-into-fable-level-territory/" rel="noopener noreferrer"&gt;Moonshot's Kimi K3 pushes Chinese AI into Fable-level territory&lt;/a&gt; - competitive analysis&lt;/li&gt;
&lt;li&gt;VentureBeat: &lt;a href="https://venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems" rel="noopener noreferrer"&gt;China's Moonshot AI releases Kimi K3, the largest open-source model ever&lt;/a&gt; - release details and weights publication timeline&lt;/li&gt;
&lt;li&gt;Simon Willison: &lt;a href="https://simonwillison.net/2026/Jul/16/kimi-k3/" rel="noopener noreferrer"&gt;Kimi K3, and what we can still learn from the pelican benchmark&lt;/a&gt; - independent first assessment&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>kimi</category>
      <category>llm</category>
    </item>
    <item>
      <title>tokensave: An MCP Server That Saved Me Millions of Tokens While Coding</title>
      <dc:creator>Tobias Reithmeier</dc:creator>
      <pubDate>Fri, 17 Jul 2026 00:00:00 +0000</pubDate>
      <link>https://dev.to/comic_sans/tokensave-an-mcp-server-that-saved-me-millions-of-tokens-while-coding-4phl</link>
      <guid>https://dev.to/comic_sans/tokensave-an-mcp-server-that-saved-me-millions-of-tokens-while-coding-4phl</guid>
      <description>&lt;p&gt;For a few weeks now, a tool has been running in my projects that is invisible in daily use - until you look at the counter. tokensave has saved me over 12 million tokens so far. It's an MCP server that gives Claude Code a local code graph to query, instead of letting the model read whole files and search through the repository as plain text. Here's what it does, how it works under the hood, what a reproducible benchmark says about it - and where its limits are.&lt;/p&gt;

&lt;h2&gt;
  
  
  A quick detour: why tokens are the currency
&lt;/h2&gt;

&lt;p&gt;Language models don't read and write in characters or words, but in tokens - word fragments of roughly three to four characters. Everything an AI assistant gets to see counts: your question, the system instructions, and above all every file and every search result it loads into its context. As an order of magnitude, a source file with 1,000 lines costs around 10,000 tokens - per request, because the context window is transmitted anew with every API call.&lt;/p&gt;

&lt;p&gt;That has two consequences. First, money: billing is per token, usually priced per million. If your assistant reads half a repository for every code question, you pay for it. Second, quality: a model's context window is finite - between 200,000 and one million tokens on current models - and the fuller it is with irrelevant code, the less room remains for what actually matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  How a model splits text - and why code is expensive
&lt;/h2&gt;

&lt;p&gt;The splitting into tokens is done by a tokenizer, usually based on Byte Pair Encoding (BPE): the algorithm learns from large text corpora which character sequences frequently occur together and merges them into single tokens. Common English words become one token; rare ones fall apart into several. Three things follow from this that are worth knowing when you're trying to save:&lt;/p&gt;

&lt;p&gt;First, code tokenizes differently from prose. Indentation, brackets, and special characters are tokens of their own, and an identifier like &lt;code&gt;getUserAccountBalance&lt;/code&gt; breaks into several fragments - source code is often more expensive per character than plain text. Second, German is more expensive than English: tokenizers are trained predominantly on English text, so long German compound words split more finely. And third, token counts are not comparable across model generations - a new tokenizer can count the same text roughly a third differently. The rule of thumb "1,000 lines ≈ 10,000 tokens" is therefore exactly that: a rule of thumb.&lt;/p&gt;

&lt;h2&gt;
  
  
  Less context isn't just cheaper - it makes answers better
&lt;/h2&gt;

&lt;p&gt;One could object: context windows keep growing, so why save at all? Because large contexts have a documented quality problem. In 2023, Liu et al. showed in "Lost in the Middle: How Language Models Use Long Contexts" that language models use information at the beginning and end of a long context far more reliably than information in the middle - retrieval accuracy measurably drops there.&lt;/p&gt;

&lt;p&gt;For code questions, that means: if you dump three complete files into the model's context, 95 percent of which are irrelevant, you don't just pay for the irrelevant tokens - you also risk the relevant passage drowning in the middle of the context. Fewer but better-targeted tokens are therefore not just about thrift; they're a quality lever.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is tokensave?
&lt;/h2&gt;

&lt;p&gt;Anyone who has watched an AI assistant work through a codebase knows the pattern: to answer "where is this function called?", it first reads three files end to end, then greps its way through half the repository. That works - but it burns thousands of tokens every time, most of which contribute nothing to the answer.&lt;/p&gt;

&lt;p&gt;tokensave (&lt;a href="https://github.com/aovestdipaperino/tokensave" rel="noopener noreferrer"&gt;github.com/aovestdipaperino/tokensave&lt;/a&gt;, MIT license, around 450 stars on GitHub) tackles exactly that: instead of searching code as text, it answers questions about the structure of the code - from a local graph that is built once and then queried on demand. According to the project, it supports 34 programming languages, from Rust to Python, TypeScript, and Swift.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the code graph works
&lt;/h2&gt;

&lt;p&gt;During indexing, tokensave breaks the codebase down into its building blocks: functions, methods, types, and modules become nodes; their relationships - who calls whom, who imports what - become edges. The result is stored as a libSQL database (a SQLite fork) in the project folder. If you look inside, you'll find more than just the core tables &lt;code&gt;nodes&lt;/code&gt;, &lt;code&gt;edges&lt;/code&gt;, and &lt;code&gt;files&lt;/code&gt;: an FTS5 full-text index for symbol search, a &lt;code&gt;vectors&lt;/code&gt; table with embeddings for semantic similarity, and fingerprint tables that let the incremental sync detect which files have changed - a sync after small changes takes a measured 21 milliseconds on my machine.&lt;/p&gt;

&lt;p&gt;This architecture explains why queries are so cheap: the expensive part - parsing, indexing, embedding - happens once at init. After that, every question is just a database lookup. Asked the classic way, "where is &lt;code&gt;parse_config&lt;/code&gt; called?", the assistant has to read files or load grep hits with their surroundings into context. With tokensave, it asks the graph the same question and gets back a list of call sites, plus, on request, exactly the affected code sections: a few hundred tokens instead of thousands.&lt;/p&gt;

&lt;p&gt;On top of that sits a whole toolbox: symbol search, call chains across multiple levels, impact analysis ("what breaks if I change this?"), dead-code detection. For more complex structural questions, you can even query the database directly with SQL.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the graph doesn't see
&lt;/h2&gt;

&lt;p&gt;Honesty requires saying this too: a statically built graph only knows what can be resolved statically. Dynamic dispatch, reflection, duck typing in Python, macros - anything decided at runtime never becomes an edge. My own website project illustrates it: the graph contains 1,115 nodes but only 252 edges, all of them function calls. In a small Python and JavaScript codebase with a lot of dynamic glue, the relationship network stays sparse; in a large, statically typed Rust or Java codebase, the ratio would look very different.&lt;/p&gt;

&lt;p&gt;In practice that means: for "where is symbol X defined and who calls it directly?", the graph is excellent. For "which handler gets registered at runtime via this config file?", it doesn't help - that's still a job for classic reading. tokensave doesn't replace code understanding; it makes the most common navigation questions cheap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup in five minutes
&lt;/h2&gt;

&lt;p&gt;The setup itself is unremarkable:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Install&lt;/strong&gt; via crates.io, Homebrew, or as a prebuilt binary; ends up at e.g. &lt;code&gt;~/.local/bin/tokensave&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;tokensave init&lt;/code&gt;&lt;/strong&gt; inside the project folder creates &lt;code&gt;.tokensave/tokensave.db&lt;/code&gt; and indexes the code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;tokensave install&lt;/code&gt;&lt;/strong&gt; registers the tool as an MCP server in the agent configuration - after that, the &lt;code&gt;tokensave_*&lt;/code&gt; tools are available in Claude Code.&lt;/li&gt;
&lt;li&gt;After external changes, especially after a &lt;code&gt;git pull&lt;/code&gt;: &lt;strong&gt;&lt;code&gt;tokensave sync&lt;/code&gt;&lt;/strong&gt;. Otherwise the graph works off a stale state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;tokensave upgrade&lt;/code&gt;&lt;/strong&gt; fetches new versions from GitHub.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step 4 is the only one you really need to remember - more on why below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measured: the benchmark
&lt;/h2&gt;

&lt;p&gt;Claimed savings are one thing; reproducible measurements are another. tokensave ships its own command for this: &lt;code&gt;tokensave bench&lt;/code&gt; asks ten standard questions about the current project ("How is configuration loaded?", "How are errors defined and propagated?", …) and compares the token cost of the graph answer against the baseline, i.e. reading the relevant files.&lt;/p&gt;

&lt;p&gt;The result in my website project: &lt;strong&gt;94 percent mean savings&lt;/strong&gt; - 53,900 tokens for the baseline versus 3,000 tokens across all ten questions, so roughly 5,000 to 6,000 tokens per question the classic way against about 300 via the graph. If you want it even more concrete: &lt;code&gt;tokensave discover&lt;/code&gt; analyzes your own Claude Code history and shows which past navigation steps a graph query would have handled more cheaply.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it delivers day to day
&lt;/h2&gt;

&lt;p&gt;The savings aren't an estimate; they're right there in the built-in counter: over 12 million tokens across all my projects, growing daily. A pleasant side effect: answers to code questions arrive noticeably faster, because less context has to pass through the model - and they tend to be more precise, for exactly the "Lost in the Middle" reason above.&lt;/p&gt;

&lt;h2&gt;
  
  
  The prompt-caching nuance
&lt;/h2&gt;

&lt;p&gt;An obvious objection: the API providers cache anyway. True - with prompt caching, an unchanged context prefix is stored, and cache reads cost only about a tenth of the normal input price. That softens the "paid anew on every request" from the token detour, but it doesn't replace tokensave, because the two levers work in different places: caching makes &lt;em&gt;re-reading&lt;/em&gt; the same bloated context cheaper - and even a small change at the beginning invalidates the cache. tokensave keeps the context from getting bloated in the first place. One is a discount on a large bill; the other makes the bill small. In practice, you benefit from both at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stumbles
&lt;/h2&gt;

&lt;p&gt;The biggest pitfall is staleness. When the code changes outside your own session - the classic case being a &lt;code&gt;git pull&lt;/code&gt; - the graph keeps working off the old state. To be fair: current versions detect commits since the last sync and warn explicitly in the status check ("1 commit since last sync"). But if you miss the warning, you still get quietly wrong answers until &lt;code&gt;tokensave sync&lt;/code&gt; runs - and until then, you'll be hunting the bug in the wrong place.&lt;/p&gt;

&lt;p&gt;Two smaller limitations on top of that: queries are capped per task, so complex questions have to be bundled instead of tried one at a time. And tokensave only answers questions about code structure - for web search, documentation, or anything beyond the repository, it naturally does nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local, but not network-free: the worldwide counter
&lt;/h2&gt;

&lt;p&gt;The code graph, the search, the embeddings - all of that runs locally; no source code leaves the machine. Still, tokensave isn't entirely network-free, and that belongs in any honest review: by default, it reports the number of saved tokens to a public "worldwide counter". What gets transmitted is a single HTTP POST with one number (something like &lt;code&gt;{"amount": 4823}&lt;/code&gt;); the country of origin is derived server-side from the IP - no code, no file names, no user ID. Add to that version checks against the GitHub releases API for &lt;code&gt;upgrade&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Both are documented and can be turned off (&lt;code&gt;tokensave disable-upload-counter&lt;/code&gt;). I've deliberately left the upload on - partly because the global counter is a nice argument for the tool itself: it currently stands at around 17.6 billion tokens saved across all users worldwide.&lt;/p&gt;

&lt;h2&gt;
  
  
  My experience
&lt;/h2&gt;

&lt;p&gt;A tool that writes itself into the agent configuration and reads along with every code question doesn't get the benefit of the doubt from me. So before putting it into production use, I looked at the repo metadata, the commit history, and its runtime behavior - the result was an actively maintained open-source project whose network behavior matches exactly what the documentation states: the counter upload, the version checks, nothing else.&lt;/p&gt;

&lt;p&gt;Since then, exactly what you'd hope for from a tool like this has happened: nothing. No incident, no surprise - just a counter that keeps growing and code questions that get answered faster. By now, tokensave is the first tool I reach for when researching code in any of my projects.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;tokensave is no miracle cure, but a well-thought-out specialist tool: it does one thing - answering code questions without full-text reads - and it does that measurably well, in my case with 94 percent savings in the reproducible benchmark and over 12 million tokens saved in daily use. The limits are known and manageable: the graph only sees static structure, it wants a sync after external changes, and anyone who insists on complete network abstinence turns off the counter upload. If you regularly work with AI assistants in larger codebases, it's worth a look. Just that &lt;code&gt;tokensave sync&lt;/code&gt; after a &lt;code&gt;git pull&lt;/code&gt; - that one you really need to remember.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;tokensave on GitHub: &lt;a href="https://github.com/aovestdipaperino/tokensave" rel="noopener noreferrer"&gt;github.com/aovestdipaperino/tokensave&lt;/a&gt; - source code, license, worldwide-counter documentation, and commit history&lt;/li&gt;
&lt;li&gt;Liu et al. (2023): &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;Lost in the Middle: How Language Models Use Long Contexts&lt;/a&gt; - study on retrieval accuracy in long contexts&lt;/li&gt;
&lt;li&gt;Anthropic: &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching" rel="noopener noreferrer"&gt;Prompt caching documentation&lt;/a&gt; - how prefix caching works and how it's priced&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>mcp</category>
    </item>
  </channel>
</rss>
