<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Prashanth Velidandi</title>
    <description>The latest articles on DEV Community by Prashanth Velidandi (@pmv_inferx).</description>
    <link>https://dev.to/pmv_inferx</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3382940%2F7c22cb83-8e48-49dd-8ec0-2aa7fb44b3e5.jpg</url>
      <title>DEV Community: Prashanth Velidandi</title>
      <link>https://dev.to/pmv_inferx</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/pmv_inferx"/>
    <language>en</language>
    <item>
      <title>Fight Open Source with Open Source</title>
      <dc:creator>Prashanth Velidandi</dc:creator>
      <pubDate>Mon, 14 Sep 2026 12:08:03 +0000</pubDate>
      <link>https://dev.to/pmv_inferx/fight-open-source-with-open-source-56ga</link>
      <guid>https://dev.to/pmv_inferx/fight-open-source-with-open-source-56ga</guid>
      <description>&lt;p&gt;Dario Amodei proposed something unusual for Silicon Valley: slow down.&lt;/p&gt;

&lt;p&gt;His argument is not that artificial intelligence should stop. It is that frontier capabilities are advancing faster than our ability to understand and secure them, and that we should deliberately create more time for safety to catch up.&lt;/p&gt;

&lt;p&gt;He proposes three broad steps: independent evaluators embedded inside frontier AI companies, coordination among AI companies in democratic countries, and eventually international coordination, most importantly between the United States and China.&lt;/p&gt;

&lt;p&gt;I agree with the problem more than I agree with the prescription.&lt;/p&gt;

&lt;p&gt;AI safety matters. Independent evaluation matters. Responsible development matters.&lt;/p&gt;

&lt;p&gt;But legitimate concerns about AI safety should not become a strategy that unintentionally slows American innovation, concentrates AI in a handful of companies, or assumes that technological competition between the United States and China can simply be negotiated away.&lt;/p&gt;

&lt;p&gt;We will get to more powerful AI.&lt;/p&gt;

&lt;p&gt;The question is how we get there—and which country builds the ecosystem around it.&lt;/p&gt;

&lt;p&gt;Start With What We Can Do Tomorrow&lt;/p&gt;

&lt;p&gt;The strongest part of Amodei’s proposal is also the simplest: independent evaluation.&lt;/p&gt;

&lt;p&gt;Do it.&lt;/p&gt;

&lt;p&gt;Anthropic has committed to giving independent evaluators employee-like access to examine its safety practices, training processes and incidents. Other frontier laboratories should seriously consider doing the same.&lt;/p&gt;

&lt;p&gt;This does not require China.&lt;/p&gt;

&lt;p&gt;It does not require an international treaty.&lt;/p&gt;

&lt;p&gt;It does not even require waiting for Congress.&lt;/p&gt;

&lt;p&gt;A company that believes independent evaluation makes its models safer can invite independent evaluators inside tomorrow.&lt;/p&gt;

&lt;p&gt;There is something powerful about companies voluntarily demonstrating that safety and technological progress do not have to be enemies.&lt;/p&gt;

&lt;p&gt;If these systems work, governments can eventually establish durable standards around them.&lt;/p&gt;

&lt;p&gt;The second proposal—coordination among frontier companies in democratic countries—is more complicated.&lt;/p&gt;

&lt;p&gt;America is a market economy. Competitors coordinating how quickly they develop technology raises legitimate questions about antitrust law, enforcement, standards and government authority. Amodei himself acknowledges that some forms of coordination would require government support.&lt;/p&gt;

&lt;p&gt;And government policy changes.&lt;/p&gt;

&lt;p&gt;Administrations change. Congress changes. Courts intervene. Companies enter and leave markets.&lt;/p&gt;

&lt;p&gt;If America decides that some form of coordinated pacing is necessary, the framework must be transparent, legally durable and democratically accountable. This is too important to depend indefinitely on voluntary agreements among a few CEOs.&lt;/p&gt;

&lt;p&gt;But the third proposal is where the problem becomes much larger.&lt;/p&gt;

&lt;p&gt;The China Problem&lt;/p&gt;

&lt;p&gt;Let’s be realistic about the AI competition.&lt;/p&gt;

&lt;p&gt;Many countries are doing important AI research, but the two dominant technological powers are the United States and China.&lt;/p&gt;

&lt;p&gt;China should never be underestimated.&lt;/p&gt;

&lt;p&gt;It has enormous engineering talent, industrial capacity, capital and increasingly sophisticated domestic technology. MacroPolo’s AI Talent Tracker found that researchers originating from China accounted for 47 percent of the world’s top-tier AI researchers in its 2022 dataset. At the same time, the United States remained the leading destination for elite AI talent.&lt;/p&gt;

&lt;p&gt;That combination should tell America something important.&lt;/p&gt;

&lt;p&gt;China has the people to compete.&lt;/p&gt;

&lt;p&gt;And China has made technological self-reliance and leadership strategic national objectives.&lt;/p&gt;

&lt;p&gt;The political systems are also fundamentally different.&lt;/p&gt;

&lt;p&gt;China is an authoritarian state with a highly centralized political structure. When Beijing identifies a technology as strategically important, it can coordinate national policy, infrastructure, financing, universities and industry in ways that are much harder in the United States.&lt;/p&gt;

&lt;p&gt;America is a democracy and a decentralized market economy. Congress debates. Courts intervene. States disagree. Companies compete. Citizens object. Administrations change.&lt;/p&gt;

&lt;p&gt;That friction is part of America.&lt;/p&gt;

&lt;p&gt;So America cannot simply imitate China’s strategy.&lt;/p&gt;

&lt;p&gt;It has to be more strategic.&lt;/p&gt;

&lt;p&gt;International dialogue is still worthwhile. The United States and China should communicate about catastrophic risks, military applications, autonomous systems and other areas where misunderstanding could be dangerous.&lt;/p&gt;

&lt;p&gt;But dialogue is different from dependence.&lt;/p&gt;

&lt;p&gt;An American AI strategy cannot depend on the assumption that every restriction will be interpreted identically, implemented identically and honored indefinitely by every participant.&lt;/p&gt;

&lt;p&gt;Amodei recognizes this problem himself. He writes that a global pacing agreement would require extremely strong verification because the incentive to defect could be enormous.&lt;/p&gt;

&lt;p&gt;That is precisely the problem.&lt;/p&gt;

&lt;p&gt;Before asking America to substantially slow technological development, we need to know what happens if someone else does not.&lt;/p&gt;

&lt;p&gt;Containment Is Not a Complete AI Strategy&lt;/p&gt;

&lt;p&gt;Export controls can make access to advanced computing more difficult and expensive. They can buy time and constrain capacity.&lt;/p&gt;

&lt;p&gt;But they are not a complete technological strategy.&lt;/p&gt;

&lt;p&gt;China has enormous incentives to develop domestic chips, improve software efficiency and find alternative architectures precisely because access to American technology is constrained.&lt;/p&gt;

&lt;p&gt;The same applies to model distillation and open weights.&lt;/p&gt;

&lt;p&gt;Once powerful models become software that can be downloaded, modified, fine-tuned and deployed around the world, containment becomes extraordinarily complicated.&lt;/p&gt;

&lt;p&gt;And trying to restrict foreign open models after developers and enterprises have integrated them into products can create another problem: America may end up weakening its own ecosystem.&lt;/p&gt;

&lt;p&gt;The better long-term question is not simply:&lt;/p&gt;

&lt;p&gt;How do we prevent developers from using Chinese AI?&lt;/p&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;p&gt;Why aren’t we giving them an American alternative they prefer?&lt;/p&gt;

&lt;p&gt;That is the strategic opening.&lt;/p&gt;

&lt;p&gt;Fight Open Source with Open Source&lt;/p&gt;

&lt;p&gt;Don’t fight open-source AI by trying to wish it away.&lt;/p&gt;

&lt;p&gt;Fight open source with open source.&lt;/p&gt;

&lt;p&gt;Make American open AI something developers around the world want to build on.&lt;/p&gt;

&lt;p&gt;Give developers models they can download, modify, fine-tune and deploy themselves. Give startups an alternative to paying frontier API prices for every token they generate. Give enterprises the ability to keep sensitive workloads inside their own infrastructure.&lt;/p&gt;

&lt;p&gt;Give researchers access.&lt;/p&gt;

&lt;p&gt;Give universities access.&lt;/p&gt;

&lt;p&gt;Let a college student fine-tune a model.&lt;/p&gt;

&lt;p&gt;Let a five-person startup build a company around one.&lt;/p&gt;

&lt;p&gt;Let an enterprise run one behind its firewall.&lt;/p&gt;

&lt;p&gt;Let researchers take one apart.&lt;/p&gt;

&lt;p&gt;Let millions of developers discover applications that no frontier laboratory could possibly anticipate.&lt;/p&gt;

&lt;p&gt;This is not an argument against frontier AI.&lt;/p&gt;

&lt;p&gt;It is an argument for both.&lt;/p&gt;

&lt;p&gt;Anthropic should thrive.&lt;/p&gt;

&lt;p&gt;OpenAI should thrive.&lt;/p&gt;

&lt;p&gt;Google should thrive.&lt;/p&gt;

&lt;p&gt;American open-source and open-weight AI should thrive too.&lt;/p&gt;

&lt;p&gt;One pushes the frontier of intelligence.&lt;/p&gt;

&lt;p&gt;The other distributes intelligence.&lt;/p&gt;

&lt;p&gt;America benefits from both.&lt;/p&gt;

&lt;p&gt;We should not weaken frontier laboratories to protect open source, and we should not weaken open source to protect the economics of frontier laboratories.&lt;/p&gt;

&lt;p&gt;Let them compete.&lt;/p&gt;

&lt;p&gt;The evidence already points in this direction. Stanford’s 2025 AI Index reported that the gap between leading open-weight and closed models narrowed sharply during 2024, while the cost of using capable models continued to fall. That means open models are not merely a philosophical alternative. They are becoming a practical competitive instrument.&lt;/p&gt;

&lt;p&gt;The United States government has recognized the same strategic logic. The U.S. AI Action Plan identifies leading American open-weight models as having geostrategic value and calls for supporting their adoption, particularly by startups, researchers and organizations that cannot send sensitive information to closed providers.&lt;/p&gt;

&lt;p&gt;That is exactly the opportunity.&lt;/p&gt;

&lt;p&gt;The objective should not simply be to have the world’s smartest model sitting behind an American API.&lt;/p&gt;

&lt;p&gt;The objective should be to have the world building with American AI.&lt;/p&gt;

&lt;p&gt;Affordability Is a Strategic Issue&lt;/p&gt;

&lt;p&gt;There is another reality that gets overlooked in debates about frontier models: economics.&lt;/p&gt;

&lt;p&gt;The most advanced AI systems are extraordinarily expensive to build and operate. Frontier APIs make that intelligence accessible without requiring customers to own the infrastructure, but using them at scale can still become expensive.&lt;/p&gt;

&lt;p&gt;For many startups, developers, universities, small businesses and organizations around the world, cost matters enormously.&lt;/p&gt;

&lt;p&gt;Open models change that equation.&lt;/p&gt;

&lt;p&gt;Organizations can choose their infrastructure. They can optimize inference. They can fine-tune models for specific workloads. They can run models locally when privacy requires it.&lt;/p&gt;

&lt;p&gt;That is not merely a technical preference.&lt;/p&gt;

&lt;p&gt;It affects adoption.&lt;/p&gt;

&lt;p&gt;If America wants its AI ecosystem to become the world’s default ecosystem, affordability matters.&lt;/p&gt;

&lt;p&gt;Make American intelligence powerful.&lt;/p&gt;

&lt;p&gt;But also make it accessible.&lt;/p&gt;

&lt;p&gt;The country whose technology millions of developers can afford to experiment with gains something that cannot easily be purchased later: an ecosystem.&lt;/p&gt;

&lt;p&gt;And ecosystems create strategic advantages that individual models cannot. Developers build skills around them. Startups form around them. Universities teach them. Enterprises integrate them. New tools and standards emerge from them.&lt;/p&gt;

&lt;p&gt;Once that network is established, it becomes difficult for a rival to displace.&lt;/p&gt;

&lt;p&gt;Don’t Lose the Public&lt;/p&gt;

&lt;p&gt;There is another constituency the AI industry cannot afford to ignore: ordinary Americans.&lt;/p&gt;

&lt;p&gt;Gallup found in 2025 that seven in ten Americans opposed construction of AI data centers in their local area, including 48 percent who strongly opposed them.&lt;/p&gt;

&lt;p&gt;People have legitimate concerns about electricity, water, land, environmental effects and local costs.&lt;/p&gt;

&lt;p&gt;Now consider the message the public sometimes hears from the AI industry.&lt;/p&gt;

&lt;p&gt;We need enormous amounts of electricity.&lt;/p&gt;

&lt;p&gt;We need enormous data centers.&lt;/p&gt;

&lt;p&gt;We need them quickly.&lt;/p&gt;

&lt;p&gt;And the technology they power might someday destroy humanity.&lt;/p&gt;

&lt;p&gt;Those messages do not fit comfortably together.&lt;/p&gt;

&lt;p&gt;Researchers should investigate catastrophic AI risks. Companies should red-team powerful systems. Independent evaluators should test them aggressively. Governments should prepare for credible threats.&lt;/p&gt;

&lt;p&gt;But serious safety research and apocalyptic public messaging are not the same thing.&lt;/p&gt;

&lt;p&gt;Repeatedly telling people that AI might “kill us all” risks doing a disservice to the technology, particularly when the industry simultaneously needs public support for enormous infrastructure expansion.&lt;/p&gt;

&lt;p&gt;People need to see why AI is worth building.&lt;/p&gt;

&lt;p&gt;Show them scientific discoveries.&lt;/p&gt;

&lt;p&gt;Show them better medicine.&lt;/p&gt;

&lt;p&gt;Show them new companies.&lt;/p&gt;

&lt;p&gt;Show them productivity.&lt;/p&gt;

&lt;p&gt;Show them education.&lt;/p&gt;

&lt;p&gt;Show them what an individual developer can create with capabilities that previously belonged only to enormous corporations.&lt;/p&gt;

&lt;p&gt;And perhaps most importantly, let people participate.&lt;/p&gt;

&lt;p&gt;AI should not feel like something being built behind closed doors by five companies while everyone else is asked to accept the consequences.&lt;/p&gt;

&lt;p&gt;Open models can help change that relationship. They give developers, researchers and businesses a stake in the technology rather than asking them merely to consume it.&lt;/p&gt;

&lt;p&gt;America’s Advantage Is America&lt;/p&gt;

&lt;p&gt;China can coordinate from the top.&lt;/p&gt;

&lt;p&gt;America can innovate from everywhere.&lt;/p&gt;

&lt;p&gt;That difference should be treated as an advantage, not a weakness.&lt;/p&gt;

&lt;p&gt;America has extraordinary universities, capital markets, entrepreneurs, researchers, immigrants, chip companies, cloud providers, frontier laboratories, startups and one of the world’s largest developer communities.&lt;/p&gt;

&lt;p&gt;Use all of it.&lt;/p&gt;

&lt;p&gt;Build the world’s best frontier models.&lt;/p&gt;

&lt;p&gt;Build the world’s best open models.&lt;/p&gt;

&lt;p&gt;Build the chips.&lt;/p&gt;

&lt;p&gt;Build the infrastructure.&lt;/p&gt;

&lt;p&gt;Make inference cheaper.&lt;/p&gt;

&lt;p&gt;Invest in safety.&lt;/p&gt;

&lt;p&gt;Invite independent evaluation.&lt;/p&gt;

&lt;p&gt;Attract the world’s best scientists.&lt;/p&gt;

&lt;p&gt;Give startups room to experiment.&lt;/p&gt;

&lt;p&gt;Give enterprises control.&lt;/p&gt;

&lt;p&gt;And give developers the freedom to build.&lt;/p&gt;

&lt;p&gt;Never underestimate China’s technological ambition. Never underestimate the resources and talent it can bring to this competition.&lt;/p&gt;

&lt;p&gt;But America’s answer should not be to become more like China.&lt;/p&gt;

&lt;p&gt;It should be to become more American.&lt;/p&gt;

&lt;p&gt;Open competition has produced extraordinary American technologies before. The internet, personal computing and the modern software industry all became more powerful because innovation spread beyond a small number of institutions.&lt;/p&gt;

&lt;p&gt;AI should not be different.&lt;/p&gt;

&lt;p&gt;The United States should not measure success only by whether an American company trains the most capable model. It should measure success by whether developers, researchers, startups and enterprises around the world choose American technology as the foundation for what they build next.&lt;/p&gt;

&lt;p&gt;Fight Open Source with Open Source&lt;/p&gt;

&lt;p&gt;AI development will not stop because one company slows down.&lt;/p&gt;

&lt;p&gt;It will not stop because one country regulates it.&lt;/p&gt;

&lt;p&gt;And it is increasingly difficult to imagine that knowledge this economically and strategically valuable will simply disappear.&lt;/p&gt;

&lt;p&gt;We will get there.&lt;/p&gt;

&lt;p&gt;That does not mean racing recklessly.&lt;/p&gt;

&lt;p&gt;It means recognizing that safety and progress are not opposites.&lt;/p&gt;

&lt;p&gt;Independent evaluation can coexist with rapid innovation.&lt;/p&gt;

&lt;p&gt;Frontier models can coexist with open models.&lt;/p&gt;

&lt;p&gt;Closed commercial systems can coexist with locally deployed intelligence.&lt;/p&gt;

&lt;p&gt;Safety can coexist with competition.&lt;/p&gt;

&lt;p&gt;Anthropic should succeed.&lt;/p&gt;

&lt;p&gt;American open source should succeed.&lt;/p&gt;

&lt;p&gt;Thousands of AI startups we have not heard of yet should succeed.&lt;/p&gt;

&lt;p&gt;And America should create an environment where all of them can.&lt;/p&gt;

&lt;p&gt;Because this competition is larger than any company’s valuation, any single model release or any benchmark.&lt;/p&gt;

&lt;p&gt;It is about which technological ecosystem the world chooses to build upon.&lt;/p&gt;

&lt;p&gt;The wrong response to open-source AI is to treat it as a threat that must be contained.&lt;/p&gt;

&lt;p&gt;The right response is to build something better, cheaper, safer and more useful.&lt;/p&gt;

&lt;p&gt;Fight open source with open source.&lt;/p&gt;

&lt;p&gt;Fight competition with competition.&lt;/p&gt;

&lt;p&gt;Make AI safer without making innovation inaccessible.&lt;/p&gt;

&lt;p&gt;Win the developers.&lt;/p&gt;

&lt;p&gt;Win the enterprises.&lt;/p&gt;

&lt;p&gt;Win the researchers.&lt;/p&gt;

&lt;p&gt;Win the public.&lt;/p&gt;

&lt;p&gt;And let the world build with American AI.&lt;/p&gt;

&lt;p&gt;We will get there—but America should make sure the world gets there with us.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>claude</category>
    </item>
    <item>
      <title>From RAG to Skill Function: A New Architecture for Enterprise AI Knowledge</title>
      <dc:creator>Prashanth Velidandi</dc:creator>
      <pubDate>Tue, 30 Jun 2026 20:58:30 +0000</pubDate>
      <link>https://dev.to/pmv_inferx/from-rag-to-skill-function-a-new-architecture-for-enterprise-ai-knowledge-4e76</link>
      <guid>https://dev.to/pmv_inferx/from-rag-to-skill-function-a-new-architecture-for-enterprise-ai-knowledge-4e76</guid>
      <description>&lt;p&gt;Enterprise AI knowledge systems have a scaling problem.&lt;/p&gt;

&lt;p&gt;RAG was the answer for years. Retrieve relevant chunks, feed them to the model, generate an answer. It works — until it doesn't. Retrieval misses, chunking breaks context, multi-document reasoning fails, pipelines grow complex. And as knowledge bases grow, the problems compound.&lt;/p&gt;

&lt;p&gt;Long context models offered a partial fix. Skip retrieval entirely, load the whole document. Better understanding, simpler architecture. But you're still paying full prefill cost on every query, and a single context window can't hold an entire enterprise knowledge base.&lt;/p&gt;

&lt;p&gt;We built something different. We're calling it &lt;strong&gt;Skill Function&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem with RAG
&lt;/h2&gt;

&lt;p&gt;RAG's core limitation isn't the retrieval algorithm — it's the fundamental architecture. The model can only reason over what gets retrieved. If the retrieval step misses something, the model never sees it. Wrong answer, not because the model can't reason, but because it never had the chance.&lt;/p&gt;

&lt;p&gt;The specific failure modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval quality limits answer quality.&lt;/strong&gt; Chunking breaks document structure, tables, cross-references, and long-range dependencies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-document reasoning is hard.&lt;/strong&gt; Missing even one relevant chunk leads to incomplete answers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production pipelines are complex.&lt;/strong&gt; Multiple retrieval stages, reranking, metadata filtering, hybrid search — each adds failure surface.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No deep document understanding.&lt;/strong&gt; RAG retrieves passages, not comprehension. It works for lookup, fails for analysis.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Long Context Solves Some of This
&lt;/h2&gt;

&lt;p&gt;Recent models with 128k–1M token context windows change the equation. Load the entire document, skip retrieval, let the model reason over everything.&lt;/p&gt;

&lt;p&gt;The improvements are real:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No retrieval errors&lt;/li&gt;
&lt;li&gt;Document structure preserved&lt;/li&gt;
&lt;li&gt;Better cross-section reasoning&lt;/li&gt;
&lt;li&gt;Simpler architecture — no vector database, no chunking pipeline&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But long context introduces new problems at scale:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cost.&lt;/strong&gt; Every query reprocesses the entire document, even when only a small section is relevant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency.&lt;/strong&gt; Prefill time on 100k+ tokens adds up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scalability.&lt;/strong&gt; Enterprise knowledge bases contain thousands of documents. No single context window holds all of it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Long context is a meaningful improvement for individual documents. It doesn't solve enterprise-scale knowledge.&lt;/p&gt;




&lt;h2&gt;
  
  
  Skill Function
&lt;/h2&gt;

&lt;p&gt;A Skill Function is a protected AI capability hosted as a callable endpoint in the cloud.&lt;/p&gt;

&lt;p&gt;The architecture has two components:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Document Skill&lt;/strong&gt; — A Document Skill specializes in a single document or related document collection. Instead of retrieving chunks, it loads the entire document into its context. Deep understanding, no retrieval, full structure preserved. Each Document Skill runs in its own isolated context — its own model, its own knowledge, its own reasoning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Orchestrator Skill&lt;/strong&gt; — An Orchestrator Skill organizes multiple Document Skills into a hierarchy. It maintains a summary of each sub-skill's expertise. When a query arrives, the Orchestrator determines which Document Skills are relevant, invokes them on demand, and synthesizes their results.&lt;/p&gt;

&lt;p&gt;An Orchestrator can invoke Document Skills or other Orchestrator Skills — forming a hierarchical knowledge tree that scales to arbitrarily large knowledge bases.&lt;/p&gt;




&lt;h2&gt;
  
  
  How It Works
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;User sends a query to the root Orchestrator Skill&lt;/li&gt;
&lt;li&gt;Orchestrator selects relevant sub-skills based on their summaries&lt;/li&gt;
&lt;li&gt;Selected sub-skills process the query using full document context&lt;/li&gt;
&lt;li&gt;Orchestrator aggregates results and produces the final answer&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Only the skills relevant to the query execute. Everything else stays idle. Context stays clean.&lt;/p&gt;

&lt;p&gt;Instead of forwarding the entire conversation history, each Skill Function receives only the current query and a concise summary of relevant conversation history. This eliminates context pollution and keeps each skill focused.&lt;/p&gt;




&lt;h2&gt;
  
  
  Skill Function vs RAG
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;RAG&lt;/th&gt;
&lt;th&gt;Skill Function&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge Unit&lt;/td&gt;
&lt;td&gt;Document chunks&lt;/td&gt;
&lt;td&gt;Specialized Document Skills&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge Access&lt;/td&gt;
&lt;td&gt;Vector retrieval&lt;/td&gt;
&lt;td&gt;AI skill routing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Document Understanding&lt;/td&gt;
&lt;td&gt;Partial chunks&lt;/td&gt;
&lt;td&gt;Complete documents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-document Reasoning&lt;/td&gt;
&lt;td&gt;Retrieve multiple chunks&lt;/td&gt;
&lt;td&gt;Coordinate multiple skills&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scalability&lt;/td&gt;
&lt;td&gt;Larger vector databases&lt;/td&gt;
&lt;td&gt;Hierarchical skill tree&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context Usage&lt;/td&gt;
&lt;td&gt;Retrieved chunks&lt;/td&gt;
&lt;td&gt;Only relevant skills execute&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Engineering&lt;/td&gt;
&lt;td&gt;Chunking, embeddings, retrieval tuning&lt;/td&gt;
&lt;td&gt;Skill organization and orchestration&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;RAG treats enterprise knowledge as a searchable database. Skill Function treats it as a network of specialized AI experts coordinated through hierarchical orchestration.&lt;/p&gt;




&lt;h2&gt;
  
  
  Skill Function vs Claude-style Skills
&lt;/h2&gt;

&lt;p&gt;Claude-style skills (SKILL.md files) load multiple skills into a shared context window. More skills means more context competition, slower responses, and a hard ceiling on how many skills can run together.&lt;/p&gt;

&lt;p&gt;Skill Function assigns a dedicated context per skill. Each Document Skill operates independently with its own long-context environment. Knowledge doesn't compete for a shared window.&lt;/p&gt;

&lt;p&gt;Because contexts are isolated, skills compose recursively without exhausting a global context:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A Document Skill can be called by an Orchestrator Skill&lt;/li&gt;
&lt;li&gt;An Orchestrator Skill can call other Orchestrator Skills&lt;/li&gt;
&lt;li&gt;The hierarchy scales to arbitrary depth&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Claude-style Skills:&lt;/strong&gt; One global context → simplicity, but limited scaling and context competition.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Skill Function:&lt;/strong&gt; Many isolated contexts → hierarchical composition, scalable knowledge depth, controllable execution.&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Looks Like in Practice
&lt;/h2&gt;

&lt;p&gt;Upload a PDF. We automatically convert it into skill experts — each section becomes its own Document Skill with its own model, context, and reasoning.&lt;/p&gt;

&lt;p&gt;On our platform, you can combine those experts into an Orchestrator Skill. Skills call other skills. Your query automatically reaches the right expert.&lt;/p&gt;

&lt;p&gt;The whole thing is exposed as an MCP server.&lt;/p&gt;

&lt;p&gt;For example: take your company knowledge across legal, finance, HR, and product — turn each into a Document Skill, combine them into one Orchestrator, and query across your entire company knowledge base. The right expert answers every time.&lt;/p&gt;

&lt;p&gt;No vector database. No embeddings. No retrieval step. No document size limit. 70-90% cheaper than loading everything into one context window.&lt;/p&gt;




&lt;h2&gt;
  
  
  Try It
&lt;/h2&gt;

&lt;p&gt;We're testing this now. Try it free at &lt;a href="https://inferx.net" rel="noopener noreferrer"&gt;inferx.net&lt;/a&gt;.&lt;br&gt;
or reach out to me: &lt;a href="mailto:prashanth@inferx.net"&gt;prashanth@inferx.net&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Happy to answer questions in the comments.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>Why AI Skills Are Broken - and How We Fixed the Architecture</title>
      <dc:creator>Prashanth Velidandi</dc:creator>
      <pubDate>Fri, 19 Jun 2026 14:52:54 +0000</pubDate>
      <link>https://dev.to/pmv_inferx/why-ai-skills-are-broken-and-how-we-fixed-the-architecture-4om7</link>
      <guid>https://dev.to/pmv_inferx/why-ai-skills-are-broken-and-how-we-fixed-the-architecture-4om7</guid>
      <description>&lt;p&gt;&lt;strong&gt;The promise of AI skills&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When Claude introduced SKILL.md files in late 2025, it changed how developers think about AI agents. Instead of hardcoding every instruction into a system prompt, you write a skill file, drop it in a folder, and your agent knows what to do. Simple. Elegant. Powerful.&lt;/p&gt;

&lt;p&gt;The ecosystem exploded. skills.sh now has 600,000+ skills. Developers are building, sharing, and shipping faster than ever.&lt;/p&gt;

&lt;p&gt;But there's a structural problem nobody is talking about.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;The three failures of local skills&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The monolithic model tax&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When you configure an agent today — Claude Code, Cursor, OpenClaw — you pick one model. That model handles everything. Summarizing a routine email. Reviewing complex legal contracts. Generating pricing strategy. Same model. Same price.&lt;/p&gt;

&lt;p&gt;80% of your context window is consumed by tool definitions, system prompts, and conversation history. Not your actual task. You're paying premium prices for infrastructure that delivers minimal user value.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The security problem nobody talks about&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Local skills run with full system privileges. They can read your files, invoke shell commands, access cloud credentials, and open network connections.&lt;/p&gt;

&lt;p&gt;A comprehensive study of 31,132 publicly available skills found that 26.1% contain at least one security vulnerability. Skills bundling executable scripts are 2.12x more likely to contain vulnerabilities.&lt;/p&gt;

&lt;p&gt;One malicious skill — or even a benign skill manipulated through prompt injection — can exfiltrate SSH keys, access cloud credentials, or delete critical data. The attack surface is enormous.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The context bottleneck&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Complex workflows, deep domain knowledge, and large reference materials cannot fit in a single context window. Even with a 1 million token window, models systematically lose information in the middle.&lt;/p&gt;

&lt;p&gt;The more skills you add, the worse it gets. Attention scatters. Response times increase. Accuracy drops with every turn.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Rethinking the architecture&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The root cause of all three problems is the same: skills live inside the agent's shared context window, running on your local machine with full system privileges.&lt;/p&gt;

&lt;p&gt;What if skills lived outside the context window entirely?&lt;/p&gt;

&lt;p&gt;That's what we built. &lt;strong&gt;Skill Function&lt;/strong&gt; — a cloud-native Skill-as-a-Service platform.&lt;/p&gt;

&lt;p&gt;Instead of loading a skill into the agent's context, Skill Function moves each skill to the cloud as an independent callable service.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST api.inferx.net/skills/saas-pricing

{
  "input": "B2B SaaS, $50 ACV, PLG motion, 3 tiers"
}

→ Expert output. 195ms. Instructions never leave the platform.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;p&gt;&lt;strong&gt;How it works&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Right model for each skill&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every Skill Function is bound to a pre-selected model chosen by the skill author. A simple classification skill uses a 7B model. A complex code-review skill uses a 70B model.&lt;/p&gt;

&lt;p&gt;When your agent calls the Skill Function, it no longer forces every task through your expensive flagship model. Each task gets the model it actually needs — no more, no less.&lt;/p&gt;

&lt;p&gt;Result: 70-90% lower inference cost for mixed workloads.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dedicated clean context&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Each Skill Function handles one task at a time. Input is simple: the user's request plus a short summary of relevant history. No unrelated tool definitions. No other skills' prompts. No accumulated conversation history.&lt;/p&gt;

&lt;p&gt;The skill runs in its own isolated context window — clean, focused, free of cross-talk. Performance does not degrade with every turn.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Skills call other skills&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One Skill Function can call other Skill Functions, just like traditional function calls in software. Complex workflows decompose into a directed graph of skill calls.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;orchestrator → legal-reviewer    (if legal task)
orchestrator → pricing-strategist (if pricing task)  
orchestrator → code-reviewer      (if code task)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This breaks the context barrier entirely. Not fragmentation — composition.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zero local execution&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A Skill Function is a pure knowledge skill. It cannot call tools directly — no curl, no bash, no local file access, no network egress.&lt;/p&gt;

&lt;p&gt;This eliminates the entire local attack surface. Even a successful prompt injection can only influence the skill's output text. It cannot trigger system-level actions.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;MCP-native discovery&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Skill Function exposes a standard MCP tool-calling interface. When you subscribe to a cloud skill, it automatically appears in your local agent through MCP tool discovery — just like a locally installed tool.&lt;/p&gt;

&lt;p&gt;No skill files to download. No environment variables to set. No local deployment. The agent simply sees a new tool and calls it.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;The result&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Local Skills&lt;/th&gt;
&lt;th&gt;Skill Function&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;One flagship for everything&lt;/td&gt;
&lt;td&gt;Right model per task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context&lt;/td&gt;
&lt;td&gt;Shared, fills up&lt;/td&gt;
&lt;td&gt;Isolated per skill&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security&lt;/td&gt;
&lt;td&gt;Full local privileges&lt;/td&gt;
&lt;td&gt;Zero local execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Composition&lt;/td&gt;
&lt;td&gt;File references&lt;/td&gt;
&lt;td&gt;Function calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Discovery&lt;/td&gt;
&lt;td&gt;Manual install&lt;/td&gt;
&lt;td&gt;MCP auto-discovery&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;p&gt;&lt;strong&gt;Try it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;50+ Skill Functions available today across marketing, design, engineering, finance, and research. Or import your own SKILL.md and run it as a protected callable endpoint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;inferx.net&lt;/strong&gt; — free to start.&lt;/p&gt;

&lt;p&gt;Full technical white paper: &lt;a href="https://inferx.net/skill-function-whitepaper" rel="noopener noreferrer"&gt;https://inferx.net/skill-function-whitepaper&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;For questions: &lt;a href="mailto:prashanth@inferx.net"&gt;prashanth@inferx.net&lt;/a&gt; · @InferXai&lt;/em&gt;&lt;/p&gt;




</description>
      <category>ai</category>
      <category>agentskills</category>
      <category>agents</category>
      <category>langchain</category>
    </item>
    <item>
      <title>RAG became the default answer for private knowledge access. We asked a different question: what if context didn’t need to be repeatedly retrieved at all? Persistent KV cache changed the economics completely.</title>
      <dc:creator>Prashanth Velidandi</dc:creator>
      <pubDate>Tue, 26 May 2026 12:21:50 +0000</pubDate>
      <link>https://dev.to/pmv_inferx/rag-became-the-default-answer-for-private-knowledge-access-we-asked-a-different-question-what-if-45m9</link>
      <guid>https://dev.to/pmv_inferx/rag-became-the-default-answer-for-private-knowledge-access-we-asked-a-different-question-what-if-45m9</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/pmv_inferx/we-replaced-our-rag-pipeline-with-persistent-kv-cache-heres-what-we-found-7cl" class="crayons-story__hidden-navigation-link"&gt;We Replaced Our RAG Pipeline With Persistent KV Cache. Here's What We Found.&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/pmv_inferx" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3382940%2F7c22cb83-8e48-49dd-8ec0-2aa7fb44b3e5.jpg" alt="pmv_inferx profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/pmv_inferx" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Prashanth Velidandi
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Prashanth Velidandi
                
              
              &lt;div id="story-author-preview-content-3731263" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/pmv_inferx" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3382940%2F7c22cb83-8e48-49dd-8ec0-2aa7fb44b3e5.jpg" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Prashanth Velidandi&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/pmv_inferx/we-replaced-our-rag-pipeline-with-persistent-kv-cache-heres-what-we-found-7cl" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;May 23&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/pmv_inferx/we-replaced-our-rag-pipeline-with-persistent-kv-cache-heres-what-we-found-7cl" id="article-link-3731263"&gt;
          We Replaced Our RAG Pipeline With Persistent KV Cache. Here's What We Found.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/rag"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;rag&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/serverless"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;serverless&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/machinelearning"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;machinelearning&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/webdev"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;webdev&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/pmv_inferx/we-replaced-our-rag-pipeline-with-persistent-kv-cache-heres-what-we-found-7cl" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/fire-f60e7a582391810302117f987b22a8ef04a2fe0df7e3258a5f49332df1cec71e.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;2&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/pmv_inferx/we-replaced-our-rag-pipeline-with-persistent-kv-cache-heres-what-we-found-7cl#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            3 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>We Replaced Our RAG Pipeline With Persistent KV Cache. Here's What We Found.</title>
      <dc:creator>Prashanth Velidandi</dc:creator>
      <pubDate>Sat, 23 May 2026 08:34:13 +0000</pubDate>
      <link>https://dev.to/pmv_inferx/we-replaced-our-rag-pipeline-with-persistent-kv-cache-heres-what-we-found-7cl</link>
      <guid>https://dev.to/pmv_inferx/we-replaced-our-rag-pipeline-with-persistent-kv-cache-heres-what-we-found-7cl</guid>
      <description>&lt;p&gt;RAG has become the default answer for giving LLMs access to private knowledge. And for good reason — it works. But after running it in production we kept hitting the same wall. Not retrieval accuracy. The operational tax.&lt;/p&gt;

&lt;p&gt;Re-embedding on data changes. Chunking drift. Retrieval misses on edge cases. Pipeline failures at 2am. The vector database that needs babysitting.&lt;/p&gt;

&lt;p&gt;So we ran an experiment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Hypothesis&lt;/strong&gt;&lt;br&gt;
What if instead of chunking, embedding, and retrieving — we just loaded the full document into the LLM context, cached the KV state persistently, and reused it across every query?&lt;/p&gt;

&lt;p&gt;No retrieval step. No embedding pipeline. No vector database. Just the model with full document context, warm and ready.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How It Works&lt;/strong&gt;&lt;br&gt;
The core idea is simple. When an LLM processes a prompt it generates a key-value attention cache — the internal representation of everything it has read. Normally this cache is transient. It lives in VRAM during the request and disappears after.&lt;br&gt;
We persist it.&lt;br&gt;
The initialization prompt — your document — gets processed once. The resulting KV cache gets stored externally and indexed to that document. Every subsequent query retrieves that cached state and appends the user query. The model never recomputes the document. Ever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The math&lt;/strong&gt;:&lt;br&gt;
KV_init = LLM.prefill(document)&lt;br&gt;
KV_store[document_id] = KV_init&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;# On every query:&lt;/strong&gt;&lt;br&gt;
KV_full = KV_store[document_id] + LLM.prefill(query)&lt;br&gt;
output = LLM.decode(KV_full)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What We Found&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Answer quality improved.&lt;br&gt;
No retrieval misses are possible when the full document is in context. The model has read everything. It doesn't guess which chunks are relevant — it knows the whole document. For complex multi-part questions that span different sections this is a significant improvement over chunked retrieval.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Updates became trivial.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Document changes? Re-run the prefill, store the new KV cache. Minutes not hours. No re-embedding pipeline. No re-indexing. No retrieval regression testing. Just regenerate and deploy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Operational complexity dropped.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No embedding model to maintain. No vector database to monitor. No chunking strategy to tune. No retrieval quality metrics to track. The surface area for things to break quietly got dramatically smaller.&lt;br&gt;
Latency on warm cache is effectively instant.&lt;/p&gt;

&lt;p&gt;When the KV state is already loaded the query just appends and generates. No retrieval hop, no context injection latency.&lt;br&gt;
The Honest Tradeoffs&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context window is the ceiling.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Current limit is around 120k tokens — roughly 200-300 pages. Works well for focused documents. For large corpora you need a routing layer to select the right cache per query. You've pushed the retrieval problem up one level — instead of retrieving chunks you're selecting a cache. Simpler problem but not zero.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cold cache restore adds latency.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first query after a cache restore pays a latency cost. For strict SLA requirements this matters. Warm cache is instant. Cold restore depends on your infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Initial prefill costs more than embedding.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Running a full forward pass on a large document costs more compute than embedding it. The economics work when query volume is high enough to amortize that cost. Low query, high update frequency — RAG still wins.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where This Wins&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This approach is clearly better when:&lt;/p&gt;

&lt;p&gt;You have a focused, structured document — legal contract, compliance policy, product manual, technical spec&lt;br&gt;
Query volume is high relative to update frequency&lt;br&gt;
Full context comprehension matters more than breadth&lt;br&gt;
You want to eliminate pipeline maintenance entirely&lt;br&gt;
Privacy matters — no document chunks sent to embedding APIs&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where RAG Still Wins&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Very large document collections where context limits apply&lt;br&gt;
Highly dynamic data that changes multiple times per day&lt;br&gt;
When you genuinely don't know which document is relevant at query time&lt;br&gt;
Low query volume where prefill cost doesn't amortize&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What We're Building&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We've been running this in production at InferX as part of our Sovereign Endpoints™ infrastructure. The persistent KV cache layer sits on top of our GPU snapshotting architecture — which is what makes the cold cache restore fast enough to be practical.&lt;br&gt;
We're now opening a limited beta for teams who want to test this on real workloads. Particularly interested in legal, compliance, finance, and developer tooling use cases.&lt;br&gt;
If you're running RAG in production and want to run a head-to-head comparison — we'd love to work with you.&lt;/p&gt;

&lt;p&gt;🎬 Demo dropping in 2 days — follow to see it first.&lt;br&gt;
&lt;/p&gt;
&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
      &lt;div class="c-embed__body flex items-center justify-between"&gt;
        &lt;a href="https://inferx.net/" rel="noopener noreferrer" class="c-link fw-bold flex items-center"&gt;
          &lt;span class="mr-2"&gt;inferx.net&lt;/span&gt;
          

        &lt;/a&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;


</description>
      <category>rag</category>
      <category>serverless</category>
      <category>machinelearning</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
