AI Is Eating the Web — and Our Digital Memory With It
Meta Description: As AI eats the web, the internet's collective memory is disappearing. Here's what's being lost, why it matters, and what you can do to preserve it.
TL;DR: AI systems are consuming, summarizing, and effectively replacing vast swaths of the open web. In the process, the original sources — blogs, forums, niche websites, and community archives — are dying from traffic starvation. The internet's collective memory is eroding in real time. This article explains what's happening, what's already been lost, and what individuals and organizations can do about it.
The Web Is Being Quietly Hollowed Out
Something strange happened between 2023 and 2026: the internet got quieter.
Not in terms of data volume — more bytes are flowing than ever. But the ecosystem of human-written, independently published, community-maintained content that made the web genuinely useful is collapsing at an alarming rate. As AI eats the web, the internet's collective memory is disappearing, and most people haven't noticed yet because the replacement — AI-generated summaries, chatbot answers, zero-click search results — feels like information.
It isn't. Not entirely.
Between 2023 and mid-2026, independent web traffic monitoring firms documented a 40–60% decline in organic search referrals to small and mid-sized publishers. The Wayback Machine reported a 34% spike in "save page now" requests in 2025 alone, as users began panic-archiving content they feared would vanish. Stack Overflow, once the backbone of developer knowledge-sharing, saw active question volume drop by more than half. Dozens of niche forums — communities built over decades around everything from vintage synthesizers to rare plant cultivation — shut down after losing the search traffic that justified their hosting costs.
This isn't a natural evolution of the web. It's a structural collapse driven by a specific economic mechanism: AI systems are trained on human-created content, then deployed to answer questions that would previously have sent users to that content. The original sources get no traffic. No traffic means no revenue. No revenue means the source disappears. And when it disappears, future AI systems have one less place to learn from.
It's a loop. And it's accelerating.
How AI Is Consuming the Web's Institutional Knowledge
The Training Data Problem
Every major large language model — GPT-series models, Claude, Gemini, LLaMA variants — was trained on enormous crawls of the public internet. This content includes Wikipedia, Reddit threads, personal blogs, academic papers, cooking websites, travel guides, hobbyist wikis, and millions of other sources that represent decades of accumulated human knowledge.
The irony is brutal: the richer and more useful a piece of web content was, the more likely it was to be scraped, weighted heavily in training, and used to power a system that now replaces it in search results.
[INTERNAL_LINK: how large language models are trained on web data]
Zero-Click Search and the Answer Economy
Google's AI Overviews (formerly Search Generative Experience), Microsoft's Copilot integration in Bing, and Perplexity's answer engine have fundamentally changed how people interact with search results. Studies from early 2026 show that approximately 65% of searches now end without a single click to an external website — up from roughly 50% in 2023.
For a recipe blog, a local news outlet, or a technical documentation site, this is existential. Their content is being used — just not by them.
The Forum Collapse
Forums deserve special attention because they represent something irreplaceable: longitudinal human conversation. A thread from 2009 on a woodworking forum isn't just an answer — it's a record of how people thought about a problem at a specific moment in time, with all the uncertainty, debate, and lived experience that entails.
Since 2024:
- Reddit implemented aggressive API pricing, fracturing its third-party ecosystem and driving many communities toward private or alternative platforms
- Google Groups archived and effectively killed thousands of mailing list archives
- Fandom wikis began consolidating and deleting "low-traffic" pages en masse
- Dozens of phpBB and vBulletin forums shut down entirely after losing ad revenue
When these communities die, the knowledge doesn't migrate cleanly to AI systems. It degrades. Context collapses. Nuance disappears. What remains is a flattened, averaged version of what humans once knew.
What's Actually Being Lost
Niche Expertise That Doesn't Scale
Ask an AI how to replace the capacitors on a 1987 Roland Juno-106 synthesizer. You'll get a plausible-sounding answer. Ask someone who spent fifteen years on a dedicated synth repair forum, and you'll get the right answer — including the specific capacitor brand that fails at high temperatures, the forum post from 2011 where a technician discovered the workaround, and the warning about a batch of counterfeit parts circulating in 2019.
That layered, experiential, community-verified knowledge lives in forums. And forums are dying.
Local and Hyperlocal Information
Local news has been in crisis for years, but AI has accelerated the collapse. When someone searches "best ramen in [small city]" and gets an AI summary scraped from a review site that itself scraped from a blog that hasn't been updated since 2022, the information chain is three steps removed from reality — and nobody's checking.
[INTERNAL_LINK: local news collapse and AI-generated content]
Historical Web Content
The internet of the 1990s and early 2000s was a genuinely different place — chaotic, personal, experimental. GeoCities pages, early web forums, Usenet archives, personal homepages. Much of this content exists only in the Internet Archive Wayback Machine — a nonprofit that is perpetually underfunded and has faced multiple legal challenges from publishers attempting to restrict its archiving activities.
In 2025, a federal court ruling partially restricted the Wayback Machine's ability to serve archived copies of certain copyrighted materials. The long-term implications are still being litigated, but the precedent is chilling.
The "Vibes" of the Early Web
This sounds abstract, but it matters: the early web had personality. Individual expression. Weirdness. A teenager's hand-coded fan page about their favorite band carried information that no database could fully capture — enthusiasm, context, cultural moment. As AI eats the web, the internet's collective memory is disappearing not just as data, but as texture.
The Economics of Digital Extinction
Here's the mechanism in plain terms:
| Stage | What Happens | Effect on Web Ecosystem |
|---|---|---|
| 1. AI Training | Models scrape and learn from web content | Content creators receive nothing |
| 2. AI Deployment | Users get answers from AI instead of clicking links | Traffic to source sites drops 40-60% |
| 3. Revenue Collapse | Ad revenue, affiliate income, subscriptions dry up | Publishers cut content, reduce staff |
| 4. Site Closure | Economically unviable sites shut down | Content permanently lost |
| 5. Training Degradation | Future AI models have less high-quality data | AI outputs become less accurate over time |
This is sometimes called the "ouroboros problem" — AI consuming the very ecosystem it depends on for future improvement.
What's Being Done (And Whether It's Working)
Digital Preservation Efforts
The Internet Archive remains the most critical institution in this space. Its Wayback Machine has archived over 800 billion web pages, and its various collections preserve software, music, books, and video. If you value the web's memory, Internet Archive deserves your financial support — it operates as a nonprofit and is perpetually under-resourced relative to its mission.
Assessment: Genuinely essential. Irreplaceable. Chronically underfunded. The legal threats it faces are the single biggest risk to web preservation today.
Archive Team is a volunteer collective that specifically targets at-risk sites — when a platform announces it's shutting down, Archive Team mobilizes to crawl and preserve it. They've saved content from GeoCities, Vine, Google+, and hundreds of other defunct platforms.
Assessment: Scrappy, effective, and genuinely heroic. Limited by volunteer capacity and bandwidth costs.
Emerging Technical Solutions
IPFS (InterPlanetary File System) offers a decentralized approach to content storage where files are addressed by their content rather than their location. If a file exists anywhere on the IPFS network, it's accessible — no single point of failure.
Assessment: Technically promising, but adoption remains niche. Content on IPFS still needs to be "pinned" by someone paying for storage, which recreates some of the economic problems it's meant to solve.
Solid (Social Linked Data), Tim Berners-Lee's project for decentralized personal data storage, continues to develop slowly. It's a long-term play, not a near-term solution.
Policy and Legal Approaches
The EU's AI Act (fully enforced from 2025) includes provisions requiring AI companies to disclose training data sources, with some compensation mechanisms still being negotiated. In the United States, the situation is murkier — several high-profile copyright lawsuits against AI companies are working through the courts, but legislative action remains stalled.
Some publishers have negotiated direct licensing deals with AI companies. The New York Times, Associated Press, and several academic publishers have reached agreements with OpenAI and Google. But these deals primarily benefit large institutions — the independent blogger, the niche forum, the local news site gets nothing.
What You Can Do Right Now
This isn't a problem only governments and corporations can solve. Individual action matters.
For Readers and Internet Users
-
Use the Wayback Machine proactively. If you find a valuable page, save it:
web.archive.org/save/[URL] - Support independent publishers directly — subscribe, donate, buy merchandise. Remove the middleman
- Click through to sources even when AI gives you an answer. Traffic signals value
- Download content you care about — videos, articles, forum threads — using tools like yt-dlp (for video) or SingleFile (a browser extension for saving complete web pages)
For Content Creators and Publishers
- Self-host your archives — don't rely solely on third-party platforms
- Use structured data and semantic markup to make your content more attributable and traceable
- Maintain an email list — it's the one channel AI can't intermediate
- Export your community data regularly — if you run a forum, schedule monthly database exports
- Consider static site generators like Hugo or Eleventy — they produce portable, archivable HTML that survives platform changes
For Organizations and Institutions
- Fund digital preservation nonprofits
- Advocate for mandatory AI training data transparency
- Implement robots.txt protections where appropriate (with the caveat that compliance is voluntary and inconsistently honored)
- Partner with library systems and academic institutions on archiving initiatives
[INTERNAL_LINK: how to preserve your website content long-term]
Key Takeaways
- AI systems are trained on web content, then deployed to replace the traffic that made that content economically viable — a self-destructive loop
- Zero-click search has reached ~65% of queries, devastating independent publishers who depend on search traffic
- Forums, local news, and niche expertise sites are the most vulnerable — and the hardest to replace
- The Internet Archive is the most critical preservation institution and needs both funding and legal protection
- Individual actions matter: save pages, click through to sources, support independent publishers directly
- Policy solutions are emerging but slow — EU AI Act provisions and ongoing copyright litigation may reshape incentives, but not quickly enough for sites dying today
- The loss isn't just data — it's context, texture, and the record of how humans actually thought and communicated
The Stakes Are Higher Than They Look
When we talk about AI eating the web, it's tempting to frame it as a business story — publishers losing revenue, platforms shutting down. But the deeper issue is epistemic.
The internet became humanity's most comprehensive external memory. It preserved arguments, mistakes, discoveries, jokes, grief, and expertise across languages, cultures, and decades. That record is how future humans — and future AI systems — will understand what it was like to be alive right now.
As AI eats the web, the internet's collective memory is disappearing not just as an economic problem, but as a civilizational one. We are allowing a system trained on the richness of human expression to gradually consume and replace that richness with averaged, flattened, unattributed summaries.
The web doesn't have to die this way. But preventing it requires treating digital preservation as infrastructure — as essential as roads or electrical grids — rather than a niche concern for librarians and archivists.
Start saving things. Start clicking through. Start paying for the content you value. The web you remember is being archived right now, whether you participate or not.
📣 Take Action Today
If this article changed how you think about the web's future, here's your immediate next step: go to Internet Archive and make a donation, even a small one. Then install SingleFile in your browser and start saving the pages that matter to you. These two actions take less than ten minutes and directly contribute to preserving the web's memory.
Share this article with someone who still thinks the internet is fine. It isn't — but it's fixable, if enough people pay attention.
Frequently Asked Questions
Q: Is AI actually "eating" the web, or is this just normal technological change?
A: It's both, but the scale and speed matter. Every technology shift displaces previous ones — search engines changed how people found information, social media changed how they shared it. But the current AI transition is unique because it directly monetizes existing web content without compensating creators, while simultaneously reducing the traffic that made creating that content economically viable. The loop is structurally different from previous shifts.
Q: Can AI companies just train on new, AI-generated content instead?
A: This is called "model collapse," and it's a documented problem. When AI models are trained primarily on AI-generated content, output quality degrades over successive generations — the models lose the nuance, diversity, and accuracy that came from training on genuine human expression. The web's human-created content isn't just training data; it's the quality signal that keeps AI systems grounded in reality.
Q: What happened to all the content from platforms that have already shut down?
A: It varies. Some of it was archived by the Wayback Machine or Archive Team before shutdown. Some was exported by users who anticipated the closure. A significant portion is simply gone — particularly content from platforms that shut down suddenly, restricted archiving before closure, or existed before systematic preservation efforts were in place. Vine, early MySpace, Google+, and thousands of smaller platforms represent genuine permanent losses.
Q: Does using robots.txt actually protect my content from AI scraping?
A: Partially. Major AI companies have committed to honoring robots.txt directives, and most do — for future crawls. Content already scraped before you added protections is likely already in training datasets. Additionally, compliance is voluntary; smaller or less scrupulous AI developers may ignore directives entirely. robots.txt is a useful tool but not a complete solution.
Q: Is there any realistic scenario where the web's ecosystem recovers?
A: Yes, but it requires multiple things happening simultaneously: legal frameworks that require AI companies to compensate content creators (similar to music licensing models), technical standards that make content attribution and compensation automatable, renewed cultural investment in independent publishing, and robust public funding for digital preservation institutions. None of these are impossible — but none are happening fast enough to prevent significant losses in the near term.
Top comments (0)