--
title: "Your Old Tweets May Be Feeding AI Training Data"
description: "Large models need astronomical amounts of text."
tags: ["tweets", "feeding", "training", "data"]
canonical_url: https://digital-footprint-health.shop/blog/ai-scraping-old-tweets-training-data
Large models need astronomical amounts of text. Public web pages, books, forums, and open datasets are all sources, and public social posts are high-value corpus because of their volume and real conversational context.
Where your old tweets may land. - The public posts themselves. If set to public, the odds of being scraped are high, often without separate consent.
Can ordinary users retract. What you can do first is reduce the source: set old tweets private or delete them to lower the chance of future scraping; for what platforms already grabbed, watch for objection, deletion, or opt-out mechanisms they and regulators provide, though effectiveness varies by jurisdiction. The more practical move is to clear the truly dangerous hard-private facts first, those matter more than being read by a model.
A often-ignored angle: training data also shapes model bias. The emotional and extreme expressions in public tweets subtly enter the model tone. A heated line you posted years ago may reappear, in another form, inside some AI answer.
If you want to actively opt out of training sets. There is no one-click opt-out yet, but you can stack a few moves: batch-privatize or delete historical public posts to lower future scrape odds; watch platform-published objection and deletion channels and object to clearly violating scrapes; for already-trained models, influence is limited, so focus on source control. Treat it as a long-term action like cleaning hard-private facts, more practical than hoping for a magic switch someday.
The part worth keeping: Since public scraping became a flashpoint, several platforms updated terms, added opt-out or objection forms, or restricted bulk access through their APIs. Full write-up is on the source blog.
Top comments (0)