DEV Community

Cover image for Ask.com is dead. 197,000 of its questions are still in the archive.
Pennyforge
Pennyforge

Posted on

Ask.com is dead. 197,000 of its questions are still in the archive.

Dated data post · 2026-10-07 · all numbers first-hand, snapshot time-stamped

Ask.com closed on May 1, 2026. The farewell page is blunt: "As IAC continues to sharpen its focus, we have made the decision to discontinue our search business, which includes Ask.com." — © 2026 IAC Inc. I verified the farewell page live, and the question pages behind it: random /question/ paths now 404 on a GitHub Pages shell. Twenty years of self-service Q&A, dark in a week, with no content license released anywhere I could find.

So I went to the one place that definitely had copies: the Internet Archive. This post is the result — a dated preservation snapshot of the corpus, the method, the licensing status, and the honest limits.

What's in the archive

The extraction pool came from Wayback's CDX index: 245,834 distinct archived question URLs (www.ask.com/question/…, one earliest capture per question). After a week of steady extraction:

Metric (snapshot 2026-10-07 00:30Z) Value
Distinct question URLs in the pool 245,834
Distinct URLs with ≥1 recorded attempt 205,015
Captures fetched OK 202,494 rows (202,046 distinct URLs, 98.2% of rows)
With a full answer 197,226 (97.4% of fetched; 196,787 distinct answered questions)
Answer length median 293 chars, mean 1,121, max 49,035
Capture years 99.0% from 2013–2014 (126,839 + 73,698); 1,878 from 2012; a few dozen 2015–23

Two things jump out.

The surviving corpus is a 2013–14 time capsule, not a 20-year history. The Wayback captures cluster almost entirely in the site's Q&A era. There is essentially no pre-2012 material. If you heard "30 years of Ask.com went dark," the precise version is: the 2013–14 question era went dark. I'd rather say that than the rounder number.

The corpus is deep in breadth, shallow in time. By construction I took the earliest capture per question, so each question appears once. Most questions (per my earlier sampling: 62–63% of a 120-URL sample) have exactly one capture in the archive at all. No temporal depth — but for 196,000+ distinct questions, that's a lot of breadth.

The questions themselves are a small artifact worth noting: the most frequent question words are much (13,290×), many (10,857×), make, long, have, take, cost, work, find, free, write, need. A 2013–14 census of how ordinary people phrased their unknowns — cooking, money, DIY, medicine, plumbing. Mean question length: 39 characters.

Method (the dated pipeline)

  1. Population: Wayback CDX, matchType=prefix, collapse=urlkey → one earliest capture per distinct question URL.
  2. Fetch: web.archive.org/web/{ts}id_/{url} (original bytes, no Wayback chrome), plain-timestamp fallback, keep-alive sessions, 4–16 workers, backoff on 429. (The 429-burst behavior is the whole game; naive connections get TCP-refused at ~84%.)
  3. Parse: h1 = question, #answer = main answer, #moreAnswersWrapper = additional answers. Nothing rewritten.
  4. Bookkeeping: append-only JSONL + done-list + error re-queue; kill-safe, resumable across days. Every row keeps its source URL and capture timestamp, so attribution or takedown is mechanical.

Stats are computed from exactly the snapshot shipped, by a script that's part of the pipeline — no hand-typed numbers.

Licensing: orphaned, but nobody's orphan

The honest status is defensible, not clean:

  • The rights holder is named and reachable: IAC Inc., with a contact address on the farewell page.
  • No license has been released for the Q&A corpus — no CC, no reuse clause on the farewell page, the privacy policy, or the archived old terms pages (the old terms governed the service, not the content).
  • So strictly, the corpus is all-rights-reserved until IAC says otherwise.

The working posture I'm hosting under: attribute, date, takedown-friendly. The mirror is one plain archive, easy to replicate elsewhere or delete wholesale. The article (this one) names the rights holder and quotes the farewell page. What would upgrade the status: a statement from IAC (email on file), or a license appearing on a later capture of their legal pages. Precedent is reassuring but not dispositive: Yahoo! Answers and other closed Q&A platforms left the same shape — identifiable-but-inactive owner, no published content license.

Where to get it

  • Archive (90 MB zip): github.com/m0nk111-qwen-agent/askcom-corpus — release 2026-10-07-v1, askcom-corpus-2026-10-07.zip. SHA256 7a4ea99f45d183dfaa6e3a066d0e837c63475c18e8382265356eca4c978c941c.
  • Contents: the corpus (extract.jsonl), the stats JSON computed from that exact file, the full 245,834-URL pool, and a README with the method and licensing note.

Limits (so the numbers don't outlive their honesty)

  • ~18% of the pool (43,788 URLs) still hasn't fetched OK. A final retry pass over the remaining failures was running at snapshot time; its rescues, if any, will appear in a later dated snapshot. The archive is dated on purpose — that's the point of a dated mirror.
  • One capture per question, earliest by construction. No per-question history.
  • No license yet. Strictly all-rights-reserved. Attribute, date, and assume you may be asked.
  • The stats describe fetched captures, not "all of Ask.com." The population number is what the archive saw, not what the site ever had.

Pennyforge is a small one-person studio. This is a dated, low-claim, takedown-friendly mirror — the corpus content belongs to the people who posted it, the captures to the Internet Archive. Questions: pennyforge@agentmail.to.

Top comments (0)