DEV Community

Hay Equipos
Hay Equipos

Posted on

How to get all archived URLs and snapshots of a website from the Wayback Machine

You need to know what a website used to look like: every page a domain ever had before a migration, the version of a pricing page from last spring, or what an expired domain hosted before you buy it. The Wayback Machine has the answer, but its calendar view shows one page at a time and gives you nothing you can sort or filter. You want a table.

The Wayback Machine Snapshots and Archived URLs actor, published by Hay Equipos on the Apify Store, reads the Internet Archive's own public CDX and availability APIs and returns clean rows. It has three modes:

  1. Snapshots: the capture history of one exact page, optionally thinned to one capture per hour, day, month or year, or to captures where the content changed.
  2. Archived URLs: every distinct URL the archive holds under a domain or path, with its first capture.
  3. Closest: the one archived copy nearest to a date.

What you get back

One row per capture or archived URL. Illustrative example:

{
  "input": "example.com",
  "mode": "snapshots",
  "url": "https://www.example.com/",
  "capturedAt": "2019-03-04T11:22:33Z",
  "timestamp": "20190304112233",
  "statusCode": 200,
  "mimeType": "text/html",
  "digest": "ABCDEFGHIJKLMNOPQRSTUVWXYZ234567",
  "lengthBytes": 8421,
  "archiveUrl": "https://web.archive.org/web/20190304112233/https://www.example.com/",
  "rawArchiveUrl": "https://web.archive.org/web/20190304112233id_/https://www.example.com/",
  "scrapedAt": "2026-09-27T07:21:05.981Z"
}
Enter fullscreen mode Exit fullscreen mode
capturedAt url statusCode mimeType digest
2017-01-12T08:00:00Z https://www.example.com/ 200 text/html 7KQ2...
2018-02-03T09:30:00Z https://www.example.com/ 200 text/html M4XZ...
2019-03-04T11:22:33Z https://www.example.com/ 200 text/html ABCD...

Every row has two links: archiveUrl (the normal Wayback page) and rawArchiveUrl (the archived copy without the Wayback toolbar, which is what you want to feed a parser or an AI model). digest is the Archive's fingerprint of the content, so two captures with the same digest are identical. lengthBytes is the compressed size the Archive stores. In Archived URLs mode each row also has firstCapturedAt. In Closest mode each row also has requestedDate.

Step by step in the Apify Console

  1. Open the actor on the Apify Store (link at the end) and click Try for free.
  2. In URLs or domains, add one per line: a page (example.com/about), a domain (example.com) or a path prefix (example.com/blog/). The scheme and www are optional, and pasted Wayback links work too.
  3. Choose What to get: Snapshots, Archived URLs, or Closest.
  4. Optional: set From date and To date as 2019, 2019-06 or 2019-06-30. A year or month covers the whole period.
  5. For Snapshots, use Keep one snapshot per to thin the history (hour, day, month, year, or content change).
  6. Turn on Only successful captures to drop redirects and errors, and use Content type filter (text/html, application/pdf, or a pattern like image/.*) to narrow results.
  7. For Archived URLs, turn on Include subdomains to also list blog.example.com and similar. For Closest, set Closest to date, or leave it empty for the most recent capture.
  8. Set Maximum rows per URL or domain (default 500) and Maximum rows in total (default 5,000), then click Start and export from the Output tab.

Example: every archived HTML page under a docs folder.

{
  "urls": ["crawlee.dev/docs/"],
  "mode": "urls",
  "mimeType": "text/html",
  "onlySuccessful": true,
  "maxRowsPerInput": 2000
}
Enter fullscreen mode Exit fullscreen mode

Calling it from code

The actor id is pistachio_implementation/wayback-machine-snapshots. Keep your token in APIFY_TOKEN.

curl, one capture per year for a home page:

curl -X POST \
  "https://api.apify.com/v2/acts/pistachio_implementation~wayback-machine-snapshots/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["apify.com"], "mode": "snapshots", "from": "2015", "to": "2025", "collapse": "year", "onlySuccessful": true}'
Enter fullscreen mode Exit fullscreen mode

Python with apify-client, finding the copy of a page closest to a date:

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])

run = client.actor("pistachio_implementation/wayback-machine-snapshots").call(
    run_input={
        "urls": ["example.com/pricing", "example.com/terms"],
        "mode": "closest",
        "closestTo": "2020-01-15",
    }
)

for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(row["input"], row["capturedAt"], row["rawArchiveUrl"])

summary = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("RUN_SUMMARY")
print(summary["value"]["errors"])  # inputs with no captures
Enter fullscreen mode Exit fullscreen mode

Pricing

Pay per event, with no start fee, no subscription and no platform usage charged on top:

  • $0.0005 per archive row saved ($0.50 per 1,000 rows).
  • Inputs with no captures, invalid inputs and Archive errors are free. They are listed in RUN_SUMMARY in the run's key value store.

For example, the yearly history of 100 home pages over 10 years is about 1,000 rows, so about $0.50. You can set a maximum charge per run and the actor stops cleanly when it is reached.

Limits and what it does not do

  • Capture records and links only, not page content. Open or fetch rawArchiveUrl to get the page as it was captured.
  • The Archive's API can be slow, often 10 to 30 seconds per request on a large domain, and sometimes answers that it is busy. The actor makes one request at a time, at least 1.2 seconds apart, and backs off and retries. Very large domains take minutes.
  • Rows come oldest first. If you only want recent captures, set the From date.
  • Sites that asked the Archive to exclude them return no rows.
  • It only knows what the Wayback Machine captured. Pages the Archive never crawled will not appear.

Try it on the Apify Store: https://apify.com/pistachio_implementation/wayback-machine-snapshots

Top comments (0)