You need to know what a website used to look like: every page a domain ever had before a migration, the version of a pricing page from last spring, or what an expired domain hosted before you buy it. The Wayback Machine has the answer, but its calendar view shows one page at a time and gives you nothing you can sort or filter. You want a table.
The Wayback Machine Snapshots and Archived URLs actor, published by Hay Equipos on the Apify Store, reads the Internet Archive's own public CDX and availability APIs and returns clean rows. It has three modes:
- Snapshots: the capture history of one exact page, optionally thinned to one capture per hour, day, month or year, or to captures where the content changed.
- Archived URLs: every distinct URL the archive holds under a domain or path, with its first capture.
- Closest: the one archived copy nearest to a date.
What you get back
One row per capture or archived URL. Illustrative example:
{
"input": "example.com",
"mode": "snapshots",
"url": "https://www.example.com/",
"capturedAt": "2019-03-04T11:22:33Z",
"timestamp": "20190304112233",
"statusCode": 200,
"mimeType": "text/html",
"digest": "ABCDEFGHIJKLMNOPQRSTUVWXYZ234567",
"lengthBytes": 8421,
"archiveUrl": "https://web.archive.org/web/20190304112233/https://www.example.com/",
"rawArchiveUrl": "https://web.archive.org/web/20190304112233id_/https://www.example.com/",
"scrapedAt": "2026-09-27T07:21:05.981Z"
}
| capturedAt | url | statusCode | mimeType | digest |
|---|---|---|---|---|
2017-01-12T08:00:00Z |
https://www.example.com/ | 200 | text/html | 7KQ2... |
2018-02-03T09:30:00Z |
https://www.example.com/ | 200 | text/html | M4XZ... |
2019-03-04T11:22:33Z |
https://www.example.com/ | 200 | text/html | ABCD... |
Every row has two links: archiveUrl (the normal Wayback page) and rawArchiveUrl (the archived copy without the Wayback toolbar, which is what you want to feed a parser or an AI model). digest is the Archive's fingerprint of the content, so two captures with the same digest are identical. lengthBytes is the compressed size the Archive stores. In Archived URLs mode each row also has firstCapturedAt. In Closest mode each row also has requestedDate.
Step by step in the Apify Console
- Open the actor on the Apify Store (link at the end) and click Try for free.
- In URLs or domains, add one per line: a page (
example.com/about), a domain (example.com) or a path prefix (example.com/blog/). The scheme andwwware optional, and pasted Wayback links work too. - Choose What to get: Snapshots, Archived URLs, or Closest.
- Optional: set From date and To date as
2019,2019-06or2019-06-30. A year or month covers the whole period. - For Snapshots, use Keep one snapshot per to thin the history (hour, day, month, year, or content change).
- Turn on Only successful captures to drop redirects and errors, and use Content type filter (
text/html,application/pdf, or a pattern likeimage/.*) to narrow results. - For Archived URLs, turn on Include subdomains to also list
blog.example.comand similar. For Closest, set Closest to date, or leave it empty for the most recent capture. - Set Maximum rows per URL or domain (default 500) and Maximum rows in total (default 5,000), then click Start and export from the Output tab.
Example: every archived HTML page under a docs folder.
{
"urls": ["crawlee.dev/docs/"],
"mode": "urls",
"mimeType": "text/html",
"onlySuccessful": true,
"maxRowsPerInput": 2000
}
Calling it from code
The actor id is pistachio_implementation/wayback-machine-snapshots. Keep your token in APIFY_TOKEN.
curl, one capture per year for a home page:
curl -X POST \
"https://api.apify.com/v2/acts/pistachio_implementation~wayback-machine-snapshots/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls": ["apify.com"], "mode": "snapshots", "from": "2015", "to": "2025", "collapse": "year", "onlySuccessful": true}'
Python with apify-client, finding the copy of a page closest to a date:
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("pistachio_implementation/wayback-machine-snapshots").call(
run_input={
"urls": ["example.com/pricing", "example.com/terms"],
"mode": "closest",
"closestTo": "2020-01-15",
}
)
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
print(row["input"], row["capturedAt"], row["rawArchiveUrl"])
summary = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("RUN_SUMMARY")
print(summary["value"]["errors"]) # inputs with no captures
Pricing
Pay per event, with no start fee, no subscription and no platform usage charged on top:
- $0.0005 per archive row saved ($0.50 per 1,000 rows).
- Inputs with no captures, invalid inputs and Archive errors are free. They are listed in
RUN_SUMMARYin the run's key value store.
For example, the yearly history of 100 home pages over 10 years is about 1,000 rows, so about $0.50. You can set a maximum charge per run and the actor stops cleanly when it is reached.
Limits and what it does not do
-
Capture records and links only, not page content. Open or fetch
rawArchiveUrlto get the page as it was captured. - The Archive's API can be slow, often 10 to 30 seconds per request on a large domain, and sometimes answers that it is busy. The actor makes one request at a time, at least 1.2 seconds apart, and backs off and retries. Very large domains take minutes.
- Rows come oldest first. If you only want recent captures, set the From date.
- Sites that asked the Archive to exclude them return no rows.
- It only knows what the Wayback Machine captured. Pages the Archive never crawled will not appear.
Try it on the Apify Store: https://apify.com/pistachio_implementation/wayback-machine-snapshots
Top comments (0)