List Wayback Machine Snapshots to Track Site History
Export every Wayback Machine capture of a URL or domain as JSON — timestamp, status, digest — to reconstruct site history and find what changed.

TL;DR — Wayback mode of the Internet Archive Scraper exports every archived capture of a URL or domain — timestamp, status, digest, playback link — as JSON. Collapse by digest to see only real changes. Built for researchers reconstructing site history.
Why site history is a research input
A company's old pricing page, a product's removed claim, a competitor's repositioned homepage — the evidence exists in the Wayback Machine's hundreds of billions of captures, described on archive.org. The question "what did this page say on March 3rd?" is answerable — if you can list captures without clicking through a calendar UI.
The job is a capture inventory: every snapshot, deduped by content hash, as a table you can filter and dereference.
How does this compare to the alternatives?
| Wayback calendar UI | Raw CDX API scripting | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free | Free + your time | Pay per result |
| Reliability | Click per capture | You maintain it | Maintained |
| Setup time | Zero | Hours | Minutes |
| Maintenance | Manual | Yours | Handled |
The calendar UI answers one URL at a time. The actor answers a domain's worth of URLs as a dataset.
How to list snapshots in 4 steps
How do I list captures for one URL?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/internet-archive-scraper").call(
run_input={
"mode": "wayback",
"urls": ["apify.com/pricing"],
"waybackMatchType": "exact",
"waybackStatusCode": "200",
"maxResults": 200,
}
)
caps = list(client.dataset(run["defaultDatasetId"]).iterate_items())exact matches the URL as written; waybackStatusCode: "200" skips redirects and error captures.
How do I cover a whole site section?
prefix matches a path subtree, host the whole host, domain including subdomains:
run = client.actor("thirdwatch/internet-archive-scraper").call(
run_input={
"mode": "wayback",
"urls": ["competitor.com"],
"waybackMatchType": "domain",
"waybackCollapse": "digest",
"maxResults": 1000,
}
)How do I collapse to real changes only?
waybackCollapse: "digest" keeps one capture per distinct content hash — the moments the page actually changed:
import pandas as pd
df = pd.DataFrame(caps)
df["timestamp"] = pd.to_datetime(df["timestamp"], format="%Y%m%d%H%M%S")
print(df.sort_values("timestamp")[["timestamp", "original"]].tail(20))That list is your change log — each row is a capture worth opening.
How do I bound the window?
waybackFrom/waybackTo accept Wayback timestamps (year or full yyyyMMdd):
run_input={
"mode": "wayback",
"urls": ["apify.com"],
"waybackFrom": "2025",
"waybackTo": "20251231",
"waybackCollapse": "timestamp:6",
}timestamp:6 collapse keeps one capture per month — a fast site-history skeleton.
Sample output
{"record_type": "wayback_snapshot", "timestamp": "2026-03-14T09:22:41Z",
"original": "https://apify.com/pricing", "statuscode": "200",
"digest": "A1B2C3D4E5", "playback_url": "https://web.archive.org/web/20260314092241/https://apify.com/pricing"}digest is the content hash — identical digests mean identical bytes. playback_url opens the capture directly.
Common pitfalls
Not every capture is a real change — crawl artifacts and session parameters create noise, which is what digest collapse is for. Very popular URLs have thousands of captures; bound with waybackFrom/To and collapse before raising maxResults. A 200 capture can still be a parked page or redirect-stub — spot-check playback URLs. The actor returns the capture list; rendering archived pages stays with web.archive.org.
Related use cases
Frequently asked questions
What does wayback mode return?
One record per archived capture: timestamp, original URL, status code, digest, and a playback link. Pass a URL, path prefix, host, or whole domain via waybackMatchType.
Can I deduplicate identical captures?
Yes. waybackCollapse='digest' collapses captures whose content hash is unchanged, so you see only captures where the page actually changed.
How do I limit to a date range or status?
waybackFrom and waybackTo bound the period; waybackStatusCode filters to e.g. 200 captures only, skipping redirects and errors.
Does the actor fetch the archived pages?
It lists captures with their playback URLs. Fetching page content is a separate request to web.archive.org — the list tells you exactly what exists and when.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.