Skip to main content
Thirdwatchthirdwatch
Other

List Wayback Machine Snapshots to Track Site History

Export every Wayback Machine capture of a URL or domain as JSON — timestamp, status, digest — to reconstruct site history and find what changed.

Sep 21, 2026 · 3 min read · 536 words
See the scraper →

TL;DR — Wayback mode of the Internet Archive Scraper exports every archived capture of a URL or domain — timestamp, status, digest, playback link — as JSON. Collapse by digest to see only real changes. Built for researchers reconstructing site history.

Why site history is a research input

A company's old pricing page, a product's removed claim, a competitor's repositioned homepage — the evidence exists in the Wayback Machine's hundreds of billions of captures, described on archive.org. The question "what did this page say on March 3rd?" is answerable — if you can list captures without clicking through a calendar UI.

The job is a capture inventory: every snapshot, deduped by content hash, as a table you can filter and dereference.

How does this compare to the alternatives?

Wayback calendar UI Raw CDX API scripting Thirdwatch actor
Cost Free Free + your time Pay per result
Reliability Click per capture You maintain it Maintained
Setup time Zero Hours Minutes
Maintenance Manual Yours Handled

The calendar UI answers one URL at a time. The actor answers a domain's worth of URLs as a dataset.

How to list snapshots in 4 steps

How do I list captures for one URL?

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/internet-archive-scraper").call(
    run_input={
        "mode": "wayback",
        "urls": ["apify.com/pricing"],
        "waybackMatchType": "exact",
        "waybackStatusCode": "200",
        "maxResults": 200,
    }
)
caps = list(client.dataset(run["defaultDatasetId"]).iterate_items())

exact matches the URL as written; waybackStatusCode: "200" skips redirects and error captures.

How do I cover a whole site section?

prefix matches a path subtree, host the whole host, domain including subdomains:

run = client.actor("thirdwatch/internet-archive-scraper").call(
    run_input={
        "mode": "wayback",
        "urls": ["competitor.com"],
        "waybackMatchType": "domain",
        "waybackCollapse": "digest",
        "maxResults": 1000,
    }
)

How do I collapse to real changes only?

waybackCollapse: "digest" keeps one capture per distinct content hash — the moments the page actually changed:

import pandas as pd
df = pd.DataFrame(caps)
df["timestamp"] = pd.to_datetime(df["timestamp"], format="%Y%m%d%H%M%S")
print(df.sort_values("timestamp")[["timestamp", "original"]].tail(20))

That list is your change log — each row is a capture worth opening.

How do I bound the window?

waybackFrom/waybackTo accept Wayback timestamps (year or full yyyyMMdd):

run_input={
    "mode": "wayback",
    "urls": ["apify.com"],
    "waybackFrom": "2025",
    "waybackTo": "20251231",
    "waybackCollapse": "timestamp:6",
}

timestamp:6 collapse keeps one capture per month — a fast site-history skeleton.

Sample output

{"record_type": "wayback_snapshot", "timestamp": "2026-03-14T09:22:41Z",
 "original": "https://apify.com/pricing", "statuscode": "200",
 "digest": "A1B2C3D4E5", "playback_url": "https://web.archive.org/web/20260314092241/https://apify.com/pricing"}

digest is the content hash — identical digests mean identical bytes. playback_url opens the capture directly.

Common pitfalls

Not every capture is a real change — crawl artifacts and session parameters create noise, which is what digest collapse is for. Very popular URLs have thousands of captures; bound with waybackFrom/To and collapse before raising maxResults. A 200 capture can still be a parked page or redirect-stub — spot-check playback URLs. The actor returns the capture list; rendering archived pages stays with web.archive.org.

Related use cases

Frequently asked questions

What does wayback mode return?

One record per archived capture: timestamp, original URL, status code, digest, and a playback link. Pass a URL, path prefix, host, or whole domain via waybackMatchType.

Can I deduplicate identical captures?

Yes. waybackCollapse='digest' collapses captures whose content hash is unchanged, so you see only captures where the page actually changed.

How do I limit to a date range or status?

waybackFrom and waybackTo bound the period; waybackStatusCode filters to e.g. 200 captures only, skipping redirects and errors.

Does the actor fetch the archived pages?

It lists captures with their playback URLs. Fetching page content is a separate request to web.archive.org — the list tells you exactly what exists and when.

Related

Try it yourself

100 free credits, no credit card.

About 30 real searches. Add the MCP to Claude or Cursor in two minutes.