Skip to main content
Thirdwatchthirdwatch
Business & local data

Monitor Competitor Website Changes With Wayback Data

List a competitor's archived page captures as JSON and pinpoint when pricing pages, messaging, and positioning actually changed.

Sep 21, 2026 · 3 min read · 534 words
See the scraper →

TL;DR — Wayback mode exports every archived capture of a competitor's pages — timestamp, status, content digest, playback URL. Collapse by digest and you get a change log of when their pricing, positioning, and product pages actually changed.

Why Wayback beats "I think they changed something"

Competitor intelligence usually runs on vibes — someone notices a new headline six weeks late. The Wayback Machine keeps an objective record: hundreds of billions of captures (archive.org), timestamped and content-hashed. When a rival drops a price tier or softens a claim, the capture history says when.

The job-to-be-done: a deduplicated timeline of real page versions, so "they changed pricing in March" is a fact, not a rumor.

How does this compare to the alternatives?

Manual spot checks Live page monitors Thirdwatch actor (Wayback)
Cost Free, sporadic Subscription Pay per result
Reliability Misses silent changes Real-time only Retrospective, complete
Setup time Zero Moderate Minutes
Maintenance Memory Yours Handled

Live monitors catch changes going forward; Wayback reconstructs what already happened. You want both — this actor covers the past.

How to track a competitor's history in 4 steps

How do I build the change timeline?

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/internet-archive-scraper").call(
    run_input={
        "mode": "wayback",
        "urls": ["competitor.com/pricing"],
        "waybackMatchType": "exact",
        "waybackStatusCode": "200",
        "waybackCollapse": "digest",
        "maxResults": 300,
    }
)
versions = list(client.dataset(run["defaultDatasetId"]).iterate_items())

digest collapse returns one record per distinct page version — each row is a real change.

How do I read the timeline?

import pandas as pd
df = pd.DataFrame(versions)
df["timestamp"] = pd.to_datetime(df["timestamp"])
print(df[["timestamp", "playback_url"]].sort_values("timestamp"))

Consecutive timestamps are version boundaries: the pricing page on record at each interval is the one the previous capture shows.

How do I cover the whole site, not one page?

run_input={
    "mode": "wayback",
    "urls": ["competitor.com"],
    "waybackMatchType": "domain",
    "waybackCollapse": "digest",
    "waybackFrom": "2026",
    "maxResults": 1000,
}

Domain match plus digest collapse gives a site-wide change inventory for the year — product pages, docs, careers, everything the archive saw change.

How do I diff the interesting pair?

before, after = df["playback_url"].iloc[-2], df["playback_url"].iloc[-1]
# fetch both URLs, extract text, difflib the paragraphs
import difflib, httpx
from bs4 import BeautifulSoup
texts = [BeautifulSoup(httpx.get(u).text, "html.parser").get_text() for u in (before, after)]
print("\n".join(difflib.unified_diff(texts[0].split(), texts[1].split(), lineterm="")[:60]))

Only diff where digest differs — the collapse already did that triage.

Sample output

{"record_type": "wayback_snapshot", "timestamp": "2026-04-02T17:40:11Z",
 "original": "https://competitor.com/pricing", "statuscode": "200",
 "digest": "F0E1D2C3",
 "playback_url": "https://web.archive.org/web/20260402174011/https://competitor.com/pricing"}

digest is the dedupe key; playback_url is the exhibit.

Common pitfalls

Wayback captures lag — a change last week may not be archived yet, so this is retrospective, not alerting. Heavy JS pages sometimes capture poorly; check the playback link before drawing conclusions. Popular domains produce thousands of captures — always combine waybackCollapse with a date bound. The actor lists captures; the diff itself is your follow-up step.

Related use cases

Frequently asked questions

How current is Wayback data for competitor tracking?

Wayback captures happen on the archive's crawl schedule — days to weeks between captures for typical sites. It answers 'when did this change' retrospectively, not in real time; pair it with a live monitor for real-time alerts.

Can I watch just the pricing or careers page?

Yes. Pass that exact URL with waybackMatchType 'exact' and collapse by digest — you get one record per distinct version of that page.

What about pages that were deleted?

Deleted pages are exactly what Wayback preserves. If the archive crawled them, their captures remain listed and playable even after removal from the live site.

How do I see what actually changed between captures?

The record gives each capture a playback URL and content digest. Fetch the two captures flanking a digest change and diff the text — the digest tells you which pairs are worth comparing.

Related

Try it yourself

100 free credits, no credit card.

About 30 real searches. Add the MCP to Claude or Cursor in two minutes.