Research Historical Media Collections on Archive.org
Pull archive.org item records — old broadcasts, scanned books, audio — as JSON to size and profile historical media collections for research.

TL;DR — The Internet Archive Scraper exports archive.org item records — identifier, title, media type, downloads, collections — as JSON. Researchers use it to size, profile, and inventory historical media collections without clicking through a search box.
Why historical media research needs data, not browsing
Digital-humanities and media-history questions are quantitative: how much 1970s local TV survives, which subjects dominate a scanned-journal corpus, how access counts distribute. The Internet Archive hosts the material — tens of millions of items per its about page — but the interface answers one query at a time.
The research need is a corpus table: every item matching your criteria, with the fields that let you count, group, and cite.
How does this compare to the alternatives?
| Site search | Bulk metadata torrents | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free | Free + heavy ETL | Pay per result |
| Reliability | Manual | Snapshot in time | Current, bounded |
| Setup time | Zero | Days | Minutes |
| Maintenance | None | Yours | Handled |
Bulk dumps are for archive-scale projects. For a defined corpus — one collection, one subject range — query-time export is proportionate.
How to profile a collection in 4 steps
How do I pull the corpus?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/internet-archive-scraper").call(
run_input={
"mode": "search",
"collection": "tvarchive",
"mediaTypes": ["movies"],
"yearFrom": 1960,
"yearTo": 1979,
"sort": "publicdate asc",
"maxResults": 1000,
}
)
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())How do I profile coverage?
import pandas as pd
df = pd.DataFrame(items)
print(df["mediatype"].value_counts())
print(df["collections"].explode().value_counts().head(10))The distribution of media types and sub-collections is the collection's shape — your methods section's first table.
How do I deepen selected items?
ids = [i["identifier"] for i in items if i["downloads"] > 100][:200]
run = client.actor("thirdwatch/internet-archive-scraper").call(
run_input={"mode": "item", "identifiers": ids, "maxFilesPerItem": 50}
)Two passes — cheap search for the census, item mode for the subsample that needs full metadata.
How do I cite the dataset?
Each record's identifier plus details_url is the citation unit. Export the corpus table as CSV alongside your notes and the dataset is reproducible by anyone with the same filters.
Sample output
{"record_type": "item", "identifier": "CSPAN3_20190824_120000",
"title": "20th anniversary special on Apollo 11 moon landing",
"mediatype": "movies", "downloads": 58,
"collections": ["TV-BBC"],
"details_url": "https://archive.org/details/CSPAN3_20190824_120000"}Common pitfalls
Catalog metadata reflects what uploaders wrote — dates and creators need verification for formal claims. Very large filters truncate at maxResults; slice by year or sub-collection. Sort order changes coverage shape — downloads desc oversamples popular items for a census. The actor returns the archive's fields faithfully; empty fields mean the source record is empty.
Related use cases
Frequently asked questions
What kinds of collections can I profile?
Any archive.org collection or query slice — old TV broadcasts, scanned journals, radio audio, software. mediaTypes, collection, creator, subject, language, and year filters define the corpus precisely.
Can I measure collection size and coverage?
Yes. Search returns item records with media type, downloads, and collections; grouping those fields profiles what a collection contains and how used it is.
How do I get deeper than title-level metadata?
Item mode returns full item metadata and file lists per identifier — descriptions, dates, formats — capped by maxFilesPerItem.
Is archive.org metadata reliable for citation?
It is the archive's own catalog record — good enough to cite as 'archive.org item X'. Verify dates and creators against the item page for formal publication.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.