Search Internet Archive Collections for Research
Query archive.org's catalog of books, film, audio, and software as structured JSON — filter by media type, creator, subject, and year for research.

TL;DR — The Internet Archive Scraper searches archive.org's catalog of books, film, audio, and software and returns structured item records — identifier, title, media type, downloads, collections. Built for researchers who need the library as data, not as a search box.
Why search the Internet Archive as data
The Internet Archive holds one of the largest public media catalogs in existence — tens of millions of texts plus audio, video, and software, described on its about page. For research questions like "which Apollo-era broadcasts survive online" or "how large is the public-domain piano-roll corpus", the answer is a filtered catalog query — not an afternoon of clicking.
The job-to-be-done is a dataset: identifiers you can dereference, download counts you can rank, and collection tags you can group.
How does this compare to the alternatives?
| Website search | Raw API scripting | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free | Free + your time | Pay per result |
| Reliability | Manual paging | You maintain parsing | Maintained |
| Setup time | Zero | Hours | Minutes |
| Maintenance | None | Yours | Handled |
The site search is built for humans; the raw API returns paginated JSON you wire up yourself. The actor returns flat records ready for pandas.
How to search the catalog in 4 steps
How do I run a basic catalog search?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/internet-archive-scraper").call(
run_input={
"mode": "search",
"query": "apollo 11",
"sort": "downloads desc",
"maxResults": 50,
}
)
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())maxResults caps the records; sort by downloads desc to surface canonical items first.
How do I filter to one media type or creator?
run = client.actor("thirdwatch/internet-archive-scraper").call(
run_input={
"mode": "search",
"query": "",
"creator": "NASA",
"mediaTypes": ["movies"],
"yearFrom": 1969,
"yearTo": 1975,
"maxResults": 100,
}
)creator, subject, collection, language, and yearFrom/yearTo combine — use them to carve a corpus rather than a query.
How do I read an item's full metadata and file list?
Switch to item mode with identifiers from the search:
run = client.actor("thirdwatch/internet-archive-scraper").call(
run_input={
"mode": "item",
"identifiers": ["CSPAN3_20190824_120000"],
"maxFilesPerItem": 50,
}
)record_type: "item_metadata" rows carry the item's full metadata and file inventory — the step between "found it" and "download it".
How do I build the corpus as a table?
import pandas as pd
df = pd.DataFrame(items)
print(df.groupby("mediatype")["downloads"].agg(["count", "sum"]))Group by mediatype or explode collections to profile what the archive actually holds on your topic.
Sample output
{"record_type": "item", "identifier": "CSPAN3_20190824_120000",
"title": "20th anniversary special on Apollo 11 moon landing",
"mediatype": "movies", "downloads": 58,
"collections": ["TV-BBC"],
"details_url": "https://archive.org/details/CSPAN3_20190824_120000"}identifier is the stable key — it feeds item mode and constructs details_url.
Common pitfalls
Search relevance ranking is archive.org's own — for corpus work prefer filters over query tuning. Big collections can exceed maxResults; page through with narrower filters rather than one huge pull. downloads reflects archive.org's counter, not your quality bar — sort by it, don't trust it. The actor returns archive.org's schema faithfully, so fields vary by media type.
Related use cases
- Build an arXiv research dataset — the academic-paper counterpart.
- Analyze Hacker News discussions with Python
- Guide to scraping business data
- Blog hub
Frequently asked questions
What does search mode return?
One record per archive.org item: identifier, title, media type, download count, collections, and a details URL. Filters like mediaTypes, creator, subject, language, and year range narrow the result set before it reaches you.
Can I search inside a specific collection?
Yes. The collection field restricts results to one archive.org collection, and the query field accepts archive.org's own query syntax like creator:"NASA" AND subject:"mars" for precise work.
How do I sort results?
The sort field supports downloads, public date, added date, item size, and identifier ordering in both directions — 'downloads desc' surfaces the most-used items first.
Is the Internet Archive free to query?
Yes. archive.org is a nonprofit library and its metadata APIs are public. The actor packages them into a clean, filterable dataset so you skip request plumbing.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.