Skip to main content
Thirdwatchthirdwatch
Other

Search Internet Archive Collections for Research

Query archive.org's catalog of books, film, audio, and software as structured JSON — filter by media type, creator, subject, and year for research.

Sep 21, 2026 · 3 min read · 555 words
See the scraper →

TL;DR — The Internet Archive Scraper searches archive.org's catalog of books, film, audio, and software and returns structured item records — identifier, title, media type, downloads, collections. Built for researchers who need the library as data, not as a search box.

Why search the Internet Archive as data

The Internet Archive holds one of the largest public media catalogs in existence — tens of millions of texts plus audio, video, and software, described on its about page. For research questions like "which Apollo-era broadcasts survive online" or "how large is the public-domain piano-roll corpus", the answer is a filtered catalog query — not an afternoon of clicking.

The job-to-be-done is a dataset: identifiers you can dereference, download counts you can rank, and collection tags you can group.

How does this compare to the alternatives?

Website search Raw API scripting Thirdwatch actor
Cost Free Free + your time Pay per result
Reliability Manual paging You maintain parsing Maintained
Setup time Zero Hours Minutes
Maintenance None Yours Handled

The site search is built for humans; the raw API returns paginated JSON you wire up yourself. The actor returns flat records ready for pandas.

How to search the catalog in 4 steps

How do I run a basic catalog search?

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/internet-archive-scraper").call(
    run_input={
        "mode": "search",
        "query": "apollo 11",
        "sort": "downloads desc",
        "maxResults": 50,
    }
)
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())

maxResults caps the records; sort by downloads desc to surface canonical items first.

How do I filter to one media type or creator?

run = client.actor("thirdwatch/internet-archive-scraper").call(
    run_input={
        "mode": "search",
        "query": "",
        "creator": "NASA",
        "mediaTypes": ["movies"],
        "yearFrom": 1969,
        "yearTo": 1975,
        "maxResults": 100,
    }
)

creator, subject, collection, language, and yearFrom/yearTo combine — use them to carve a corpus rather than a query.

How do I read an item's full metadata and file list?

Switch to item mode with identifiers from the search:

run = client.actor("thirdwatch/internet-archive-scraper").call(
    run_input={
        "mode": "item",
        "identifiers": ["CSPAN3_20190824_120000"],
        "maxFilesPerItem": 50,
    }
)

record_type: "item_metadata" rows carry the item's full metadata and file inventory — the step between "found it" and "download it".

How do I build the corpus as a table?

import pandas as pd
df = pd.DataFrame(items)
print(df.groupby("mediatype")["downloads"].agg(["count", "sum"]))

Group by mediatype or explode collections to profile what the archive actually holds on your topic.

Sample output

{"record_type": "item", "identifier": "CSPAN3_20190824_120000",
 "title": "20th anniversary special on Apollo 11 moon landing",
 "mediatype": "movies", "downloads": 58,
 "collections": ["TV-BBC"],
 "details_url": "https://archive.org/details/CSPAN3_20190824_120000"}

identifier is the stable key — it feeds item mode and constructs details_url.

Common pitfalls

Search relevance ranking is archive.org's own — for corpus work prefer filters over query tuning. Big collections can exceed maxResults; page through with narrower filters rather than one huge pull. downloads reflects archive.org's counter, not your quality bar — sort by it, don't trust it. The actor returns archive.org's schema faithfully, so fields vary by media type.

Related use cases

Frequently asked questions

What does search mode return?

One record per archive.org item: identifier, title, media type, download count, collections, and a details URL. Filters like mediaTypes, creator, subject, language, and year range narrow the result set before it reaches you.

Can I search inside a specific collection?

Yes. The collection field restricts results to one archive.org collection, and the query field accepts archive.org's own query syntax like creator:"NASA" AND subject:"mars" for precise work.

How do I sort results?

The sort field supports downloads, public date, added date, item size, and identifier ordering in both directions — 'downloads desc' surfaces the most-used items first.

Is the Internet Archive free to query?

Yes. archive.org is a nonprofit library and its metadata APIs are public. The actor packages them into a clean, filterable dataset so you skip request plumbing.

Related

Try it yourself

100 free credits, no credit card.

About 30 real searches. Add the MCP to Claude or Cursor in two minutes.