Skip to main content
Thirdwatchthirdwatch
Other

Build a Public-Domain Media Catalog From Archive.org

Export archive.org item metadata — titles, media types, files, downloads — as JSON to assemble a rights-clear public-domain media catalog.

Sep 21, 2026 · 3 min read · 536 words
See the scraper →

TL;DR — The Internet Archive Scraper turns archive.org into a queryable catalog: search mode finds items, item mode returns full metadata and file lists. Filter to public-domain collections and you have a rights-clear media dataset as JSON.

Why build a catalog instead of browsing

Public-domain media is an ingredient — for apps, datasets, classrooms, model training. The Internet Archive hosts millions of qualifying items across books, audio, film, and images (archive.org). But "what exists, in what format, under what license" is a database question, and the site is a search box.

The deliverable is a table: identifier, title, media type, file inventory, license signal — filterable, joinable, refreshable.

How does this compare to the alternatives?

Site browsing Bulk metadata dumps Thirdwatch actor
Cost Free Free + heavy lifting Pay per result
Reliability One item at a time Terabyte-scale dumps Bounded, current
Setup time Zero Days of ETL Minutes
Maintenance Manual Yours Handled

Archive.org publishes bulk metadata dumps — they are enormous and stale the moment you untar them. Query-time export gives you just the slice you need, fresh.

How to build the catalog in 4 steps

How do I enumerate a collection?

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/internet-archive-scraper").call(
    run_input={
        "mode": "search",
        "query": "",
        "collection": "gutenberg",
        "mediaTypes": ["texts"],
        "sort": "downloads desc",
        "maxResults": 500,
    }
)
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())

An empty query with a collection filter enumerates rather than searches — the right shape for catalog work.

How do I attach file lists?

Feed search identifiers into item mode:

ids = [i["identifier"] for i in items[:100]]
run = client.actor("thirdwatch/internet-archive-scraper").call(
    run_input={"mode": "item", "identifiers": ids, "maxFilesPerItem": 50}
)
meta = list(client.dataset(run["defaultDatasetId"]).iterate_items())

item_metadata records carry the per-file inventory — formats, sizes, and the license field where the archive has one.

How do I filter to usable formats?

import pandas as pd
df = pd.DataFrame(meta)
epub = df[df["files"].apply(lambda fs: any("epub" in (f.get("name") or "").lower() for f in (fs or [])))]

Keep the items whose file list actually contains a format your project can serve.

How do I dedupe and version the catalog?

identifier is the stable primary key. Re-run monthly, join on it, and downloads deltas show which items are gaining use:

catalog = df.drop_duplicates("identifier").set_index("identifier")

Sample output

{"record_type": "item", "identifier": "aliceinwonderl0000carr",
 "title": "Alice's Adventures in Wonderland", "mediatype": "texts",
 "downloads": 4213, "collections": ["gutenberg"],
 "details_url": "https://archive.org/details/aliceinwonderl0000carr"}

Item mode then expands each identifier into a metadata record with its file list — the two-pass shape that keeps runs cheap.

Common pitfalls

Public domain status varies by jurisdiction and edition — the license field helps, but legal review is yours. downloads is a popularity counter, not a quality filter. Very large collections need pagination by yearFrom/yearTo slices or subjects. The actor returns archive.org's fields as-is; missing fields reflect the source item, not a bug.

Related use cases

Frequently asked questions

Is everything on archive.org public domain?

No. The archive hosts public-domain, Creative Commons, and lendable modern works. Check each item's license field in item-mode metadata before reuse — the actor returns it, but the rights decision is yours.

How do I get file-level detail?

Run item mode with identifiers from search results. Each item_metadata record lists the item's files with names and formats, capped by maxFilesPerItem.

Can I filter to a curated collection like Project Gutenberg?

Yes. Set the collection field to the collection identifier, optionally combined with mediaTypes and a subject or creator query.

How large can the catalog get?

As large as your maxResults and filters allow. For a serious corpus, iterate by year or subject so each run stays bounded and resumable.

Related

Try it yourself

100 free credits, no credit card.

About 30 real searches. Add the MCP to Claude or Cursor in two minutes.