Skip to main content
Thirdwatchthirdwatch
Other

Find Open-Access PDFs With Google Scholar Data

Export Scholar results including direct PDF links where indexed — locate open-access versions of papers as JSON instead of chasing paywalls.

Sep 21, 2026 · 2 min read · 464 words
See the scraper →

TL;DR — The Google Scholar Scraper returns pdfUrl — the open-access copy Scholar indexed — alongside each paper's metadata. Query a reading list, filter to records with pdfUrl, and skip the paywall chase entirely.

Why the open version is already indexed

Scholar quietly links millions of papers to free copies — arXiv preprints, repository deposits, publisher OA versions. The information exists; it's scattered one result at a time. A queryable export turns "does a free copy exist?" into a field check.

The job-to-be-done: for a reading list or citation corpus, which papers have an open PDF — and where.

How does this compare to the alternatives?

Manual per-paper search Browser extensions Thirdwatch actor
Cost Free, slow Free, per-page Pay per result
Reliability One at a time Only while browsing Batch export
Setup time Zero Extension install Minutes
Maintenance Every paper Per lookup Re-run the list

Extensions help when you're already on a page. For a hundred-paper list, you want the links in a table.

How to locate open PDFs in 4 steps

How do I check a reading list?

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/google-scholar-scraper").call(
    run_input={
        "queries": [
            '"Scaling Instruction-Finetuned Language Models"',
            '"LoRA: Low-Rank Adaptation of Large Language Models"',
        ],
        "maxResults": 5,
    }
)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())

How do I split open vs paywalled?

import pandas as pd
df = pd.DataFrame(rows).drop_duplicates("title")
open_oa = df[df.pdfUrl.notna()][["title", "pdfUrl", "venue", "year"]]
closed = df[df.pdfUrl.isna()][["title", "url", "venue"]]
print(f"{len(open_oa)} open / {len(closed)} closed")

How do I fetch the open copies?

import httpx, pathlib
for r in open_oa.to_dict("records"):
    try:
        data = httpx.get(r["pdfUrl"], timeout=30, follow_redirects=True).content
        pathlib.Path("pdfs").mkdir(exist_ok=True)
        pathlib.Path("pdfs", r["title"][:50] + ".pdf").write_bytes(data)
    except Exception:
        pass

How do I handle the closed ones?

url is the canonical page — institutional access, author copy, or interlibrary loan start there. The split itself is the useful deliverable: you know which papers need a workaround and which don't.

Sample output

{"title": "LoRA: Low-Rank Adaptation of Large Language Models",
 "url": "https://arxiv.org/abs/2106.09685",
 "authors": "E Hu, Y Shen", "venue": "ICLR", "year": "2022",
 "citedBy": 11000, "pdfUrl": "https://arxiv.org/pdf/2106.09685",
 "query": "\"LoRA: Low-Rank Adaptation\"", "position": 1}

Common pitfalls

pdfUrl absence doesn't mean closed — Scholar sometimes indexes a version without linking the PDF. Repository links rot; fetch promptly and keep the record. Publisher-hosted "free" PDFs may still carry license terms — read them before redistribution. The actor locates; rights and access rules remain yours.

Related use cases

Frequently asked questions

Where do the pdfUrl links point?

To the PDF Scholar indexed — usually an arXiv, institutional repository, or publisher open-access copy. It is the legal open version, not a paywall bypass.

What fraction of papers have a pdfUrl?

It varies by field — heavy in CS and physics (arXiv), thinner in paywalled disciplines. Filter records where pdfUrl is present; the rest point to their canonical page.

Can I search for a specific paper's free version?

Yes — query the exact title. If Scholar knows an open copy, its record carries pdfUrl alongside the canonical url.

Does the actor download PDFs?

No — it returns the links. Fetching is a separate step on the hosting repository, which keeps provenance clean.

Related

Try it yourself

100 free credits, no credit card.

About 30 real searches. Add the MCP to Claude or Cursor in two minutes.