Find Open-Access PDFs With Google Scholar Data
Export Scholar results including direct PDF links where indexed — locate open-access versions of papers as JSON instead of chasing paywalls.

TL;DR — The Google Scholar Scraper returns
pdfUrl— the open-access copy Scholar indexed — alongside each paper's metadata. Query a reading list, filter to records withpdfUrl, and skip the paywall chase entirely.
Why the open version is already indexed
Scholar quietly links millions of papers to free copies — arXiv preprints, repository deposits, publisher OA versions. The information exists; it's scattered one result at a time. A queryable export turns "does a free copy exist?" into a field check.
The job-to-be-done: for a reading list or citation corpus, which papers have an open PDF — and where.
How does this compare to the alternatives?
| Manual per-paper search | Browser extensions | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free, slow | Free, per-page | Pay per result |
| Reliability | One at a time | Only while browsing | Batch export |
| Setup time | Zero | Extension install | Minutes |
| Maintenance | Every paper | Per lookup | Re-run the list |
Extensions help when you're already on a page. For a hundred-paper list, you want the links in a table.
How to locate open PDFs in 4 steps
How do I check a reading list?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/google-scholar-scraper").call(
run_input={
"queries": [
'"Scaling Instruction-Finetuned Language Models"',
'"LoRA: Low-Rank Adaptation of Large Language Models"',
],
"maxResults": 5,
}
)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())How do I split open vs paywalled?
import pandas as pd
df = pd.DataFrame(rows).drop_duplicates("title")
open_oa = df[df.pdfUrl.notna()][["title", "pdfUrl", "venue", "year"]]
closed = df[df.pdfUrl.isna()][["title", "url", "venue"]]
print(f"{len(open_oa)} open / {len(closed)} closed")How do I fetch the open copies?
import httpx, pathlib
for r in open_oa.to_dict("records"):
try:
data = httpx.get(r["pdfUrl"], timeout=30, follow_redirects=True).content
pathlib.Path("pdfs").mkdir(exist_ok=True)
pathlib.Path("pdfs", r["title"][:50] + ".pdf").write_bytes(data)
except Exception:
passHow do I handle the closed ones?
url is the canonical page — institutional access, author copy, or interlibrary loan start there. The split itself is the useful deliverable: you know which papers need a workaround and which don't.
Sample output
{"title": "LoRA: Low-Rank Adaptation of Large Language Models",
"url": "https://arxiv.org/abs/2106.09685",
"authors": "E Hu, Y Shen", "venue": "ICLR", "year": "2022",
"citedBy": 11000, "pdfUrl": "https://arxiv.org/pdf/2106.09685",
"query": "\"LoRA: Low-Rank Adaptation\"", "position": 1}Common pitfalls
pdfUrl absence doesn't mean closed — Scholar sometimes indexes a version without linking the PDF. Repository links rot; fetch promptly and keep the record. Publisher-hosted "free" PDFs may still carry license terms — read them before redistribution. The actor locates; rights and access rules remain yours.
Related use cases
Frequently asked questions
Where do the pdfUrl links point?
To the PDF Scholar indexed — usually an arXiv, institutional repository, or publisher open-access copy. It is the legal open version, not a paywall bypass.
What fraction of papers have a pdfUrl?
It varies by field — heavy in CS and physics (arXiv), thinner in paywalled disciplines. Filter records where pdfUrl is present; the rest point to their canonical page.
Can I search for a specific paper's free version?
Yes — query the exact title. If Scholar knows an open copy, its record carries pdfUrl alongside the canonical url.
Does the actor download PDFs?
No — it returns the links. Fetching is a separate step on the hosting repository, which keeps provenance clean.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.