Scrape Google Scholar for a Literature Review
Export Google Scholar results — titles, authors, venues, citation counts, PDF links — as JSON to systematize the search phase of a literature review.

TL;DR — The Google Scholar Scraper exports Scholar search results — title, authors, venue, year, citedBy, pdfUrl — as JSON. Run your full query list, dedupe on title, and the screening table for a literature review writes itself.
Why the search phase deserves a script
A systematic review's credibility rests on a reproducible search: every query logged, every result screenable. Google Scholar indexes the broadest slice of scholarly literature — but its interface exports one citation at a time.
The job-to-be-done is the result set as data: the queries you ran, the papers they returned, ranked and dated — an artifact your methods section can cite and a co-reviewer can re-run.
How does this compare to the alternatives?
| Manual Scholar search | Reference-manager import | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free, slow | Free, clunky | Pay per result |
| Reliability | One page at a time | RIS export quirks | Clean JSON |
| Setup time | Zero | Per-tool quirks | Minutes |
| Maintenance | Redo per update | Re-export | Re-run queries |
Reference managers import citations for known papers. The discovery step — "what does Scholar return for this string" — is what the actor captures.
How to run the review search in 4 steps
How do I export a query's results?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/google-scholar-scraper").call(
run_input={
"queries": ["retrieval augmented generation evaluation"],
"fromYear": 2023,
"maxResults": 50,
}
)
papers = list(client.dataset(run["defaultDatasetId"]).iterate_items())How do I batch the whole search strategy?
queries = [
"retrieval augmented generation evaluation",
"RAG hallucination benchmark",
"retrieval augmented factuality",
]
run = client.actor("thirdwatch/google-scholar-scraper").call(
run_input={"queries": queries, "fromYear": 2023, "maxResults": 50}
)Each record keeps its query — your PRISMA-style flow counts come free.
How do I build the screening table?
import pandas as pd
df = pd.DataFrame(papers).drop_duplicates("title")
df = df.sort_values(["query", "citedBy"], ascending=[True, False])
df[["title", "authors", "venue", "year", "citedBy", "url", "pdfUrl"]].to_csv("screening.csv", index=False)Dedupe on title (Scholar surfaces the same paper under multiple strings), then screen by venue and citations first.
How do I find open-access PDFs?
open_access = df[df["pdfUrl"].notna()]
print(f"{len(open_access)} with direct PDF")pdfUrl is the repository or publisher PDF link Scholar indexed — the fastest path to full text for screening.
Sample output
{"title": "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks",
"url": "https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html",
"authors": "P Lewis, E Perez, A Piktus", "venue": "NeurIPS",
"year": "2020", "snippet": "…a general-purpose fine-tuning recipe…",
"citedBy": 8412, "pdfUrl": "https://arxiv.org/pdf/2005.11401",
"query": "retrieval augmented generation evaluation", "position": 1}Common pitfalls
Scholar's counts merge editions — a paper's citedBy is a lower bound across versions. Results personalize slightly; keep hl-consistent queries and note the run date in your protocol. Some venues return no url or pdfUrl — that reflects Scholar's index, not a failure. The actor exports metadata; eligibility decisions stay with the reviewers.
Related use cases
Frequently asked questions
What fields come back per paper?
Title, URL, authors, venue, year, snippet, citedBy count, pdfUrl where Scholar exposes one, plus the query and rank position — the fields a screening table needs.
Can I restrict by publication year?
Yes. fromYear and toYear bound results — the standard date-window filter for systematic and rapid reviews.
Can I batch multiple search strings?
The queries array takes your full search strategy — every record carries its query so results stay attributable to the exact string that found them.
Does it fetch the papers themselves?
It returns metadata and pdfUrl links. Downloading PDFs happens separately, on the publisher or repository hosting them.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.