Skip to main content
Thirdwatchthirdwatch
Other

Scrape Google Scholar for a Literature Review

Export Google Scholar results — titles, authors, venues, citation counts, PDF links — as JSON to systematize the search phase of a literature review.

Sep 21, 2026 · 3 min read · 513 words
See the scraper →

TL;DR — The Google Scholar Scraper exports Scholar search results — title, authors, venue, year, citedBy, pdfUrl — as JSON. Run your full query list, dedupe on title, and the screening table for a literature review writes itself.

Why the search phase deserves a script

A systematic review's credibility rests on a reproducible search: every query logged, every result screenable. Google Scholar indexes the broadest slice of scholarly literature — but its interface exports one citation at a time.

The job-to-be-done is the result set as data: the queries you ran, the papers they returned, ranked and dated — an artifact your methods section can cite and a co-reviewer can re-run.

How does this compare to the alternatives?

Manual Scholar search Reference-manager import Thirdwatch actor
Cost Free, slow Free, clunky Pay per result
Reliability One page at a time RIS export quirks Clean JSON
Setup time Zero Per-tool quirks Minutes
Maintenance Redo per update Re-export Re-run queries

Reference managers import citations for known papers. The discovery step — "what does Scholar return for this string" — is what the actor captures.

How to run the review search in 4 steps

How do I export a query's results?

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/google-scholar-scraper").call(
    run_input={
        "queries": ["retrieval augmented generation evaluation"],
        "fromYear": 2023,
        "maxResults": 50,
    }
)
papers = list(client.dataset(run["defaultDatasetId"]).iterate_items())

How do I batch the whole search strategy?

queries = [
    "retrieval augmented generation evaluation",
    "RAG hallucination benchmark",
    "retrieval augmented factuality",
]
run = client.actor("thirdwatch/google-scholar-scraper").call(
    run_input={"queries": queries, "fromYear": 2023, "maxResults": 50}
)

Each record keeps its query — your PRISMA-style flow counts come free.

How do I build the screening table?

import pandas as pd
df = pd.DataFrame(papers).drop_duplicates("title")
df = df.sort_values(["query", "citedBy"], ascending=[True, False])
df[["title", "authors", "venue", "year", "citedBy", "url", "pdfUrl"]].to_csv("screening.csv", index=False)

Dedupe on title (Scholar surfaces the same paper under multiple strings), then screen by venue and citations first.

How do I find open-access PDFs?

open_access = df[df["pdfUrl"].notna()]
print(f"{len(open_access)} with direct PDF")

pdfUrl is the repository or publisher PDF link Scholar indexed — the fastest path to full text for screening.

Sample output

{"title": "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks",
 "url": "https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html",
 "authors": "P Lewis, E Perez, A Piktus", "venue": "NeurIPS",
 "year": "2020", "snippet": "…a general-purpose fine-tuning recipe…",
 "citedBy": 8412, "pdfUrl": "https://arxiv.org/pdf/2005.11401",
 "query": "retrieval augmented generation evaluation", "position": 1}

Common pitfalls

Scholar's counts merge editions — a paper's citedBy is a lower bound across versions. Results personalize slightly; keep hl-consistent queries and note the run date in your protocol. Some venues return no url or pdfUrl — that reflects Scholar's index, not a failure. The actor exports metadata; eligibility decisions stay with the reviewers.

Related use cases

Frequently asked questions

What fields come back per paper?

Title, URL, authors, venue, year, snippet, citedBy count, pdfUrl where Scholar exposes one, plus the query and rank position — the fields a screening table needs.

Can I restrict by publication year?

Yes. fromYear and toYear bound results — the standard date-window filter for systematic and rapid reviews.

Can I batch multiple search strings?

The queries array takes your full search strategy — every record carries its query so results stay attributable to the exact string that found them.

Does it fetch the papers themselves?

It returns metadata and pdfUrl links. Downloading PDFs happens separately, on the publisher or repository hosting them.

Related

Try it yourself

100 free credits, no credit card.

About 30 real searches. Add the MCP to Claude or Cursor in two minutes.