Skip to main content
Thirdwatchthirdwatch
Other

Build a Research Reading List From Scholar Data

Turn Google Scholar queries into a prioritized, deduplicated reading list — citations, venues, PDF links, and snippets as structured JSON.

Sep 21, 2026 · 2 min read · 473 words
See the scraper →

TL;DR — The Google Scholar Scraper turns a topic's queries into a structured paper list — citations, venues, snippets, PDF links. Dedupe across queries, sort by impact, and your next reading list is a ranked table, not 40 open tabs.

Why reading lists fail without structure

Topic exploration generates tabs: forty Scholar results, no ranking rationale, no dedup. The list that survives is the one that's a table — every paper scored by citations, filtered by year, with a snippet explaining why it matched.

The job-to-be-done is a merge of queries into one prioritized list — the papers multiple search angles agree on rising first.

How does this compare to the alternatives?

Browser bookmarks Reference managers Thirdwatch actor
Cost Free Free Pay per result
Reliability Rot fast Library-centric Query-driven
Setup time Zero Import friction Minutes
Maintenance Manual Sync quirks Re-run queries

Reference managers organize papers you've chosen. The actor builds the candidate set you choose from.

How to build the list in 4 steps

How do I cover a topic from multiple angles?

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/google-scholar-scraper").call(
    run_input={
        "queries": [
            "vector database indexing",
            "approximate nearest neighbor search",
            "embedding retrieval at scale",
        ],
        "fromYear": 2022,
        "maxResults": 30,
    }
)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())

How do I dedupe and rank?

import pandas as pd
df = pd.DataFrame(rows)
counts = df.groupby("title")["query"].nunique()
df["hits"] = df.title.map(counts)
top = (df.drop_duplicates("title")
       .sort_values(["hits", "citedBy"], ascending=False))
print(top[["title", "venue", "year", "citedBy", "hits"]].head(15))

Papers surfaced by multiple queries (hits > 1) are the topic's center of gravity — read those first.

How do I attach reading context?

top["pdf"] = top["pdfUrl"]
top[["title", "authors", "year", "citedBy", "pdf", "snippet"]].to_csv(
    "reading_list.csv", index=False)

snippet carries Scholar's matched excerpt — usually enough to triage without opening the paper.

How do I keep it current?

Monthly reruns with fromYear bumped catch new entrants; diff on title to append rather than rebuild.

Sample output

{"title": "Billion-scale similarity search with GPUs",
 "url": "https://arxiv.org/abs/1702.08734",
 "authors": "J Johnson, M Douze, H Jégou", "venue": "IEEE TBD",
 "year": "2019", "snippet": "…GPU-accelerated nearest neighbor…",
 "citedBy": 3900, "pdfUrl": "https://arxiv.org/pdf/1702.08734",
 "query": "approximate nearest neighbor search", "position": 3}

Common pitfalls

Citation counts favor older work — weight year when the field moves fast. Title dedupe misses editions; a fuzzy match on title+first author catches them. snippet is query-biased — it shows the match, not the abstract. The actor supplies the candidate set; the reading remains your craft.

Related use cases

Frequently asked questions

How do I prioritize what to read first?

Sort the exported records by citedBy within each query — high-citation anchors first — then scan snippets for topical fit. Venue and year fields let you weight recency versus authority.

Can I merge results across many queries?

Yes. Every record carries its query; dedupe on title across the union to see which papers multiple angles surfaced — a strong read-first signal.

Does it include abstracts?

The snippet field gives Scholar's query-matched excerpt. Full abstracts live at the paper's url — the export tells you which are worth opening.

How do I keep the list fresh?

Re-run monthly with fromYear set to the current year and diff titles — new arrivals append to the list without disturbing what you've triaged.

Related

Try it yourself

100 free credits, no credit card.

About 30 real searches. Add the MCP to Claude or Cursor in two minutes.