Monitor New Papers in Your Research Field Automatically
Run Google Scholar queries on a schedule and export new papers — title, authors, venue, PDF link — as JSON so field awareness stops depending on memory.

TL;DR — The Google Scholar Scraper exports Scholar results as structured JSON. Run your field's queries weekly, diff titles against your stored set, and new papers land in a table — not an inbox you skim.
Why field awareness needs automation
Keeping up with a field means scanning the same queries every week and spotting what's new. Scholar alerts email you links; they don't give you a deduped dataset, a diff, or a filterable table.
The job-to-be-done: a weekly run whose output answers one question — which titles weren't here last week?
How does this compare to the alternatives?
| Scholar email alerts | Journal RSS feeds | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free | Free | Pay per result |
| Reliability | Unstructured | Venue-scoped only | Structured JSON |
| Setup time | Zero | Per journal | Minutes |
| Maintenance | Inbox archaeology | Feed sprawl | Scheduled runs |
Alerts and RSS are fine for casual awareness. For a lab or analyst workflow, records beat emails.
How to monitor new papers in 4 steps
How do I define the watch set?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
QUERIES = ["speculative decoding LLM", "KV cache compression"]
run = client.actor("thirdwatch/google-scholar-scraper").call(
run_input={"queries": QUERIES, "fromYear": 2026, "maxResults": 50}
)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())fromYear keeps the universe to the current literature — the set that changes week to week.
How do I diff for arrivals?
import json, pathlib
seen_file = pathlib.Path("seen_titles.json")
seen = set(json.loads(seen_file.read_text() or "[]"))
new = [r for r in rows if r["title"] not in seen]
seen |= {r["title"] for r in rows}
seen_file.write_text(json.dumps(sorted(seen)))
print(f"{len(new)} new papers this run")How do I triage what arrived?
for r in sorted(new, key=lambda x: -(x.get("citedBy") or 0))[:10]:
print(r["citedBy"], r["venue"], "-", r["title"], r["pdfUrl"] or r["url"])Sort by citations to surface breakout preprints first; pdfUrl gets you to full text without paywall detours where Scholar exposes it.
How do I run it unattended?
Save the input as an Apify Task with a weekly schedule — the dataset lands every week and your diff script consumes it.
Sample output
{"title": "Fast Inference from Transformers via Speculative Decoding",
"url": "https://arxiv.org/abs/2211.17192",
"authors": "Y Leviathan, M Kalman", "venue": "arXiv",
"year": "2026", "snippet": "…accelerating sampling without changing outputs…",
"citedBy": 2140, "pdfUrl": "https://arxiv.org/pdf/2211.17192",
"query": "speculative decoding LLM", "position": 2}Common pitfalls
Scholar indexes lag publication by days to weeks — brand-new papers appear late. Title edits between preprint versions create false "new" hits; dedupe loosely if that matters to you. Broad queries return noise — tighten with quoted phrases or author names. The actor fetches the list; reading remains gloriously yours.
Related use cases
Frequently asked questions
How is this different from Scholar's email alerts?
Alerts email links; this returns structured records — title, authors, venue, year, citedBy, pdfUrl — that feed a spreadsheet, filter pipeline, or reading-list tool directly.
How do I get only recent papers?
Set fromYear to the current year, and diff each run's titles against your stored set — new titles are the week's arrivals.
Can I watch several subtopics at once?
Yes — queries takes a list, and every record carries its query so subtopic streams stay separated in one run.
Does it find preprints?
Scholar indexes arXiv and other preprint servers alongside formal venues — the venue field distinguishes them.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.