Skip to main content
Thirdwatchthirdwatch
Other

Monitor New Papers in Your Research Field Automatically

Run Google Scholar queries on a schedule and export new papers — title, authors, venue, PDF link — as JSON so field awareness stops depending on memory.

Sep 21, 2026 · 2 min read · 487 words
See the scraper →

TL;DR — The Google Scholar Scraper exports Scholar results as structured JSON. Run your field's queries weekly, diff titles against your stored set, and new papers land in a table — not an inbox you skim.

Why field awareness needs automation

Keeping up with a field means scanning the same queries every week and spotting what's new. Scholar alerts email you links; they don't give you a deduped dataset, a diff, or a filterable table.

The job-to-be-done: a weekly run whose output answers one question — which titles weren't here last week?

How does this compare to the alternatives?

Scholar email alerts Journal RSS feeds Thirdwatch actor
Cost Free Free Pay per result
Reliability Unstructured Venue-scoped only Structured JSON
Setup time Zero Per journal Minutes
Maintenance Inbox archaeology Feed sprawl Scheduled runs

Alerts and RSS are fine for casual awareness. For a lab or analyst workflow, records beat emails.

How to monitor new papers in 4 steps

How do I define the watch set?

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
QUERIES = ["speculative decoding LLM", "KV cache compression"]
run = client.actor("thirdwatch/google-scholar-scraper").call(
    run_input={"queries": QUERIES, "fromYear": 2026, "maxResults": 50}
)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())

fromYear keeps the universe to the current literature — the set that changes week to week.

How do I diff for arrivals?

import json, pathlib
seen_file = pathlib.Path("seen_titles.json")
seen = set(json.loads(seen_file.read_text() or "[]"))
new = [r for r in rows if r["title"] not in seen]
seen |= {r["title"] for r in rows}
seen_file.write_text(json.dumps(sorted(seen)))
print(f"{len(new)} new papers this run")

How do I triage what arrived?

for r in sorted(new, key=lambda x: -(x.get("citedBy") or 0))[:10]:
    print(r["citedBy"], r["venue"], "-", r["title"], r["pdfUrl"] or r["url"])

Sort by citations to surface breakout preprints first; pdfUrl gets you to full text without paywall detours where Scholar exposes it.

How do I run it unattended?

Save the input as an Apify Task with a weekly schedule — the dataset lands every week and your diff script consumes it.

Sample output

{"title": "Fast Inference from Transformers via Speculative Decoding",
 "url": "https://arxiv.org/abs/2211.17192",
 "authors": "Y Leviathan, M Kalman", "venue": "arXiv",
 "year": "2026", "snippet": "…accelerating sampling without changing outputs…",
 "citedBy": 2140, "pdfUrl": "https://arxiv.org/pdf/2211.17192",
 "query": "speculative decoding LLM", "position": 2}

Common pitfalls

Scholar indexes lag publication by days to weeks — brand-new papers appear late. Title edits between preprint versions create false "new" hits; dedupe loosely if that matters to you. Broad queries return noise — tighten with quoted phrases or author names. The actor fetches the list; reading remains gloriously yours.

Related use cases

Frequently asked questions

How is this different from Scholar's email alerts?

Alerts email links; this returns structured records — title, authors, venue, year, citedBy, pdfUrl — that feed a spreadsheet, filter pipeline, or reading-list tool directly.

How do I get only recent papers?

Set fromYear to the current year, and diff each run's titles against your stored set — new titles are the week's arrivals.

Can I watch several subtopics at once?

Yes — queries takes a list, and every record carries its query so subtopic streams stay separated in one run.

Does it find preprints?

Scholar indexes arXiv and other preprint servers alongside formal venues — the venue field distinguishes them.

Related

Try it yourself

100 free credits, no credit card.

About 30 real searches. Add the MCP to Claude or Cursor in two minutes.