Track Citation Counts for Research Impact Over Time
Export Google Scholar citation counts per paper as JSON on a schedule — measure how citations accrue across a paper list without manual lookups.

TL;DR — The Google Scholar Scraper returns
citedByper paper as JSON. Query your paper list on a schedule, store snapshots, and citation growth becomes a measured series — for tenure files, grant reports, and lab dashboards.
Why citation tracking beats checking Scholar by hand
Impact evidence arrives on a deadline: a tenure packet, a progress report, a grant renewal. "I checked Scholar" isn't a series — a dated count per paper is.
The job-to-be-done is refreshable snapshots: same queries, same fields, new counts — joined by title into a growth table.
How does this compare to the alternatives?
| Manual Scholar checks | Citation databases | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free, tedious | Institutional license | Pay per result |
| Reliability | No history | Curated, narrower | Broad, dated |
| Setup time | Zero | Library access | Minutes |
| Maintenance | Every report | Vendor | Scheduled runs |
Scopus and Web of Science count fewer, cleaner citations. Scholar counts more, sooner — for trend reporting, breadth is the feature.
How to track citations in 4 steps
How do I snapshot a paper list?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
titles = [
'"Attention Is All You Need"',
'"BERT: Pre-training of Deep Bidirectional Transformers"',
]
run = client.actor("thirdwatch/google-scholar-scraper").call(
run_input={"queries": titles, "maxResults": 5}
)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())Quote the title — position 1 is almost always the target paper.
How do I store the series?
import json, datetime, pathlib
snap = {"date": datetime.date.today().isoformat(),
"counts": {r["title"]: r["citedBy"] for r in rows}}
pathlib.Path("citations.jsonl").open("a").write(json.dumps(snap) + "\n")Weekly appends build the series with zero rework.
How do I chart growth?
import pandas as pd
snaps = pd.read_json("citations.jsonl", lines=True)
growth = pd.json_normalize(snaps.set_index("date")["counts"]).T
growth.columns = snaps["date"]
print(growth.assign(delta=growth.iloc[:, -1] - growth.iloc[:, 0]))How do I scope by year?
fromYear/toYear bound the search — useful when the same title exists across editions and you want the canonical recent version.
Sample output
{"title": "Attention Is All You Need",
"url": "https://arxiv.org/abs/1706.03762",
"authors": "A Vaswani, N Shazeer", "venue": "NeurIPS",
"year": "2017", "citedBy": 130000,
"pdfUrl": "https://arxiv.org/pdf/1706.03762",
"query": "\"Attention Is All You Need\"", "position": 1}Common pitfalls
Scholar merges versions — citedBy is the union across arXiv, proceedings, and journal forms, which is exactly why trends work well. Homonymous titles occasionally return the wrong paper at position 1; sanity-check authors. Counts update on Scholar's schedule, not yours — compare weekly, not daily. The actor reads public result pages; volume modesty keeps it reliable.
Related use cases
Frequently asked questions
How do I track one specific paper's citations?
Query the exact paper title in quotes — the top result is usually the paper itself, and its citedBy field is the current count. Re-run on a schedule for the time series.
Can I track a whole lab's papers?
Batch the titles in queries — each record returns citedBy with the query attached, so a single run refreshes the whole list.
How accurate is Scholar's citedBy?
Scholar counts broadly — preprints and non-traditional venues included — so it's typically higher than Scopus or Web of Science. Track it as a trend, not an absolute.
How often do counts update?
Scholar updates continuously. Weekly runs catch meaningful movement; daily runs mostly show noise.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.