Build a Dataset of Fact-Check Notes That Actually Show
Export only the Community Notes currently displayed on X — rated-helpful corrections with text, reasons, and sources — as clean JSON.

TL;DR — The X Community Notes Scraper filters to notes currently displayed on X — the rated-helpful corrections users actually see.
onlyShowingOnXplusincludeStatusgives a clean, consensus-verified fact-check dataset as JSON.
Why the visible set is the dataset that matters
Most Community Notes never display — they stall awaiting ratings or fail consensus. A dataset of all notes mixes signal with drafts; the showing set is what corrected anything in the real world.
The job-to-be-done: the corpus where every record is a correction that earned display — for research baselines, product features, and accountability reporting.
How does this compare to the alternatives?
| Full TSV + self-filter | Note-by-note browsing | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free + join complexity | Free, unscalable | Pay per result |
| Reliability | You wire status joins | Complete, manual | Filtered at source |
| Setup time | Hours | Zero | Minutes |
| Maintenance | Pipeline | None | Handled |
The status data lives in a separate public file — the actor already joins it, so isShowingOnX arrives as a field, not a second ETL job.
How to build the visible corpus in 4 steps
How do I pull only showing notes?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/x-community-notes-scraper").call(
run_input={
"onlyShowingOnX": True,
"includeStatus": True,
"classification": "misleading",
"sinceDate": "2026-07-01",
"maxResults": 1000,
}
)
notes = list(client.dataset(run["defaultDatasetId"]).iterate_items())How do I verify the filter worked?
assert all(n["isShowingOnX"] for n in notes)
import pandas as pd
df = pd.DataFrame(notes)
print(df["statusLabel"].value_counts())Every row should read "Currently rated helpful" — the assertion is the QA gate.
How do I keep the evidence links?
tweetUrl and noteUrl per record — the correction's exhibit and its target. trustworthySources flags whether the note cites sources — the quality signal researchers weight.
How do I version it?
snapshotDate pins the underlying dataset — write it into your export filename so the corpus stays reproducible.
Sample output
{"noteId": "1834567890123456789", "tweetId": "1834560000000000000",
"tweetUrl": "https://x.com/i/status/1834560000000000000",
"summary": "This chart reverses the axes; the actual figures show a decline…",
"classification": "MISINFORMED_OR_POTENTIALLY_MISLEADING",
"isMisleading": true,
"reasons": ["misleadingFactualError", "misleadingMissingImportantContext"],
"trustworthySources": true, "isMediaNote": true,
"createdAt": "2026-08-19T11:30:00Z",
"noteUrl": "https://x.com/i/birdwatch/n/1834567890123456789",
"status": "CURRENTLY_RATED_HELPFUL", "isShowingOnX": true,
"statusLabel": "Currently rated helpful"}Common pitfalls
Showing status is a point-in-time snapshot — notes can lose consensus later; re-run with snapshotDate tracked for longitudinal honesty. The visible set skews toward sourced, factual corrections — that's the consensus mechanism's bias, knowable via the full set. High-volume pulls take time; bound with sinceDate. The actor reads the public dataset; visibility on X itself remains X's call.
Related use cases
Frequently asked questions
What does 'showing on X' mean exactly?
Notes rated helpful enough to display under the post — status CURRENTLY_RATED_HELPFUL. onlyShowingOnX filters to just these, so your dataset contains corrections users actually saw.
Why filter to showing notes?
Most written notes never reach consensus and never display. For downstream use — training, display, reporting — the visible set is the quality-filtered one.
Can I also get media-only corrections?
Yes — mediaNotesOnly restricts to notes on image/video posts, the most common misinformation vector.
How do I pin a reproducible snapshot?
snapshotDate selects the dataset version — the field that makes the corpus citeable and re-runnable.