Export X Community Notes for Misinformation Research
Pull X's public Community Notes dataset as filtered JSON — classification, reasons, showing status — for misinformation and platform-governance research.

TL;DR — The X Community Notes Scraper turns X's public fact-check dataset into filtered JSON: summary, classification, reasons, and whether the note is actually showing. Built for misinformation and governance research without the bulk-file wrangling.
Why Community Notes is a research-grade dataset
Community Notes is the largest public experiment in crowdsourced fact-checking — X publishes the underlying data openly, and researchers have used it to study consensus, correction speed, and political asymmetry. The raw files are large multi-gigabyte TSVs; the questions researchers actually ask are filtered slices.
The job-to-be-done: query the dataset — misleading-only, this-month-only, showing-on-X-only — without downloading and parsing the whole thing each time.
How does this compare to the alternatives?
| Raw TSV downloads | Social-listening tools | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free + your ETL | Subscriptions | Pay per result |
| Reliability | Gigabytes per pull | Platform-dependent | Filtered slices |
| Setup time | Hours of parsing | Procurement | Minutes |
| Maintenance | Rebuild pipeline | Vendor | Handled |
How to export a research slice in 4 steps
How do I pull a filtered corpus?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/x-community-notes-scraper").call(
run_input={
"classification": "misleading",
"onlyShowingOnX": True,
"includeStatus": True,
"sinceDate": "2026-08-01",
"maxResults": 1000,
}
)
notes = list(client.dataset(run["defaultDatasetId"]).iterate_items())How do I profile what gets flagged?
import pandas as pd
df = pd.DataFrame(notes)
reasons = df["reasons"].explode().value_counts()
print(reasons.head(10))The reasons distribution — factualError vs missingContext vs manipulatedMedia — is the corpus's typology.
How do I bound by date or posts?
sinceDate limits creation date; tweetIds targets specific posts; snapshotDate pins a historical dataset version for reproducibility — the field a methods section needs.
How do I join back to posts?
tweetUrl is a handle-free permalink per record — openable, linkable, and joinable against any post metadata you hold.
Sample output
{"noteId": "1834567890123456789", "tweetId": "1834560000000000000",
"tweetUrl": "https://x.com/i/status/1834560000000000000",
"summary": "The video predates the event by three years…",
"classification": "MISINFORMED_OR_POTENTIALLY_MISLEADING",
"isMisleading": true, "reasons": ["misleadingOutdatedInformation"],
"trustworthySources": true, "isMediaNote": true,
"createdAt": "2026-08-14T09:20:00Z",
"noteUrl": "https://x.com/i/birdwatch/n/1834567890123456789",
"status": "CURRENTLY_RATED_HELPFUL", "isShowingOnX": true,
"statusLabel": "Currently rated helpful"}Common pitfalls
The public dataset updates on X's schedule — snapshotDate pins reproducibility when versions matter. isShowingOnX requires includeStatus — status arrives from a separate public file. Note text is community-written — quality and bias vary; that's the research object, not a data bug. The actor filters at source; huge unfiltered pulls still take time.
Related use cases
Frequently asked questions
Where does the data come from?
X's own public Community Notes dataset — the same TSV files X publishes for transparency. The actor parses them into filtered, fielded JSON so you skip the bulk download.
What filters are available?
classification (misleading/not_misleading/all), searchText, tweetIds, sinceDate, mediaNotesOnly, onlyShowingOnX, includeStatus, includeHistorical, snapshotDate, and maxResults.
What fields come per note?
noteId, tweetId, tweetUrl, summary, classification, isMisleading, reasons array, trustworthySources, believable, harmful, validationDifficulty, media/collaboration flags, createdAt, noteUrl — plus status and isShowingOnX when includeStatus is on.
Do I need X credentials?
No — the source is X's public dataset download. No login, cookies, or API key.