Analyze Community Notes Consensus Patterns in Bulk
Export thousands of Community Notes as JSON — classification, reasons, sources, status — to study what earns 'rated helpful' and what never shows.

TL;DR — The X Community Notes Scraper exports filtered slices of the public notes dataset — with per-note reasons, source flags, and showing status — as JSON. That's the raw material for studying when crowdsourced correction actually reaches users.
Why consensus patterns matter
Community Notes' promise is that good corrections surface. Whether they do — and which kinds of claims get corrected versus contested forever — is an empirical question answerable from the public dataset X publishes.
The job-to-be-done is a corpus with the denominator: not just showing notes, but all notes with each one's status — so "what fraction of corrections reach users" is computable.
How does this compare to the alternatives?
| X's bulk TSV files | Academic snapshots | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free + ETL | Free, dated | Pay per result |
| Reliability | Full, heavy | Stale at publication | Filtered, current |
| Setup time | Hours | Search for one | Minutes |
| Maintenance | Pipeline | None (fixed) | Handled |
How to build the analysis corpus in 4 steps
How do I pull a denominator corpus?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/x-community-notes-scraper").call(
run_input={
"classification": "all",
"includeStatus": True,
"includeHistorical": True,
"sinceDate": "2026-06-01",
"maxResults": 1000,
}
)
notes = list(client.dataset(run["defaultDatasetId"]).iterate_items())How do I compute the showing rate?
import pandas as pd
df = pd.DataFrame(notes)
rate = df.groupby("isMisleading")["isShowingOnX"].mean()
print(rate)The gap between misleading-claim corrections that show and those that stall is the consensus finding.
How do I profile correction types?
reasons = df.explode("reasons")
tbl = reasons.groupby("reasons")["isShowingOnX"].agg(["count", "mean"])
print(tbl.sort_values("count", ascending=False).head(10))Reason categories correlate with consensus odds — sourced factual errors resolve; subjective framings linger contested.
How do I control for sources and media?
print(df.groupby(["trustworthySources", "isMediaNote"])["isShowingOnX"].mean())Source-cited notes on media posts behave differently than unsourced ones — the dataset's design lets you measure it.
Sample output
{"noteId": "1834567890123456789", "tweetId": "1834560000000000000",
"summary": "Official figures show the opposite trend…",
"classification": "MISINFORMED_OR_POTENTIALLY_MISLEADING",
"isMisleading": true,
"reasons": ["misleadingFactualError"],
"trustworthySources": true, "believable": "yes",
"validationDifficulty": "easy",
"isMediaNote": false, "isCollaborativeNote": false,
"createdAt": "2026-07-02T14:05:00Z",
"status": "NEEDS_MORE_RATINGS", "isShowingOnX": false,
"statusLabel": "Needs more ratings"}Common pitfalls
Status reflects the snapshot — a contested note may show later; longitudinal runs catch transitions. believable/harmful/validationDifficulty are contributor judgments, not ground truth. Volume is large — filter first, analyze second. The actor reads X's public data; rater-level data lives in a different X file beyond this export.
Related use cases
Frequently asked questions
What does a bulk analysis reveal?
Which reason categories dominate, how often notes cite trustworthy sources, what share actually show on X, and how classification splits — the dataset's shape at scale rather than anecdote.
How do I get the showing-rate denominator?
Run with includeStatus on and no onlyShowingOnX filter — you get all notes plus each one's current visibility. The ratio is the consensus rate.
Can I study media vs text posts?
isMediaNote flags notes on media posts; mediaNotesOnly restricts to them. Both directions are one filter away.
How do I reproduce an analysis later?
snapshotDate pins the dataset version — publish it in your methods and anyone re-running the same filters gets the same corpus.