Skip to main content
Thirdwatchthirdwatch
Social & news

Export X Community Notes for Misinformation Research

Pull X's public Community Notes dataset as filtered JSON — classification, reasons, showing status — for misinformation and platform-governance research.

Editorial illustration for social & news
Sep 21, 2026 · 2 min read · 444 words
View the Apify scraper →

TL;DR — The X Community Notes Scraper turns X's public fact-check dataset into filtered JSON: summary, classification, reasons, and whether the note is actually showing. Built for misinformation and governance research without the bulk-file wrangling.

Why Community Notes is a research-grade dataset

Community Notes is the largest public experiment in crowdsourced fact-checking — X publishes the underlying data openly, and researchers have used it to study consensus, correction speed, and political asymmetry. The raw files are large multi-gigabyte TSVs; the questions researchers actually ask are filtered slices.

The job-to-be-done: query the dataset — misleading-only, this-month-only, showing-on-X-only — without downloading and parsing the whole thing each time.

How does this compare to the alternatives?

Raw TSV downloads Social-listening tools Thirdwatch actor
Cost Free + your ETL Subscriptions Pay per result
Reliability Gigabytes per pull Platform-dependent Filtered slices
Setup time Hours of parsing Procurement Minutes
Maintenance Rebuild pipeline Vendor Handled

How to export a research slice in 4 steps

How do I pull a filtered corpus?

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/x-community-notes-scraper").call(
    run_input={
        "classification": "misleading",
        "onlyShowingOnX": True,
        "includeStatus": True,
        "sinceDate": "2026-08-01",
        "maxResults": 1000,
    }
)
notes = list(client.dataset(run["defaultDatasetId"]).iterate_items())

How do I profile what gets flagged?

import pandas as pd
df = pd.DataFrame(notes)
reasons = df["reasons"].explode().value_counts()
print(reasons.head(10))

The reasons distribution — factualError vs missingContext vs manipulatedMedia — is the corpus's typology.

How do I bound by date or posts?

sinceDate limits creation date; tweetIds targets specific posts; snapshotDate pins a historical dataset version for reproducibility — the field a methods section needs.

How do I join back to posts?

tweetUrl is a handle-free permalink per record — openable, linkable, and joinable against any post metadata you hold.

Sample output

{"noteId": "1834567890123456789", "tweetId": "1834560000000000000",
 "tweetUrl": "https://x.com/i/status/1834560000000000000",
 "summary": "The video predates the event by three years…",
 "classification": "MISINFORMED_OR_POTENTIALLY_MISLEADING",
 "isMisleading": true, "reasons": ["misleadingOutdatedInformation"],
 "trustworthySources": true, "isMediaNote": true,
 "createdAt": "2026-08-14T09:20:00Z",
 "noteUrl": "https://x.com/i/birdwatch/n/1834567890123456789",
 "status": "CURRENTLY_RATED_HELPFUL", "isShowingOnX": true,
 "statusLabel": "Currently rated helpful"}

Common pitfalls

The public dataset updates on X's schedule — snapshotDate pins reproducibility when versions matter. isShowingOnX requires includeStatus — status arrives from a separate public file. Note text is community-written — quality and bias vary; that's the research object, not a data bug. The actor filters at source; huge unfiltered pulls still take time.

Related use cases

Frequently asked questions

Where does the data come from?

X's own public Community Notes dataset — the same TSV files X publishes for transparency. The actor parses them into filtered, fielded JSON so you skip the bulk download.

What filters are available?

classification (misleading/not_misleading/all), searchText, tweetIds, sinceDate, mediaNotesOnly, onlyShowingOnX, includeStatus, includeHistorical, snapshotDate, and maxResults.

What fields come per note?

noteId, tweetId, tweetUrl, summary, classification, isMisleading, reasons array, trustworthySources, believable, harmful, validationDifficulty, media/collaboration flags, createdAt, noteUrl — plus status and isShowingOnX when includeStatus is on.

Do I need X credentials?

No — the source is X's public dataset download. No login, cookies, or API key.

Related