Skip to main content
Thirdwatchthirdwatch
Social & news

Build a Dataset of Fact-Check Notes That Actually Show

Export only the Community Notes currently displayed on X — rated-helpful corrections with text, reasons, and sources — as clean JSON.

Editorial illustration for social & news
Sep 21, 2026 · 2 min read · 473 words
View the Apify scraper →

TL;DR — The X Community Notes Scraper filters to notes currently displayed on X — the rated-helpful corrections users actually see. onlyShowingOnX plus includeStatus gives a clean, consensus-verified fact-check dataset as JSON.

Why the visible set is the dataset that matters

Most Community Notes never display — they stall awaiting ratings or fail consensus. A dataset of all notes mixes signal with drafts; the showing set is what corrected anything in the real world.

The job-to-be-done: the corpus where every record is a correction that earned display — for research baselines, product features, and accountability reporting.

How does this compare to the alternatives?

Full TSV + self-filter Note-by-note browsing Thirdwatch actor
Cost Free + join complexity Free, unscalable Pay per result
Reliability You wire status joins Complete, manual Filtered at source
Setup time Hours Zero Minutes
Maintenance Pipeline None Handled

The status data lives in a separate public file — the actor already joins it, so isShowingOnX arrives as a field, not a second ETL job.

How to build the visible corpus in 4 steps

How do I pull only showing notes?

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/x-community-notes-scraper").call(
    run_input={
        "onlyShowingOnX": True,
        "includeStatus": True,
        "classification": "misleading",
        "sinceDate": "2026-07-01",
        "maxResults": 1000,
    }
)
notes = list(client.dataset(run["defaultDatasetId"]).iterate_items())

How do I verify the filter worked?

assert all(n["isShowingOnX"] for n in notes)
import pandas as pd
df = pd.DataFrame(notes)
print(df["statusLabel"].value_counts())

Every row should read "Currently rated helpful" — the assertion is the QA gate.

How do I keep the evidence links?

tweetUrl and noteUrl per record — the correction's exhibit and its target. trustworthySources flags whether the note cites sources — the quality signal researchers weight.

How do I version it?

snapshotDate pins the underlying dataset — write it into your export filename so the corpus stays reproducible.

Sample output

{"noteId": "1834567890123456789", "tweetId": "1834560000000000000",
 "tweetUrl": "https://x.com/i/status/1834560000000000000",
 "summary": "This chart reverses the axes; the actual figures show a decline…",
 "classification": "MISINFORMED_OR_POTENTIALLY_MISLEADING",
 "isMisleading": true,
 "reasons": ["misleadingFactualError", "misleadingMissingImportantContext"],
 "trustworthySources": true, "isMediaNote": true,
 "createdAt": "2026-08-19T11:30:00Z",
 "noteUrl": "https://x.com/i/birdwatch/n/1834567890123456789",
 "status": "CURRENTLY_RATED_HELPFUL", "isShowingOnX": true,
 "statusLabel": "Currently rated helpful"}

Common pitfalls

Showing status is a point-in-time snapshot — notes can lose consensus later; re-run with snapshotDate tracked for longitudinal honesty. The visible set skews toward sourced, factual corrections — that's the consensus mechanism's bias, knowable via the full set. High-volume pulls take time; bound with sinceDate. The actor reads the public dataset; visibility on X itself remains X's call.

Related use cases

Frequently asked questions

What does 'showing on X' mean exactly?

Notes rated helpful enough to display under the post — status CURRENTLY_RATED_HELPFUL. onlyShowingOnX filters to just these, so your dataset contains corrections users actually saw.

Why filter to showing notes?

Most written notes never reach consensus and never display. For downstream use — training, display, reporting — the visible set is the quality-filtered one.

Can I also get media-only corrections?

Yes — mediaNotesOnly restricts to notes on image/video posts, the most common misinformation vector.

How do I pin a reproducible snapshot?

snapshotDate selects the dataset version — the field that makes the corpus citeable and re-runnable.

Related