Skip to main content
Thirdwatchthirdwatch
Build & connect

Build a Stack Exchange Q&A Dataset

Create a traceable multi-site Stack Exchange question dataset with stable identifiers, original HTML, clean text, tags, and engagement fields.

Editorial illustration for build & connect
Jul 21, 2026 · 1 min read · 108 words
View the Apify scraper →

Define the sites, queries, tags, result limits, and collection date before running the Stack Exchange Questions Scraper. Preserve both body_html and body_text: HTML retains code and structure, while clean text is convenient for search and analysis.

Store a source table keyed by site and question ID. Put exploded tags in a link table. If repeated snapshots are collected, keep engagement history separately rather than overwriting score, views, and answer count.

The dataset contains user-generated content. Retain source links, respect attribution requirements in downstream publishing, and avoid using owner fields for consequential profiling. A question corpus is most useful when its provenance and collection scope remain attached.

Frequently asked questions

Does this collect answer bodies?

No. This Actor focuses on questions and answer metadata; accepted answer IDs are references, not answer text.

What key works across sites?

Use the combination of site and question ID.

Related