Build a Stack Exchange Q&A Dataset
Create a traceable multi-site Stack Exchange question dataset with stable identifiers, original HTML, clean text, tags, and engagement fields.

Define the sites, queries, tags, result limits, and collection date before running the Stack Exchange Questions Scraper. Preserve both body_html and body_text: HTML retains code and structure, while clean text is convenient for search and analysis.
Store a source table keyed by site and question ID. Put exploded tags in a link table. If repeated snapshots are collected, keep engagement history separately rather than overwriting score, views, and answer count.
The dataset contains user-generated content. Retain source links, respect attribution requirements in downstream publishing, and avoid using owner fields for consequential profiling. A question corpus is most useful when its provenance and collection scope remain attached.
Frequently asked questions
Does this collect answer bodies?
No. This Actor focuses on questions and answer metadata; accepted answer IDs are references, not answer text.
What key works across sites?
Use the combination of site and question ID.