Build a Multilingual Wikidata Entity Dataset
Collect language-specific Wikidata labels, descriptions, aliases, and Wikipedia links for entity catalogs.

Multilingual entity catalogs need stable identity across languages while allowing labels and descriptions to vary. Translating an English name is not the same as retrieving a community-maintained local label. The Wikidata Entity Scraper lets each run select a language while preserving the same Q-IDs.
Collect one language per reproducible run
Use reviewed IDs when possible, or search in the target language for discovery.
{"queries":["inteligencia artificial","cambio climático"],"language":"es","maxResultsPerQuery":10}Store Q-ID, language code, label, description, aliases, and local Wikipedia URL as a language-specific record. Never use a translated label as the cross-language join key. If a local label is absent, decide explicitly whether the product should fall back to English, show the Q-ID, or leave the field blank.
Validate ambiguous search results with descriptions and claims. Different languages can rank candidates differently, and the same spelling can identify unrelated entities. Once a candidate is approved, move it into the exact-ID workflow for every language.
Run separate Tasks for priority languages and merge on Q-ID plus language. Track missing coverage rather than filling it with undocumented machine translation. Sitelink counts indicate breadth but not translation quality.
Community language data changes, so snapshots need collection dates and source links. Review sensitive or high-visibility fields before publication. This workflow provides a strong multilingual foundation while keeping identity, localization, and editorial approval as distinct responsibilities.
Report coverage by language and field type, since a translated label does not guarantee that a useful description, aliases, or local sitelink also exists.
Frequently asked questions
Does every entity have every language?
No. Labels, descriptions, aliases, and Wikipedia sitelinks have uneven coverage.
Is machine translation added automatically?
No. The Actor returns language data published by Wikidata.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.