Build an Image Dataset for Computer Vision Training
Collect image URLs, dimensions, and source pages from Google Images as JSON to bootstrap a labeled computer-vision dataset without manual saving.

TL;DR — The Google Images Scraper returns image URLs, dimensions, source pages, and domains as JSON for any query list. Filter by size, dedupe, download, label — a bootstrap path for computer-vision datasets without hand-saving thumbnails.
Why image collection is the slow part of CV
Every vision project starts the same way: someone spends a weekend right-clicking "save image as". Google's image index covers essentially the whole public web — the constraint was never supply, it is collection mechanics.
The job-to-be-done is a URL list with metadata: image address, pixel dimensions, and provenance. With those three fields you can filter junk before downloading, dedupe by URL, and credit sources.
How does this compare to the alternatives?
| Manual saving | Paid dataset vendors | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free, your weekend | $$$$ per dataset | Pay per result |
| Reliability | Repetitive | Curated | Fresh index |
| Setup time | Zero | Contract negotiation | Minutes |
| Maintenance | Redo per class | Vendor refresh | Re-run queries |
Vendor datasets are curated but generic. When your class is "warehouse dock damage" or "handloom weaving defects", you collect your own — this is the collection step.
How to build the dataset in 4 steps
How do I pull URLs for my classes?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("thirdwatch/google-images-scraper").call(
run_input={
"queries": ["dock door damage", "loading dock crack", "warehouse floor damage"],
"maxResults": 100,
"country": "us",
"language": "en",
}
)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())Each record carries query — your class label — alongside imageUrl, width, height, sourcePageUrl, domain.
How do I filter before downloading?
import pandas as pd
df = pd.DataFrame(rows).drop_duplicates("imageUrl")
usable = df[(df["width"] >= 400) & (df["height"] >= 400)]
print(usable.groupby("query").size())Small thumbnails waste download budget; width/height let you enforce a floor first.
How do I download with provenance?
import httpx, pathlib
for r in usable.to_dict("records"):
ext = r["imageUrl"].split("?")[0].rsplit(".", 1)[-1][:4]
p = pathlib.Path("data") / r["query"].replace(" ", "_")
p.mkdir(parents=True, exist_ok=True)
try:
p.joinpath(f"{r['position']:04d}.{ext}").write_bytes(
httpx.get(r["imageUrl"], timeout=15, follow_redirects=True).content)
except Exception:
continue # dead links are normal — log rate, keep goingKeep a CSV of imageUrl, sourcePageUrl, domain next to the files — provenance is your license audit trail.
How do I grow it?
Rerun weekly with new query phrasings. Union on imageUrl across runs; the same image surfacing under different phrasings is a hint it is canonical for the class.
Sample output
{"imageUrl": "https://example-warehouse.com/img/dock_damage_04.jpg",
"width": 1280, "height": 960,
"sourcePageUrl": "https://example-warehouse.com/safety-inspection",
"domain": "example-warehouse.com",
"query": "dock door damage", "position": 7}domain and sourcePageUrl are your licensing shortlist — where to look for reuse terms.
Common pitfalls
Hotlinked images rot — expect 10-30% of URLs to fail at download; collect more than you need. Google results drift by market; fix country/language for reproducibility. Duplicate content across domains is common — hash downloaded files, not just URLs. The actor gives you the index; license review per source stays yours.
Related use cases
Frequently asked questions
Does the actor download image files?
It returns image URLs, dimensions, source pages, and domains as JSON. Downloading the files is a separate step — which keeps the run fast and lets you filter before you fetch.
How many images can I collect per query?
maxResults caps the records per query. Pass a list to queries for multiple classes in one run — each record carries its query so the labels come for free.
Can I filter by country or language of the search?
Yes. country and language map to Google's gl and hl parameters, which change which images surface for the same query in different markets.
Is scraping Google Images allowed for training data?
Image URLs point to files hosted by their owners — check each source's license before training. Collecting the URL index is the easy part; rights clearance is your responsibility.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.