Collect Product Images for E-Commerce Catalog Builds
Pull product image URLs, dimensions, and source pages from Google Images as JSON — enough to fill catalog gaps and audit supplier photography.

TL;DR — The Google Images Scraper returns product image URLs with dimensions, source pages, and domains as JSON. Query by SKU or product name, filter by size, and fill catalog image gaps with documented provenance.
Why catalog image gaps hurt conversion
A product without a photo is a product that doesn't sell — and supplier feeds routinely ship SKUs with missing or broken image URLs. Google's image index is the fallback pool: manufacturers, retailers, and reviewers have already photographed almost everything.
The job-to-be-done is a candidate list per SKU: image URL, dimensions for quality filtering, and the source page so someone can confirm the match and the rights.
How does this compare to the alternatives?
| Supplier chase | Photo shoot | Thirdwatch actor | |
|---|---|---|---|
| Cost | Free, slow | $$$ per SKU | Pay per result |
| Reliability | Inbox-dependent | Perfect control | Broad coverage |
| Setup time | Days per vendor | Weeks | Minutes |
| Maintenance | Per SKU | Per reshoot | Re-run queries |
Photo shoots are right for hero SKUs. For the long tail of a catalog, an indexed image with a source trail is proportionate.
How to collect product images in 4 steps
How do I query per SKU?
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
skus = ["AquaPure AP-200 filter", "Trailblazer TB-40 backpack"]
run = client.actor("thirdwatch/google-images-scraper").call(
run_input={"queries": skus, "maxResults": 30, "country": "us", "language": "en"}
)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())How do I keep only usable images?
import pandas as pd
df = pd.DataFrame(rows).drop_duplicates("imageUrl")
good = df[(df.width >= 600) & (df.height >= 600)]
print(good.groupby("query").size())Dimensions filter out thumbnails before you spend review time.
How do I shortlist per SKU?
picks = (good.sort_values("position")
.groupby("query").head(5)
[["query", "imageUrl", "sourcePageUrl", "domain"]])Top-position records from manufacturer or major-retailer domains are your first review queue — sourcePageUrl is where the reviewer confirms the match.
How do I record provenance?
Keep the full record set next to your catalog import: imageUrl, sourcePageUrl, domain, query. When a rights question lands later, the trail answers where every image was found.
Sample output
{"imageUrl": "https://mfg-cdn.example.com/ap200/front.jpg",
"width": 1500, "height": 1500,
"sourcePageUrl": "https://manufacturer.example.com/ap-200",
"domain": "manufacturer.example.com",
"query": "AquaPure AP-200 filter", "position": 1}Common pitfalls
A returned image may show a different variant than your SKU — always confirm on sourcePageUrl before importing. Hotlinked URLs rot, so download promptly and keep the source trail. The same product image replicates across retail domains — dedupe by file hash, not URL. The actor supplies candidates with provenance; rights clearance for reuse is your call.
Related use cases
Frequently asked questions
Can I use found images directly in my catalog?
Check the source's license first — the record gives you imageUrl, domain, and sourcePageUrl to locate the owner. Supplier and manufacturer images often carry reuse terms; random ones don't.
How do I match images to my SKUs?
Build queries from SKU names plus model numbers. Each returned record carries its query, so the join back to your catalog is automatic.
What resolution can I expect?
Records include width and height so you can filter to catalog-grade sizes before downloading anything.
Does it cover regional markets?
country and language inputs map to Google's gl/hl parameters — the same query surfaces different supplier imagery per market.
Related
100 free credits, no credit card.
About 30 real searches. Add the MCP to Claude or Cursor in two minutes.