How to filter icons and tracking pixels when scraping images
A typical e-commerce page yields 100+ image candidates; a third are icons, logos, payment badges, sprites and invisible 1×1 tracking GIFs. HTML width/height attributes are unreliable (often missing or CSS-scaled), and URLs rarely say "icon".
Filter by real pixel size
Read the dimensions from the first bytes of each file — no full decode needed. In Python, Pillow opens lazily:
from io import BytesIO
from PIL import Image
import requests
def big_enough(url, min_w=100, min_h=100):
data = requests.get(url, timeout=30).content
try:
w, h = Image.open(BytesIO(data)).size # reads the header only
except Exception:
return False # not an image (soft 404 HTML) or corrupt
return w >= min_w and h >= min_h
In Node.js, image-size does the same from a Buffer.
Add URL rules
Exclude paths containing logo, icon, sprite, avatar, badge, pixel; or keep only the product CDN path (/products/, cdn.shopify.com/s/files).
Detect soft 404s and duplicates
Some servers answer image URLs with an HTML error page and status 200 — check the file signature (JPEG FF D8 FF, PNG 89 50 4E 47, GIF GIF8, WebP RIFF…WEBP). Hash the bytes (SHA-256) to drop the same file served under several URLs.
Hosted option
Website Image Downloader does all of this by default (minWidth/minHeight 100 px, signature check, SHA-256 dedupe) and reports how many images each filter removed:
curl -X POST "https://api.apify.com/v2/acts/sste~website-image-downloader/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" -H "Content-Type: application/json" \
-d '{"startUrls":[{"url":"https://www.example.com/"}],"minWidth":300,"minHeight":300,"urlMustNotContain":["logo","icon","sprite"]}'