Website Image Downloader logoWebsite Image Downloader

Home › Guides

How to filter icons and tracking pixels when scraping images

A typical e-commerce page yields 100+ image candidates; a third are icons, logos, payment badges, sprites and invisible 1×1 tracking GIFs. HTML width/height attributes are unreliable (often missing or CSS-scaled), and URLs rarely say "icon".

Filter by real pixel size

Read the dimensions from the first bytes of each file — no full decode needed. In Python, Pillow opens lazily:

from io import BytesIO
from PIL import Image
import requests

def big_enough(url, min_w=100, min_h=100):
    data = requests.get(url, timeout=30).content
    try:
        w, h = Image.open(BytesIO(data)).size  # reads the header only
    except Exception:
        return False  # not an image (soft 404 HTML) or corrupt
    return w >= min_w and h >= min_h

In Node.js, image-size does the same from a Buffer.

Add URL rules

Exclude paths containing logo, icon, sprite, avatar, badge, pixel; or keep only the product CDN path (/products/, cdn.shopify.com/s/files).

Detect soft 404s and duplicates

Some servers answer image URLs with an HTML error page and status 200 — check the file signature (JPEG FF D8 FF, PNG 89 50 4E 47, GIF GIF8, WebP RIFF…WEBP). Hash the bytes (SHA-256) to drop the same file served under several URLs.

Hosted option

Website Image Downloader does all of this by default (minWidth/minHeight 100 px, signature check, SHA-256 dedupe) and reports how many images each filter removed:

curl -X POST "https://api.apify.com/v2/acts/sste~website-image-downloader/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" -H "Content-Type: application/json" \
  -d '{"startUrls":[{"url":"https://www.example.com/"}],"minWidth":300,"minHeight":300,"urlMustNotContain":["logo","icon","sprite"]}'

Related

Try it on Apify → API docs Use with AI agents (MCP)