Scraping srcset and lazy-loaded images correctly
srcset: pick the largest candidate
srcset lists candidates with a width (800w) or density (2x) descriptor. Scrapers that read src get the thumbnail. Parse every candidate and keep the largest.
Pitfall: do not split on every comma — CDNs such as Cloudinary put commas inside URLs (/w_800,h_400/). Split on a comma followed by whitespace, or on a comma right after a descriptor.
picture
<picture><source srcset> often holds WebP/AVIF versions larger than the fallback <img>. Read both.
Lazy loading
Lazy-load libraries leave a placeholder in src (often a data: GIF) and put the real URL in data-src, data-srcset, data-lazy-src, data-original or data-bg. Some also add the real image inside <noscript>. Placeholders can even appear inside srcset — skip candidates that start with data: or blob:.
Images injected purely by JavaScript after load (infinite scroll, some single-page apps) need a headless browser; for most server-rendered sites, reading these attributes is enough and much faster.
Code
See the parsers in Python and JavaScript, or use the hosted downloader, which handles all of the above.