What crawl budget actually is
Two limits govern how much of your store Google crawls: crawl capacity (how much your server tolerates — rarely the binding constraint on Shopify) and crawl demand (how much Google wants your pages — driven by popularity, freshness and signal quality). Small catalogues effectively have no budget problem. Stores with thousands of products, faceted collection URLs, tag pages and blog archives do: demand is finite, and every wasted fetch is one a product page did not get.
Where Shopify stores waste it
- Dead sitemap entries. Products deleted but still listed burn fetches on 404s. The sitemap mirrors the catalogue — prune the catalogue and it follows.
- Redirect chains. Renamed handles and migrated URLs each cost extra fetches per crawl. Point feeds and internal links at final URLs.
- Duplicate crawl paths. Tag pages, vendor pages and faceted filters multiply URL space without adding products. Noindex the facets; keep the products canonical.
- Bloated child sitemaps. Shopify splits at 5,000 URLs per child; giant stores should confirm every child is fetchable, since one dead child poisons the index.
The sitemap role
The sitemap is the budget aiming device: a clean, honest list of live product URLs tells the crawler exactly where the value is. A sitemap full of dead entries does the opposite — it spends budget to teach the crawler your signals are unreliable, after which it crawls you less eagerly. Honesty compounds in both directions.
Indexation audit
- Sitemap fetches clean from outside your network, children included
- Zero dead or redirecting URLs inside the sitemap
- Faceted, tag and vendor-duplicate pages noindexed or canonicalised
- New products appear in the sitemap within a day
- Search Console “discovered” counts track catalogue size, not zero
Deeper: sitemap pillar · soft 404s and redirects · Related guide: could not fetch
Shipwork fetches your sitemap from outside your network, follows child sitemaps, and samples the URLs inside — reporting the dead weight wasting your crawl budget. Free, no account, no signup — paste your store address.
Check my sitemapQuestions
- How do I know crawl budget is my problem?
- Hundreds of valid products, a clean sitemap, and “discovered but not indexed” stretching for months. Small catalogues almost never have a budget problem.
- Do dead sitemap URLs really matter?
- Yes — each one spends a fetch that a product page did not get, and teaches the crawler your signals are unreliable. Prune relentlessly.
- Should I noindex tag and filter pages?
- If they duplicate collection content without adding products, yes. Keep crawl demand concentrated on canonical product and collection URLs.
- How fast should new products get indexed?
- With a clean sitemap and healthy demand, days. With a rotting sitemap, months. The sitemap is the difference.
Keep reading