Learn / sitemap

Crawl budget: why some products never get indexed

Google allots your store a crawl budget — limited time and requests — and spends it wherever your signals point. Point it at dead URLs, redirect chains and duplicate pages, and real products wait months for a first crawl. Here is how the budget works on Shopify and how to aim it at products.

What crawl budget actually is

Two limits govern how much of your store Google crawls: crawl capacity (how much your server tolerates — rarely the binding constraint on Shopify) and crawl demand (how much Google wants your pages — driven by popularity, freshness and signal quality). Small catalogues effectively have no budget problem. Stores with thousands of products, faceted collection URLs, tag pages and blog archives do: demand is finite, and every wasted fetch is one a product page did not get.

Where Shopify stores waste it

5kURLs per child sitemap
48htypical GSC cache lag
30scrawler timeout ceiling
  • Dead sitemap entries. Products deleted but still listed burn fetches on 404s. The sitemap mirrors the catalogue — prune the catalogue and it follows.
  • Redirect chains. Renamed handles and migrated URLs each cost extra fetches per crawl. Point feeds and internal links at final URLs.
  • Duplicate crawl paths. Tag pages, vendor pages and faceted filters multiply URL space without adding products. Noindex the facets; keep the products canonical.
  • Bloated child sitemaps. Shopify splits at 5,000 URLs per child; giant stores should confirm every child is fetchable, since one dead child poisons the index.

The sitemap role

The sitemap is the budget aiming device: a clean, honest list of live product URLs tells the crawler exactly where the value is. A sitemap full of dead entries does the opposite — it spends budget to teach the crawler your signals are unreliable, after which it crawls you less eagerly. Honesty compounds in both directions.

Indexation audit

  • Sitemap fetches clean from outside your network, children included
  • Zero dead or redirecting URLs inside the sitemap
  • Faceted, tag and vendor-duplicate pages noindexed or canonicalised
  • New products appear in the sitemap within a day
  • Search Console “discovered” counts track catalogue size, not zero

Deeper: sitemap pillar · soft 404s and redirects · Related guide: could not fetch

Find out if this is happening to you

Shipwork fetches your sitemap from outside your network, follows child sitemaps, and samples the URLs inside — reporting the dead weight wasting your crawl budget. Free, no account, no signup — paste your store address.

Check my sitemap

Questions

How do I know crawl budget is my problem?
Hundreds of valid products, a clean sitemap, and “discovered but not indexed” stretching for months. Small catalogues almost never have a budget problem.
Do dead sitemap URLs really matter?
Yes — each one spends a fetch that a product page did not get, and teaches the crawler your signals are unreliable. Prune relentlessly.
Should I noindex tag and filter pages?
If they duplicate collection content without adding products, yes. Keep crawl demand concentrated on canonical product and collection URLs.
How fast should new products get indexed?
With a clean sitemap and healthy demand, days. With a rotting sitemap, months. The sitemap is the difference.

Keep reading