Crawl Budget — Definition & Meaning
What is Crawl Budget?
Crawl budget is the number of URLs Googlebot will request on your site in a given period. It's a function of your site's host load capacity and Google's "crawl demand" (how important + fresh your content appears). Small sites rarely hit the limit; sites over ~10k URLs must manage it actively.
Key points
- Wasted budget: crawling parameter URLs, faceted navigation, session IDs, calendar rabbit-holes, and 404 chains.
- Reclaim budget with: robots.txt disallows on parameter patterns, noindex on thin variants, correct canonicals, and aggressive 301 consolidation.
- Freshness signals (real lastmod, frequent publishing, healthy backlinks) grow crawl demand — Google then allocates more budget.
- Watch "Crawl stats" in GSC — sudden drops signal a technical block; sudden spikes signal a duplicate-URL explosion.
- For under-10k-URL sites, focus on content quality first — budget is only a real constraint at ecommerce or news-site scale.
Example
An ecommerce site with 5,000 products but 500,000 faceted URLs (color × size × price combinations) burns 90% of Googlebot's time on duplicate junk — starving the real product pages of crawl attention.
Frequently asked questions
How do I check my crawl budget?
GSC → Settings → Crawl stats shows requests per day, response codes, and file types. Divide daily requests by total indexable URLs to get your effective crawl rate.
Does blocking a page in robots.txt save crawl budget?
Yes — Googlebot skips disallowed URLs entirely. This is the fastest way to reclaim budget from parameter and faceted URLs.
Do small sites need to worry about crawl budget?
Rarely. Under ~10k indexable URLs, Google will happily crawl the whole site every few days. Focus on content quality and internal linking instead.