Crawl budget is the amount of time and resources Google spends crawling your website, and it is one of those topics that sound so technical that nobody in the meeting wants to vote against it. Someone says "we need to improve our crawl budget", everyone nods and, suddenly, there is a three-month SEO project underway.
Before opening it, I would check something simpler: whether the search engine is arriving late to the pages that could generate business. If the answer is no, your crawl budget is probably fine and the traffic problem lives somewhere else.
In this guide I explain what crawl budget is, how to tell whether your site really has a problem, what wastes it, how to measure it with Search Console and logs, and which changes actually move it.
What is crawl budget?
Crawl budget is the set of URLs Google can and wants to crawl on your site during a given period. Its guide to crawl budget defines it as the time and resources it allocates to each host, and each hostname counts as a separate site (www.yoursite.com and blog.yoursite.com get separate budgets).
There is no dashboard where you can look it up, because it comes from two forces crossing each other: your server's capacity to handle visits and the demand your content generates for the search engine. You will also see these two pieces called crawl rate limit and crawl demand.
Crawl capacity limit
The crawl capacity limit is the ceiling Google sets so it doesn't knock your server over. It controls how many connections the crawlers open at the same time and how long they last. Every site starts with the same, fairly conservative limit, and it goes up if the server responds quickly and consistently.
If latency rises or 5xx or 429 errors show up, crawl capacity drops and so do the bot's visits. What matters here is server response time (how long it takes to deliver the HTML), more than the page load speed the user sees.
Crawl demand
Crawl demand is how much Google wants to crawl you, and it depends on the popularity of your URLs, how often they change and the inventory of your site it already knows. According to the official guide, that list of known URLs is the factor you control the most: if the search engine knows thousands of duplicate or useless addresses, it will waste time on them.
Big events, like a site migration, also trigger more demand because everything has to be reprocessed under the new addresses (something to keep in mind if you are sizing up the risk of a migration).
Do I have a crawl budget problem?
Google's own guide says who this topic is for, and the list is short:
| Type of site | When it applies according to the guide |
|---|---|
| Large sites | More than 1 million unique pages that change about once a week |
| Medium or large sites | More than 10,000 unique pages that change daily |
| Any site | A large share of its URLs shows up in Search Console as "Discovered - currently not indexed" |
If your site doesn't match any row, the official advice is to keep your sitemap up to date and check the Page indexing report. My take goes the same way: if you have 300 pages and new ones get crawled without delay, I wouldn't spend three months on this.
Watch out for the third row, because that is where everything gets mixed up. Here is a case of our own. Of the 32 Spanish-language articles we published on this blog between September 11 and October 2, 2026, by October 5 15 were indexed, Google didn't even know 10 of them ("URL is unknown to Google") and it had seen 7 without indexing them.
It looks like a textbook crawl budget problem. But this site has a few hundred URLs, and the search engine has plenty of capacity to go through all of them.
What was happening was far less glamorous: several new articles hung only from the blog listing, with no links from pages whose content already gets visits. That is called a discovery problem, and you work on it with internal links.
My rule of thumb: I talk about crawl budget when the site is large and, on top of that, there are pages that sell waiting for their turn. If only one of those two conditions holds, the SEO problem is somewhere else.
What eats up your crawl budget?
On large sites, the issue is rarely that Google crawls too little. What happens is that it crawls a lot of content nobody cares about, starting with your sales team.
Picture an ecommerce store with 200,000 products and filters for color, size, brand and price. Suddenly, the same black T-shirt in size M has 80 different URLs.
The bot can spend the day going through "black-tshirt-size-m-sorted-by-price" while new products, the ones that could actually sell, wait (it's an example, but I'd bet something similar exists in your catalog).
What eats up the budget the most, in the order I usually find it:
- Filters and parameters. Every filter combination creates a new address with almost the same content. Five filters with four values each already produce thousands of combinations per category.
- Pagination and internal search. Pages 40, 41 and 42 of a listing nobody reaches, or internal search results that ended up linked.
- Soft 404s. Empty listings that say "no results" but respond with a 200 status code. Since the code says everything is fine, the bot keeps visiting them.
- Redirect chains. Every hop is one more request, and after a few migrations they pile up on their own.
- Duplicate content. Versions with and without a trailing slash, with uppercase letters or with campaign parameters. Every copy is one more address to crawl.
None of these things was decided by someone in an SEO meeting one day. They just happened, template by template, year after year.
How do you measure crawl budget?
Heads up, here comes the dry part (I promise it's short): there is no official crawl budget metric, so what you measure is how the search engine behaves on your site. I do it in four steps.
-
Crawl Stats in Search Console. They are under Settings and Google keeps them for the last 90 days. Look at total requests, average response time and the breakdown by response code and by page type. Lots of 3xx usually points to chained redirects; lots of 404s, to broken links.
-
30 days of logs. Server logs show what Googlebot really crawled, with no interpretation. Filter real bot visits (verifying the IP belongs to Google) and group them by template: product pages, categories, filters, parameters, pagination and internal search.
-
The share going to what you don't want indexed. Let's say 1,000,000 bot visits a month. Sounds great.
Until you find that 38% went to filters and parameters you don't even want on Google. That is when the conversation shifts to what you are letting the bot crawl.
-
How long it takes to reach new things that sell. For every new product or landing page, compare the date you published the content with Google's first visit in the logs. If a new product page takes weeks while a filter gets visited daily, there is your crawl budget problem, measured in something the business cares about.
That fourth step is the one that helps me most to decide whether crawl budget deserves a project. Crawl Stats tell you how much gets crawled and how often; time to first visit tells you whether the search engine is arriving late to what matters.
How do you optimize crawl budget?
If the numbers say there is a crawl budget problem, this is the order in which I would work on it with the development team.
- Clean up the inventory. Compare three lists: what your sitemap says, what a crawler finds by following links and what the bot visits according to the logs. The differences are the diagnosis.
- Block, consolidate or remove, depending on the case. Each tool does something different, and using the wrong one is one of the most expensive mistakes on this topic:
| I want the URL to... | Use | What happens to crawling |
|---|---|---|
| Not be crawled | robots.txt | The bot stops requesting it. Careful: if it gets links, it can still show up without content |
| Not appear in search | noindex | The bot keeps requesting it to read the tag, so it doesn't save budget |
| Be consolidated with another | canonical | It's a signal to group duplicate content, and the search engine can ignore it |
| Disappear for good | 404 or 410 | It's clear there is no need to come back |
- Fix soft 404s and flatten redirect chains. Every rule should point straight to the final destination.
- Keep sitemaps honest. Segment them by page type and use lastmod only when the content really changed. If your CMS stamps today's date on 40,000 pages every night, the search engine stops trusting your lastmod (more on this in why Google isn't using your sitemap).
- Take care of the server. Fast, stable responses raise the capacity limit. Returning 304 (Not Modified) when a page hasn't changed lets the bot reuse its cached copy.
- Bring what sells closer to the homepage. Crawlers reach things a few clicks away sooner and more often. When we crawl an enterprise website we measure average click depth by template, and our rule is that everything commercial sits 3 clicks or fewer from the homepage. On a site with 200,000 URLs this is done by template: one module of 60 topic hubs with 40 links each rearranges 2,400 pages with a single change (the full logic is in silo structure).
If cleaning up your URLs turns up thousands of pages with no useful content, content pruning is the next step.
What doesn't improve crawl budget?
Some ideas about crawl budget circulate a lot in technical SEO and are worth discarding:
- Adding noindex to save crawling. The bot still has to request the page to see the tag. If you don't want it visited, the way to go is robots.txt.
- Resubmitting the sitemap every day. It helps discovery, but it doesn't give your site more budget.
- Setting crawl-delay in robots.txt. Googlebot ignores that directive.
- Celebrating that total crawling went up. The bot going through more filters improves nothing. What you want to see is less crawling of useless URLs and more of what sells.
- Expecting more authority to fix everything. Popularity raises demand, but if your site is full of duplicates, that extra demand gets spent on duplicates.
Do AI bots use the same crawl budget?
Partly. Google clarifies that the capacity limit is shared across all its crawlers: Googlebot, AdsBot and the Google Shopping crawler draw from the same pool. Google-Extended, which often comes up in this conversation, is a robots.txt control to decide whether your content trains Gemini, without being a separate bot.
Other companies' bots, like OpenAI's GPTBot (which collects content to train its models), ClaudeBot or PerplexityBot, don't spend Google's budget, but they do consume your server. And if the server slows down, crawl capacity drops for everyone.
That's why it pays to separate them by user agent in the same logs and see how much they weigh. My default is to let AI search bots through (like OAI-SearchBot or PerplexityBot), because blocking them also means they can't cite you; training is a separate decision.
Frequently asked questions
Does crawl budget affect SEO rankings?
Not directly: Google doesn't use it as a factor to rank results. The effect is indirect, because what doesn't get crawled doesn't make it into indexing, and a page outside the index can't rank.
Does a sitemap increase crawl budget?
No. A good sitemap helps new pages get found faster, but it doesn't make your site receive more bot visits.
What I would check this week
Before asking for budget for a crawl budget project, I would do three things that cost nothing. Check in Search Console whether you have many URLs under "Discovered - currently not indexed". Pick 10 new pages that should sell and see how long Google took to visit them.
And if you have logs, calculate what share of visits goes to pages you don't want on Google. If what sells gets crawled the same day and the waste is small, save that budget for a problem you actually have (and if not, a full SEO audit will sort out the rest).