Technical Video

Crawl Budget Optimization [Video]

Taha YelkenciTaha Yelkenci
SEO Consultant
· updated: 2020-01-25 3 min read
Crawl Budget Optimization [Video]

Crawl budget is the number of requests Googlebot makes to your site in a given period — or, more bluntly, the number of pages it crawls and the time it spends on the site.

By optimising crawl budget we can help Googlebot discover our important content more easily, crawl more of it, and control the time it spends.

Google re-crawls sites at regular intervals.

Gary Illyes has explained that Google builds a list of the site's URLs before the crawl begins. The best outcome is Googlebot crawling every page on that list.

Googlebot works through the list from top to bottom. URLs that receive many internal links, along with sitemaps, help Google build that list in the first place. And when it starts crawling, negative conditions in the list — 40x and 50x errors, too many 30x redirects — cause Googlebot to lower the crawl budget, crawl fewer pages and leave the site.

You can find how many pages Googlebot crawls on your site each day, and the average time it spends per page, in Search Console.

Search Console crawl stats

Why does optimising crawl budget matter?

Googlebot crawls a certain number of pages on our site every day. Because it does not allocate us an unlimited budget, there are times when it cannot crawl everything. If the budget is limited, we need to make sure Googlebot crawls the pages that matter most.

Take an "About us" page: if we link to it from every page and list it near the top of the sitemap, Googlebot may crawl it daily. But is "About us" a page we target organic traffic with? Do we want Googlebot crawling it and spending even a small slice of the budget there? No.

When optimising crawl budget, the aim is to get Googlebot to prioritise the quality content pages we actually want traffic on.

How can we optimise crawl budget?

  1. Site speed
  2. Clearing out unnecessary pages and keeping them out of the index
  3. Server analysis

Site speed optimisation

I will not go into how to improve site speed here — there are plenty of tools and articles on PageSpeed already.

Instead let me share an experience. A client of mine had content pages that needed updating every single day, and not a small number of them: more than 500 a day. Imagine changing the title of 500+ pages daily. When yesterday's index shows in Google's results, the click-through rate is very low; when the index is refreshed daily, the CTR doubled. That mattered enormously to us.

So we asked how we could raise PageSpeed and get Googlebot to crawl more pages. We went through every externally called resource and every file on the server that loads with the page, and optimised them. That work reduced Googlebot's "time spent downloading a page".

In the image below you can see the "time spent downloading a page" line falling. Look at the "pages crawled per day" line and you see it rising in inverse proportion.

Comparison of crawl statistics

To sum the chart up: when the time Googlebot spends per page falls, the number of pages it crawls rises.

Get rid of the unnecessary pages on your site

E-commerce sites in particular carry plenty of pages that abuse crawl budget: pointless content pages, needless parameters, products that read as duplicates. Using noindex and canonical tags is usually the healthiest way to keep those to a minimum.

Once you have removed the unnecessary pages and want them out of Google's index, gather them all into a file and build a temporary sitemap. Submit it to Google and you will speed up their removal from the index. Alongside that you can block the unnecessary pages in robots.txt. As a final step you can delete the temporary sitemap from Search Console.

Check your redirects

While clearing out unnecessary pages, a lot of redirects get added. If your site has hundreds of thousands of products or sub-pages, keeping track of redirected URLs becomes hard. You end up with situations like this: page A redirects to page B; a month later B redirects to C. As the number of redirects grows, so do the chains.

Try to keep those chains to a minimum. When B is redirected to C, for example, redirect A straight to C instead of letting it pass through B.

Do not use canonical tags on 404 pages

We mostly see this at companies developing their own software in-house. It does not come up with WordPress or the large e-commerce platforms, but using a canonical tag on 404 pages can get a great many of those 404s indexed. If you use a canonical on your 404 pages and no noindex, you can end up with pointless indexed pages like the ones below.

Needlessly indexed pages

Analyse your crawl budget

You have optimised your crawl budget — so how do you measure it?

  • How often does Googlebot come to the page?
  • Which page does it crawl most?
  • How much data does it download while crawling page X?

To answer those questions you need to analyse your server's daily logs. You can build tables in Excel and interpret them, but my recommendation is to use a tool.

Splunk, Logentries or Stackify will all do the job.

Log analysis is a subject in its own right. There is a whole article in it — and if there is demand, I will write one.

Taha Yelkenci

Taha Yelkenci

SEO since 2010. Founder of rankZup. Got a question? Write to me →