How to Do Log File Analysis for SEO
SEO & GEO Consultant | | 11 min read

Log file analysis for SEO is a technical SEO task that reads a server's access logs to measure how Googlebot and AI bots actually crawl a site. Each line records which bot requested which URL, when, and which status code the server returned, so the logs show where crawl budget is spent.
In my crawl budget optimisation post I wrote that log analysis deserved an article of its own, and this is the article promised there.
What is an access log and which fields does one line contain?
An access log is a text file in which the web server writes one line for every request it receives. The common format on Apache and Nginx is called "combined". The line below is an example for the site example.com.
66.249.66.1 - - [05/Oct/2026:09:14:32 +0300] "GET /blog/sample-post/ HTTP/1.1" 200 18452 "-" "Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
A combined line carries these seven pieces of information, left to right.
- IP address: the client that made the request (66.249.66.1).
- Identity and user fields: written as a hyphen when there is no authentication.
- Date and time: when the request was received, with the time zone.
- Request line: method, requested URL and protocol (GET /blog/sample-post/ HTTP/1.1).
- Status code: the server's response (200).
- Size: the response size in bytes (18452).
- Referer and user agent: the last two fields, each in quotes.
Chrome/W.X.Y.Z is the placeholder Google's documentation uses for the Chrome version. A real log line shows a version number that changes over time.
Why is log file analysis necessary?
Log file analysis is necessary because only the access logs keep every request that reaches the server, with the bot's name, the URL and the time. The Crawl Stats report in Search Console covers Google's crawlers only and lists URLs as examples. Google notes that the report might not count some requests, so its numbers can differ slightly from your logs. The table below compares the two sources on seven points.
| Point | Access logs | Crawl Stats report |
|---|---|---|
| Clients covered | Every client: search engine bots, AI bots, browsers | Google's crawlers only |
| URL detail | The full URL of every request | Example URLs for each group, not a full list |
| Bot identity | The user agent string, verified by you | Google's own data, grouped by Googlebot type |
| Crawl purpose | Not recorded | Split into discovery or refresh |
| Requests that never reach the server | Not visible | DNS errors plus unreachable pages appear in the response table |
| Period | As long as the server keeps logs | Host status is assessed over the last 90 days |
| Access | Server or hosting panel permissions | A root-level property in Search Console |
Google's help page aims the Crawl Stats report at advanced users. The same page says a site with fewer than a thousand pages should not need this level of detail.
Where do you get access logs?
You get access logs from three places: the hosting panel, the log files on the server and the logs of your content delivery network (CDN).
- Hosting panel: the Raw Access interface in cPanel downloads each domain's log as a .gz file. With the archive option on, older logs collect in the logs folder of your home directory.
- Apache or Nginx files: the file path is set by the CustomLog directive on Apache or access_log on Nginx.
- CDN logs: a CDN answers cached requests itself, so those requests never reach the origin server's log. Cloudflare exports its own logs through Logpush.
A server behind a CDN logs the CDN's IP address by default, not the visitor's. Cloudflare sends the original address in the CF-Connecting-IP header, so log that header before you verify Googlebot.
How do you verify Googlebot in the logs?
You verify Googlebot by running a reverse DNS lookup on the logged IP address or by matching the address against the IP ranges Google publishes. A user agent string can be spoofed, so a "Googlebot" line does not always come from Google. Google's guide to verifying crawler requests describes the manual check in the three steps below.
- Run a reverse DNS lookup on the IP address from your logs with the
hostcommand. - Check that the returned domain name is googlebot.com, google.com or googleusercontent.com.
- Run the
hostcommand again on the returned domain name. The resulting IP address must match the one in your logs.
host 66.249.66.1
host crawl-66-249-66-1.googlebot.com
For large files, match the IP addresses against the ranges Google publishes as JSON files. The common-crawlers.json file lists the ranges for Googlebot in CIDR format.
How do you run log file analysis from the command line?
You run log file analysis from the command line in six steps with grep, awk, sort and uniq. The examples assume a combined-format file named access.log: field one is the IP address, field seven the URL and field nine the status code.
- Count the user agents. The output ranks the clients by request count.
awk -F'"' '{print $6}' access.log | sort | uniq -c | sort -rn | head -20 - Count the requests by bot name.
grep -oE 'Googlebot|GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|Claude-SearchBot|PerplexityBot|Perplexity-User' access.log | sort | uniq -c | sort -rn - Copy the Googlebot lines into a separate file. The second command prints the reverse DNS name of each IP address. Single out any address that does not end in one of Google's three domains.
grep 'Googlebot' access.log > googlebot.log awk '{print $1}' googlebot.log | sort -u | while read ip; do echo "$ip $(host "$ip" | awk '{print $NF}')"; done - Get the distribution of status codes.
awk '{print $9}' googlebot.log | sort | uniq -c | sort -rn - List the most crawled URLs.
awk '{print $7}' googlebot.log | sort | uniq -c | sort -rn | head -20 - Isolate the parameterised URLs, meaning those with a question mark. The second command prints their count next to the total Googlebot requests.
awk '$7 ~ /\?/ {print $7}' googlebot.log | sort | uniq -c | sort -rn | head -20 awk '$7 ~ /\?/ {p++} END {print p+0, "/", NR}' googlebot.log
To read a compressed log without unpacking it, pipe the file in with zcat.
zcat access.log.gz | grep 'Googlebot' > googlebot.log
What should you look for in log file analysis?
Log file analysis should look at six things: parameterised URLs, redirect chains, error codes, uncrawled important pages, crawl frequency by section and response time. The first three show wasted crawling.
- Parameters: parameterised URLs, such as differently sorted versions of one page, make bots crawl duplicate content. Google recommends consolidating these URLs or blocking them in robots.txt.
- Redirect chains: list the URLs that return 301 or 302. A target that redirects again forms a chain. Google states that long redirect chains affect crawling negatively.
- 404 and 5xx: Google lowers its crawl capacity limit when a site responds with 5xx errors. A 404 is the correct response for a permanently removed page.
- Uncrawled important pages: write your important URLs into a file and find those missing from the logs with the command below.
- Frequency by section: the spread of requests across directories such as /blog/ or /products/ shows which part of the site attracts crawling.
- Response time: the combined format does not include it. Add %D on Apache or $request_time on Nginx to the log format.
awk '{sub(/\?.*/, "", $7); print $7}' googlebot.log | sort -u > crawled.txt
grep -vxFf crawled.txt important.txt
awk '{sub(/\?.*/, "", $7); split($7, a, "/"); print "/" a[2]}' googlebot.log | sort | uniq -c | sort -rn | head -20
The first two commands print the paths in important.txt (one path such as /blog/sample-post/ per line) that Googlebot never requested. The third counts requests by first directory. Check whether an uncrawled page receives internal links: I covered mapping internal links with Gephi in a separate post. My post on data analysis in SEO shows how to join crawl data with traffic data in one table.
How do you recognise AI bots in access logs?
You recognise AI bots in access logs by the names in their user agent strings. The names come from the companies' own documentation: OpenAI's crawler overview, Anthropic's help article and Perplexity's crawler page. The table below lists the eight names in those documents with the purpose each company states.
| Company | Name in the logs | Purpose stated by the company |
|---|---|---|
| OpenAI | GPTBot | Crawls content that may be used to train generative AI models |
| OpenAI | OAI-SearchBot | Surfaces websites in ChatGPT's search features |
| OpenAI | ChatGPT-User | Visits a page for a user action in ChatGPT, does not crawl automatically |
| Anthropic | ClaudeBot | Collects web content that could contribute to model training |
| Anthropic | Claude-User | Accesses a site when a user asks Claude a question |
| Anthropic | Claude-SearchBot | Navigates the web to improve search result quality |
| Perplexity | PerplexityBot | Surfaces websites in Perplexity search results, not used for model training |
| Perplexity | Perplexity-User | Visits a page in response to a user's question |
Google-Extended does not appear in logs. Google states that it has no separate user agent string and works as a robots.txt token. An AI bot's user agent can be spoofed too, and all three companies publish their bots' IP addresses as JSON files. OpenAI and Perplexity say robots.txt rules may not apply to requests a user starts.
Logs show that a bot requested a page, not whether your brand appears in AI answers. I explained that measurement in my post on measuring AI search visibility.
Which tools can you use for log file analysis?
Three tool options cover log file analysis: the command line, a spreadsheet and the Screaming Frog Log File Analyser.
- Command line: needs no extra software and processes the file line by line, as the commands above do.
- Spreadsheet: import the filtered Googlebot lines into Excel or Google Sheets and group them with a pivot table. Filtering on the command line first keeps the sheet small.
- Screaming Frog Log File Analyser: a desktop program for Windows, macOS and Linux. It supports the Apache and W3C Extended log formats.
According to the official Log File Analyser page, the program verifies search engine bots automatically and shows the most and least crawled URLs. It matches an imported URL list against the logs to find uncrawled and orphan pages. The free version is limited to 1,000 log lines and one project.
What are the common mistakes in log file analysis?
Six mistakes are common in log file analysis.
- Trusting the user agent without verifying the IP address. Fake Googlebot requests inflate the crawl count.
- Reading only the origin server's log on a site behind a CDN. Requests served from cache are missing from that file.
- Searching the logs for Google-Extended. Google-Extended is a robots.txt token, not a user agent.
- Searching for a user agent with its Chrome version number. Google says the version number increases over time and recommends wildcards.
- Counting every 404 line as an error. Google notes that 404 is correct for a page that is gone without a replacement.
- Counting image, CSS and JavaScript requests together with page crawls. Googlebot fetches these resources to render the page, so split the URL list by file extension.
How do you protect personal data when sharing access logs?
You protect personal data when sharing access logs by masking the IP addresses or by sending the bot lines only. An IP address is personal data: the European Commission lists it among its examples of personal data under the General Data Protection Regulation (GDPR). Turkey's Personal Data Protection Authority names IP masking in its cookie guideline as a technique that lowers identification risk.
Send an agency or consultant the filtered googlebot.log file, not the raw log. If they need the whole log, the example command below sets the last part of each IPv4 address to zero. A reverse DNS lookup is unreliable on a masked address, so verify Googlebot before you mask.
awk '{sub(/\.[0-9]+$/, ".0", $1); print}' access.log > access-masked.log
How do you measure the result of log file analysis?
You measure the result of log file analysis by running the same commands on new logs after the fix and comparing the numbers of the two periods. Pick two periods of equal length and note these four values.
- The share of parameterised URLs in Googlebot requests.
- The number of 301, 404 and 5xx responses.
- The number of uncrawled pages on the important list.
- The number of Googlebot requests per day.
awk '{print substr($4, 2, 11)}' googlebot.log | uniq -c
The command above counts Googlebot requests by day. Put the daily counts next to the total crawl requests chart in the Crawl Stats report: when both move in the same direction, the measurement is consistent.
Taha Yelkenci
SEO since 2010. Founder of rankZup. Got a question? Write to me →
← Previous post
Next post →
You may also like
January 24, 2020 · 3 min
August 10, 2020 · 3 min
September 1, 2019 · 3 min