Reading AI Crawler Logs: What GPTBot Actually Requests
AI crawlers leave very readable patterns in your server logs. How to segment the bots, what each one's fetch pattern means, and the two questions worth answering before you touch a GEO tool.

Most GEO advice starts with content. We think it should start with logs, because your server already keeps the cheapest dataset in AI visibility: a timestamped record of which assistant bots fetched which pages, how often, and in what order. Reading it answers two questions no dashboard can — is anyone's AI actually reading us, and what do they think is worth fetching.
This is the hands-on follow-up to our AI crawlers guide. Same cast of characters; this time we look at what they do after the handshake.
First: segment the bots, or the data is noise
The common mistake is treating "AI traffic" as one bucket. The bots behave nothing alike, and mixing them produces averages that mean nothing.
- Training crawlers —
GPTBot,ClaudeBot,anthropic-ai,Google-Extended,CCBot— crawl steadily and broadly. Their fetches feed future model weights, not today's answers. - Search crawlers —
OAI-SearchBot,PerplexityBot— build the index that search-backed answers cite from. These are the ones whose attention you want on new pages this week. - User-triggered fetchers —
ChatGPT-User,Claude-User,Perplexity-User— fetch a page only when a specific user's conversation needs it. Every hit here is a receipt: a real person asked about you right now.
The commands
On any access log (nginx/Caddy format works), a first pass:
# All AI crawler hits in the current log, newest last
grep -E "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|anthropic-ai|PerplexityBot|Perplexity-User|Google-Extended" access.log \
| awk '{print $NF}' | sort | uniq -c | sort -rn# All AI crawler hits in the current log, newest last
grep -E "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|anthropic-ai|PerplexityBot|Perplexity-User|Google-Extended" access.log \
| awk '{print $NF}' | sort | uniq -c | sort -rnThat gives you the raw volume ranking. The interesting cut is bot × path:
# Which pages each bot fetches
grep "GPTBot" access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -20# Which pages each bot fetches
grep "GPTBot" access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -20Two patterns to look for. A training crawler that fetches /robots.txt and then a broad, even spread of pages is indexing. A search crawler that fetches /, your sitemap, then a narrow set of specific pages is following relevance — and the pages it ignores are the ones it does not consider citable.
Three patterns worth diagnosing
1. The search bot never shows up. If OAI-SearchBot or PerplexityBot have zero hits over a month, the usual causes are a robots.txt allowlist copied from an old template (it disallows the generic pattern they fall under), or a site that rarely changes — search indexes skip static sites. The fix is the robots.txt and llms.txt work, plus a visible update cadence.
2. User-triggered hits spike on days you did nothing. ChatGPT-User and Perplexity-User only fetch when a live user's conversation needs the page. A burst with no publishing on your side usually means someone shared or discussed your product — community traffic the analytics dashboard will not label as such. Worth grepping your referrers the same day.
3. The training crawler fetches your private pages. If GPTBot is hitting /dashboard or /settings, your robots.txt does not say what you think it says — a typo'd User-agent line silently disables the whole group. Our robots.txt template keeps private paths blocked while leaving the citation-relevant pages open.
Two honest cautions
User-agent strings are claims, not identities. Anyone can curl with User-agent: GPTBot. For counting purposes this is fine; for "did a competitor scrape my whole site" paranoia, verify ownership the way the major providers document (reverse-DNS for the ones that support it) before drawing conclusions.
Volume is not a KPI. A training crawler fetching 4,000 pages tells you that you are in a future model's diet. The number that pays rent is the user-triggered and search fetch count on the pages that convert — that is the part of the pipeline where a fetch today can become a visitor tomorrow.
The two questions
Run the segmentation once, answer these, and you will know more than most sites that have bought a GEO subscription:
- Are the search-backed bots fetching my new pages within days of publishing? If not, the index that feeds citable answers does not know you exist.
- Which pages do the search bots never fetch? Those pages are invisible to the assistants regardless of how good their content is.
Both fixes are unglamorous — robots.txt corrections, internal links, an update cadence — which is exactly why the logs are worth reading before anything else.