Most log file analyses have the same four columns: bot, URL, datetime, and status code. It might also have page load time, size, and referrer at most.

But a request is more than a URL. When a bot asks for a page, it also sends request headers: metadata that tells the server what the client is, what it will accept, and sometimes what it already has in its cache.

You already know two of them. User-Agent is the one the entire bot analytics industry is built on. If-Modified-Since is how Googlebot says, “I saw this page on 25 August; only send it again if it changed.”

Almost nobody keeps the rest. Access logs throw them away by default, and as far as I know, no other log analysis product stores them. We started recording every header on both sides of every bot request in September 2026. We already have some data that’s the most interesting thing we have looked at in months.

What a request header actually is

Headers are how a client and a server exchange everything that is not the page itself.

Your browser sends them on every request: which languages you read, which image formats you can display, which page you came from. You can see yours right now by opening DevTools, going to the Network tab, and clicking the request for this page.

The server answers with its own set. The status code is the famous one, but a redirect carries Location, a cached page carries Cache-Control and Last-Modified, and cookies ride along in Set-Cookie.

What we store, and why it takes four copies

EdgeComet sits in the request path between your site and the bots that fetch it, which means we see the request twice and the response twice:

  1. What the bot sent us
  2. What we sent your origin
  3. What your origin answered
  4. What the bot finally received

It sounds like an implementation detail until the first time a customer asks why Googlebot got a 404 on a page that works fine in a browser. You can answer it by reading the exact request their server received instead of guessing.

User-Agent header

User-Agent is the only header present on 100% of all requests in our dataset. Every robots.txt rule, every log segment, every “AI crawler traffic is up 300%” chart in a client deck rests on it. It is also the only header in the request that the client writes about itself.

Here is what the major crawlers actually put in it:

TEXT
Googlebot Desktop
Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

Googlebot Smartphone
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.7977.82 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

AdsBot-Google
AdsBot-Google (+http://www.google.com/adsbot.html)

Storebot-Google
Mozilla/5.0 (X11; Linux x86_64; Storebot-Google/1.0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/147.0.0.0 Safari/537.36

bingbot
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/116.0.1938.76 Safari/537.36

GPTBot
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.4; +https://openai.com/gptbot)

ChatGPT-User
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot

ClaudeBot
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; [email protected])

PerplexityBot
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)

Applebot
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)

meta-webindexer
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36 (compatible; meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler))

LinkedInBot
LinkedInBot/1.0 (compatible; Mozilla/5.0; Apache-HttpClient +http://www.linkedin.com)

Read bingbot, GPTBot, ClaudeBot and PerplexityBot one after another. They are the same string with the name swapped:

TEXT
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; NAME/VERSION; +contact)

If-Modified-Since header

If-Modified-Since is a caching negotiation. The bot says, “I have a copy from this timestamp.” If nothing changed, the server answers 304 Not Modified with no body; everyone saves the bandwidth, and the bot spends its next request on a different page. Everyone in technical SEO has heard this is good for crawl budget.

Here is who actually sends it:

BotShare of requests with
If-Modified-Since
LinkupBot91.9%
SeznamBot86.4%
Barkrowler74.6%
Slackbot21.7%
Googlebot Smartphone15.1%
AhrefsBot15.0%
Googlebot Desktop14.6%
AdsBot-Google0.6%
bingbot0.4%
Baiduspider0.04%

Googlebot uses it on about one request in seven. The advice you have read for years treats this as Googlebot’s standard behavior. It is not; it is a minority of its fetches.

The entire AI industry sends it zero times. Not “rarely”. Zero, across GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot, ClaudeBot, Claude-User, PerplexityBot, meta-externalagent and meta-webindexer. Applebot, Amazonbot, and LinkedInBot send it zero times too. 

When people say AI crawlers are heavier on origins than Googlebot, this is one of the mechanisms. They have no way to be told: “nothing changed”. Every fetch is a full fetch.

Googlebot

The search crawler is close to the minimum viable HTTP request.
A complete Googlebot Desktop request, every header it sent:

HeaderValue
accept*/*
accept-encodinggzip, br
fromgooglebot(at)googlebot.com
user-agentMozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

And a complete Googlebot Smartphone request:

HeaderValue
accepttext/html,application/xhtml+xml,application/signed-exchange;v=b3,application/xml;q=0.9,*/*;q=0.8
accept-encodinggzip, br
amp-cache-transformgoogle;v="1..8"
fromgooglebot(at)googlebot.com
user-agentMozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.7977.82 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

The biggest crawler on the internet asks for your page with four headers and one extra line on mobile. Note application/signed-exchange;v=b3 sitting in that accept string: Googlebot is still telling every server on the web that it accepts signed exchanges.

The amp-cache-transform header

That extra line on mobile is amp-cache-transform: google;v="1..8", and it appears on 94.3% of Googlebot Smartphone requests. Googlebot Desktop never sends it.

AMP has been out of the SEO conversation for years. It is still in the mobile crawler’s default header set on nearly every request it makes, advertising which versions of the AMP cache transform Google can handle. Nobody is going to act on this, but it reminds us that crawler configuration outlives strategy by a long time.

Chrome versions in the Googlebot user agent

Googlebot is evergreen, and because the Chrome version sits inside the user agent string, you can check that claim directly. Among IP-verified Googlebot Smartphone requests: Chrome/152 on 94.8%, Chrome/153 on 2.7%. AdsBot-Google-Mobile was on Chrome/152 for 99.7% of its requests.

There is a small tail underneath: Chrome/124 on 1.0%, Chrome/151 on 0.8%, Chrome/99 on 0.7%, and seven requests still on Chrome/117. Not enough to change anything you do, but “Googlebot always runs the latest Chrome” is a rounding-up of the truth.

How old the If-Modified-Since timestamps are

Because If-Modified-Since carries a date, it is a direct readout of the bot’s memory of your page. We measured how old that timestamp was at request time:

Googlebot Desktop is re-checking things it fetched yesterday. Smartphone carries a much longer memory, out to 60 days. AdsBot is working from a copy about three weeks old. Three agents, one infrastructure, three completely different recrawl rhythms, and you can read them straight out of the header.

AdsBot-Google and where bot cookies come from

AdsBot-Google sends 7.11 headers per request, against Googlebot’s 3.14. 

We pulled the 90k AdsBot requests in the cohort apart by what they actually asked for. 99.9% of them are not pages at all. They are in-page API endpoints:

What AdsBot requestedRequests
/catalog_ajax_data/index/get39,167
/customer/section/load10,235
/livechat/getvisitor/8,738
/strix_header_banner/banner/getdata/8,396

Almost 100% of origin responses to those came back as application/json. Exactly 71 AdsBot requests in the entire cohort were HTML document URLs.

This is a complete AdsBot request for an actual page:

HeaderValue
accepttext/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8
accept-encodinggzip, br
fromgooglebot(at)googlebot.com
user-agentAdsBot-Google (+http://www.google.com/adsbot.html)

And this is a complete AdsBot request from the same crawl, moments later, for one of the AJAX endpoints that page called:

HeaderValue
accept*/*
accept-encodinggzip, br
accept-languageen-US
cache-controlmax-age=86400
cookie_lb=...; _biano=...; CookieConsent={...}; PHPSESSID=...; form_key=...
fromgooglebot(at)googlebot.com
refererhttps://example-shop.com/p/product-200x100-cm-alb
user-agentAdsBot-Google (+http://www.google.com/adsbot.html)
x-requested-withXMLHttpRequest

AdsBot’s document requests average 3.3 headers, which is Googlebot’s number, and 70% of those 71 carried nothing beyond accept, from, and user-agent.

The rich headers belong to the renderer, not to the crawler. AdsBot opens a product page in a headless browser; the page’s own JavaScript fires requests at the price, cart and live-chat endpoints, and those requests carry everything a browser attaches: referer on 100% of them, always the product page it is standing on, accept-language: en-US on 100%, cache-control: max-age=86400 on 100%, x-requested-with: XMLHttpRequest on 20.8%, and cookies on 90.1%.

Cookies surprised me most. The origin sets them: it returns set-cookie on 97.3% of those responses. They are consent banners and analytics state, CookieConsent, _ga, _gcl_au, form_key, mage-cache-storage. It loaded a page, your site handed it a consent cookie, and it echoed the cookie back on the next request, the way every browser does.

Storebot-Google

Storebot-Google sends 24 headers per request. Full Chrome client hints (sec-ch-ua, sec-ch-ua-platform, sec-ch-ua-mobile), the full sec-fetch-* set, accept-language, dnt: 1, cookies on 95.7% of requests and a referer on 95.7%.

It also cryptographically signs every single request. Here is one complete Storebot request, the whole thing:

HeaderValue
accept*/*
accept-encodinggzip, br
accept-languageen-US,en;q=0.9
dnt1
refererhttps://example-shop.com/p/product.hml
sec-ch-ua"Google Chrome";v="147", "Not.A/Brand";v="8", "Chromium";v="147"
sec-ch-ua-mobile?0
sec-ch-ua-platform"Linux"
sec-fetch-destempty
sec-fetch-modecors
sec-fetch-sitesame-origin
signatureg=:+5UdNVHs6h/EqgSd7RensWXZJUxSJUnEtr9Zy/VsR/gKrgMP6MrYxF18uXhj2CGHC/PC8HvM6SiaEsZLyk/eAA==:
signature-agentg="https://agent.bot.goog"
signature-inputg=("@target-uri" "@method" "signature-agent";key="g");created=1789380826;expires=1789381126;keyid="DYiMjA";alg="ed25519";tag="web-bot-auth"
user-agentMozilla/5.0 (X11; Linux x86_64; Storebot-Google/1.0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/147.0.0.0 Safari/537.36

Set that against the four-line Googlebot request above. Those three signature* lines are Web Bot Auth, a scheme that lets a bot prove who it is by signing each request with a private key instead of relying on a published IP range. Storebot uses it on 100% of its requests, with a 300-second validity window. No bot from OpenAI, Anthropic, Perplexity, or Meta sends it.

The takeaway for an SEO reading their logs: Storebot-Google is Google driving a browser through your checkout, not a crawler reading HTML. Treat it like a user, because everything it sends says it is one.

ChatGPT and GPTBot

OpenAI’s bots are the most talkative of the AI vendors, and they expose internal infrastructure in a way the others do not. A complete GPTBot request, fetching one of our own pages:

HeaderValue
accept*/*
accept-encodinggzip, br
fromgptbot(at)openai.com
refererhttps://edgecomet.com/
user-agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.4; +https://openai.com/gptbot)
x-openai-host-hash467038936

A complete ChatGPT-User request:

HeaderValue
accepttext/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9
accept-encodinggzip, br
accept-languageen-US,en;q=0.9
user-agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
x-envoy-expected-rq-timeout-ms15000

Two headers in there exist nowhere else on the web, and both are worth a section.

The x-openai-host-hash header

GPTBot sends x-openai-host-hash on 99.8% of requests. OAI-SearchBot sends it on 85.3%.

The value is a numeric token, and it is stable per domain. One dominant value covers 91% to 100% of all OpenAI requests to a given site. 

GPTBot and OAI-SearchBot send the same value for the same domain. OpenAI’s training crawler and its search crawler share one internal identifier for your site. They are two user agents on top of one host registry, the same way Googlebot and AdsBot are two configurations on one fetching service.

If you run several country domains, each one gets its own hash. They are treated as separate hosts internally.

OpenAI’s 15-second timeout header

ChatGPT-User sends x-envoy-expected-rq-timeout-ms: 15000 on 95.2% of requests. OAI-AdsBot sends it on 95.7%, OAI-SearchBot on 11.6%. The value was 15000 on every single occurrence in the dataset except one, which carried 13994.

Envoy is a service-mesh proxy, and this header is an internal timeout budget leaking out of OpenAI’s own network into yours. It is not a documented SLA, and we would not treat it as one. But it is the only number any AI vendor has ever effectively put in writing about how long they will wait for your page, and it appears on the bot that fetches pages live while a person waits for an answer.

Fifteen seconds is a lot for a fast site and not very much for a slow product page behind a cold cache.

GPTBot’s referer header

GPTBot sends a referer on 98.6% of its requests, which is more often than any other AI crawler we see. Two things about the value are odd.

It is a site root, not a page. An ordinary referer names the page that carried the link, a crawler walking your site would report the category page it came from. GPTBot reports https://example-shop.com/. Whatever this field holds, it is closer to a note about where the URL came from than a record of the previous hop.

And only 36.6% of them are same-site. The rest name a sibling country domain of the same brand: GPTBot fetches a page on your .de site and attaches the root of your .it site.

It is tempting to conclude OpenAI treats your ccTLDs as one property, and the host hash above says the opposite, since each country domain carries its own value. I believe: GPTBot discovers URLs on one of your country domains and fetches them on another, and the field it uses to tell you so points at a homepage rather than a page.

ChatGPT-User, by contrast, sends no referer at all, and a full browser-style accept string. It behaves like a browser because it effectively is one: a fetch made on behalf of a person who just asked a question.

Claude

ClaudeBot sends 2.11 headers per request. That is the lowest of any high-volume crawler in the dataset. Here is a complete ClaudeBot request. Nothing has been trimmed:

HeaderValue
accept*/*
accept-encodinggzip, br
user-agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; [email protected])

No from, no accept-language, no internal identifiers, no signature, no cookies. A referer appears on 7.5% of requests and is the same-site every time.

Who actually sends Claude-User requests

This is the most useful thing in our Anthropic data, and you cannot see it without the headers.

Claude-User is the agent that fetches a page because a person asked Claude about it, so it is the one Anthropic name with commercial consequences for you. In our window, roughly 6,200 requests arrived wearing it. Only 501 of them were Anthropic.

What it really wasRequestsDistinct UA stringsDistinct IPsIP verification
Anthropic’s fetcher50113passes
Claude Code on someone’s laptop71230202fails, legitimately

Anthropic’s own fetcher is the only one that verifies, and it is as plain as ClaudeBot. One user agent string, three IP addresses, two headers of its own:

HeaderValue
accept*/*
accept-encodinggzip, br
user-agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; [email protected])

The second group is Claude Code, Anthropic’s coding agent, running its WebFetch tool on a developer’s machine. It fails IP verification for an obvious reason: it is running from that person’s home or office network, not from Anthropic’s servers. The user agent carries the client version instead of a bot URL, which is why we counted 30 spellings of it:

HeaderValue
accepttext/markdown, text/html, */*
accept-encodinggzip, br
user-agentClaude-User (claude-code/2.1.270; +https://support.anthropic.com/)

It sent accept: */* on all requests where we recorded headers.
This is the only agent in our entire dataset that asks for markdown ahead of HTML. If you have been following the argument about whether to serve markdown to AI agents, note who is actually asking for it. It is not a crawler building an index. It is a developer tool fetching one page on demand..

Perplexity

PerplexityBot sends two headers of its own, and one of them is the user agent. The full request:

HeaderValue
accept-encodinggzip, br
from[email protected]
user-agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)

It does not send accept at all. Across 826 verified requests, not one carried it. Perplexity is the quietest identified crawler in the dataset.

Meta

Meta’s two agents, meta-externalagent and meta-webindexer, are the heaviest header senders in the dataset: 9.19 of their own headers, 18.3 in total. A complete meta-webindexer request:

HeaderValue
accept*/*
accept-encodinggzip, br
accept-languageen-US,en;q=0.9;q=0.9
cookie_biano=...; CookieConsent={...}; _ga=...; PHPSESSID=...; mage-cache-storage={}
refererhttps://example-shop.cz/p/zucchetti-helm-sprchova-baterie-pod-omitku-ano
sec-fetch-destempty
sec-fetch-modecors
sec-fetch-sitesame-origin
user-agentMozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36 (compatible; meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler))
x-requested-withXMLHttpRequest

It mostly does not load pages. On 87% of requests, Meta sends sec-fetch-mode: cors, sec-fetch-dest: empty, sec-fetch-site: same-origin. That is the signature of an XHR subresource fetch from inside a headless browser, not a document load. Actual navigations account for about 5%.

It refuses your cache. cache-control: no-cache plus pragma: no-cache appears on about 70% of Meta requests. Where AdsBot volunteers that a day-old copy is fine, Meta explicitly demands a fresh origin hit.

Other bots

LinkedIn reads the first 3 MB and stops. A complete LinkedInBot request:

HeaderValue
accept*/*
accept-encodinggzip, br
rangebytes=0-3145727
service-namebabylonia-nearline-ingestion
user-agentLinkedInBot/1.0 (compatible; Mozilla/5.0; Apache-HttpClient +http://www.linkedin.com)
x-li-calltree-request-idAAAAAAAAAAAAAAAAAAAAAA==

That range header appears on 76% of its requests. If your page HTML is bigger than 3 MB, LinkedIn’s preview generator never reaches the end of it. service-name is an internal LinkedIn service, babylonia-ingestion on 73% of requests and babylonia-nearline-ingestion on 27%, and the x-li-calltree-request-id is a base64-encoded string of zeros on every single request they send.

bingbot and Applebot are as plain as Googlebot. Two more complete requests:

HeaderValue
accept*/*
accept-encodinggzip, br
frombingbot(at)microsoft.com
user-agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/116.0.1938.76 Safari/537.36
HeaderValue
accepttext/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8
accept-encodinggzip, br
user-agentMozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)

What this means for you

Most of this is trivia, and good trivia is worth reading. But five things here are worth acting on.

Where to find this data

EdgeComet’s log analyzer stores request and response headers for every bot request across all four hops, available to all customers since September 2026. Our customers mostly use them for unglamorous work: proving what the origin actually received, tracing a redirect chain, showing a developer that the 404 was real.

We ended up with a research article instead, because once you store something nobody has stored before, you find out things nobody has published. The user agent is what a bot says about itself, and the industry has spent a decade building on that one line. Everything else in the request is what the bot does.