Empirical Study · MINDCEPTION Research

What AI bots actually do on your website: 4,259,100 requests, 52 days

For 52 days we logged and verified every incoming request on an herbal e-commerce store (erbecedario.it). The goal: determine how AI models really crawl live web properties, detect crawler spoofing, and test if llms.txt is actually used.

Documented on Nginx and Cloudflare edge logs · Aug 1 – Sep 22, 2026 · n=1 sample

Who actually hits the origin servers

Behind Cloudflare or a reverse proxy, the client IP is the edge. We verified the unforgeable last hop of X-Forwarded-For via rDNS and ASN belonging to each organization.

Claimed Bot (User-Agent)
Verified Requests (rDNS/ASN IP)
Googlebot (Search + AI Overviews)
196,639 genuine requests (crawl-*.googlebot.com)
ClaudeBot (Anthropic)
14,726 genuine requests (AWS ASN, tight IP set)
Amazonbot (Amazon)
12,443 genuine requests (*.amazonbot.amazon)
ChatGPT-User (OpenAI, on-demand)
7,454 genuine requests (Azure/GCP)
PerplexityBot (Perplexity)
3,131 genuine requests (16 IPs, AWS)
GPTBot + OAI-SearchBot (OpenAI)
5,066 genuine requests (2,758 + 2,308)
Google-Extended (Gemini opt-out)
801 total requests — only 23 genuine (97.1% spoofed)
GrokBot (xAI)
761 total requests — only 2 genuine (99.7% spoofed)

Three core findings from the data

1

Google AI = Googlebot

Google's entire AI presence is its standard search crawl (193,264 content pages, ~3,200/day). AI Overviews and AI Mode use the same crawler. Anyone selling a "Google AI crawl strategy" distinct from technical SEO is selling an illusion.

0

llms.txt fetched by models

Across 4.2 million requests, we observed 50 attempts to fetch /llms.txt (all 404). Not a single one came from OpenAI, Anthropic, Google, or Perplexity. Only SEO auditors like Semrush and BuiltWith check for it.

99%

User-Agent spoofing

A single GCP IP address cycled through 10 distinct AI User-Agents in one day (ClaudeBot at breakfast, GPTBot at lunch, GrokBot in the afternoon) probing for /.env and /rclone.conf. Unvalidated User-Agent dashboards count attack scanners, not AI interest.

The golden rule: Being read ≠ Being cited

Crawler Scans (Server Logs)
Generative Visibility (MINDCEPTION)
Measures raw HTML and assets fetched
Measures if your brand is recommended in generated answers
Analyzes IP, rDNS, and HTTP status codes (200, 404, 503)
Analyzes semantic entity rank, mention share, and co-citations
Infrastructure prerequisite
Market perception outcome for the customer

All of these crawlers systematically scanned the store catalog and product descriptions. Yet, in model answers to user queries, brand presence varied independently. To discover what generative engines actually say about your business, you need direct answer measurement.

Measure your brand's AI visibility

We send you an entry measurement: where your brand appears across AI answers and which competitors take its place. Delivered: 5 prompts, 4 platforms, 3 brands. Calibration phase: limited capacity by application.

Delivered within 4 working days · No commitment

Frequently asked questions about AI crawlers

Does Google AI have a dedicated crawler distinct from Googlebot?

No. In our verified logs (193,264 scanned pages), Google AI Overviews and Google AI Mode rely entirely on regular Googlebot crawls. There is no separate AI search crawler.

Do AI models fetch and use the llms.txt file?

Across 4,259,100 requests over 52 days, zero consumer models (ChatGPT, Claude, Google, Perplexity) requested llms.txt. The only requests came from SEO and audit tooling.

Is checking the User-Agent enough to identify an AI crawler?

No. User-Agent strings are easily forged. In our study, 99.7% of GrokBot traffic and over 97% of Google-Extended traffic was malicious scanners masquerading as AI bots. Verification requires rDNS and ASN checks on origin IPs.

Does being crawled by an AI bot ensure you get cited in answers?

No. Crawling is a baseline technical prerequisite, but appearance in generated answers depends on entity authority, third-party press, and corroborating external citations.