Empirical Study · MINDCEPTION Research
For 52 days we logged and verified every incoming request on an herbal e-commerce store (erbecedario.it). The goal: determine how AI models really crawl live web properties, detect crawler spoofing, and test if llms.txt is actually used.
Documented on Nginx and Cloudflare edge logs · Aug 1 – Sep 22, 2026 · n=1 sample
Behind Cloudflare or a reverse proxy, the client IP is the edge. We verified the unforgeable last hop of X-Forwarded-For via rDNS and ASN belonging to each organization.
Google's entire AI presence is its standard search crawl (193,264 content pages, ~3,200/day). AI Overviews and AI Mode use the same crawler. Anyone selling a "Google AI crawl strategy" distinct from technical SEO is selling an illusion.
Across 4.2 million requests, we observed 50 attempts to fetch /llms.txt (all 404). Not a single one came from OpenAI, Anthropic, Google, or Perplexity. Only SEO auditors like Semrush and BuiltWith check for it.
A single GCP IP address cycled through 10 distinct AI User-Agents in one day (ClaudeBot at breakfast, GPTBot at lunch, GrokBot in the afternoon) probing for /.env and /rclone.conf. Unvalidated User-Agent dashboards count attack scanners, not AI interest.
All of these crawlers systematically scanned the store catalog and product descriptions. Yet, in model answers to user queries, brand presence varied independently. To discover what generative engines actually say about your business, you need direct answer measurement.
We send you an entry measurement: where your brand appears across AI answers and which competitors take its place. Delivered: 5 prompts, 4 platforms, 3 brands. Calibration phase: limited capacity by application.
Thank you! We'll reply within 2 business days. ✅
No. In our verified logs (193,264 scanned pages), Google AI Overviews and Google AI Mode rely entirely on regular Googlebot crawls. There is no separate AI search crawler.
Across 4,259,100 requests over 52 days, zero consumer models (ChatGPT, Claude, Google, Perplexity) requested llms.txt. The only requests came from SEO and audit tooling.
No. User-Agent strings are easily forged. In our study, 99.7% of GrokBot traffic and over 97% of Google-Extended traffic was malicious scanners masquerading as AI bots. Verification requires rDNS and ASN checks on origin IPs.
No. Crawling is a baseline technical prerequisite, but appearance in generated answers depends on entity authority, third-party press, and corroborating external citations.