Skip to content
Back-to-school offer · −50% for life on the platform, for any subscription before 30 September. See the offer
CiteMe Research

What is actually crawling you

385,057 AI-agent requests across twenty sites. Seven per cent of them are not AI agents at all, the biggest crawler has never read your robots.txt, and no efficiency metric ranks the engines the same way twice.

Maxime Mansiet, Mael Bourdin, Alex Teillet · 15 September 2026

7.0%
of "AI crawler" traffic hits credential paths
0
robots.txt fetches by Meta in 71,547 requests
3,228
human referrals attributed to an engine

Summary

We logged every AI-agent request to twenty sites for eight months and classified each one by what the agent was doing: bulk-crawling, fetching a page while answering somebody, or sending a person to the site. A measurable share of what every GEO tool reports as AI crawler traffic is not an AI crawler at all.

Inside the report

  1. 01A measurable share of AI crawler traffic is credential scanning
  2. 02Two independent signals agree, and one shows the regex misses most of it
  3. 03Meta has never read robots.txt. Anthropic reads it more than content
  4. 04Crawl volume, coverage and referrals are three different leaderboards
  5. 05What engines fetch when answering is not what they crawl
  6. 06Why per-engine citation rates reverse when you control the prompts

This is not a sample of the web. These are sites that bought a generative engine optimisation product, which skews toward organisations that already suspected AI visibility mattered to them. Aggregates only: no site is identified.385,057 requests, 20 sites, January to September 2026