What is actually crawling you
385,057 AI-agent requests across twenty sites. Seven per cent of them are not AI agents at all, the biggest crawler has never read your robots.txt, and no efficiency metric ranks the engines the same way twice.
Maxime Mansiet, Mael Bourdin, Alex Teillet · 15 September 2026
Summary
We logged every AI-agent request to twenty sites for eight months and classified each one by what the agent was doing: bulk-crawling, fetching a page while answering somebody, or sending a person to the site. A measurable share of what every GEO tool reports as AI crawler traffic is not an AI crawler at all.
Inside the report
- 01A measurable share of AI crawler traffic is credential scanning
- 02Two independent signals agree, and one shows the regex misses most of it
- 03Meta has never read robots.txt. Anthropic reads it more than content
- 04Crawl volume, coverage and referrals are three different leaderboards
- 05What engines fetch when answering is not what they crawl
- 06Why per-engine citation rates reverse when you control the prompts
This is not a sample of the web. These are sites that bought a generative engine optimisation product, which skews toward organisations that already suspected AI visibility mattered to them. Aggregates only: no site is identified.385,057 requests, 20 sites, January to September 2026