AI Crawler Checker
Paste your domain and see, crawler by crawler, which ones can read your site and which are shut out. Eighteen AI bots checked, with the real effect of each block.
The costliest block is the one you do not know about
A robots.txt inherited from another era, a rule copied from a forum, a hosting default: AI crawler blocks are rarely chosen, they are inherited. The site keeps working, Google traffic does not move, and vanishing from assistant answers triggers no alert anywhere.
That is what this tool is for. It does not only tell you what is blocked, it tells you what the block costs: a shut-out training crawler changes nothing about today’s answers, a shut-out search crawler removes you from them entirely.
The three families of AI crawlers
AI bots do not all do the same thing, and conflating them leads to the wrong decisions. Training crawlers collect text for future models. Search crawlers build the knowledge base queried when an answer is generated. Live crawlers fetch a specific page because a user just asked a question.
OpenAI, Anthropic and Perplexity each operate several agents across different families, configurable separately. You can therefore refuse training while staying fully visible in answers. That is the trade-off most businesses want, and the one a blanket block destroys.
Once access is settled, measure what it produces with the citability checker.
Why this matters for GEO
Access comes before everything else: No editorial work compensates for a blocked crawler. It is the one check that invalidates all the others when it fails.
The cost of a block depends on the crawler: Refusing training is a licensing choice. Refusing the search index is a withdrawal from answers.
Defaults are rarely yours: Many blocks come from a hosting template or a copied rule, never from an explicit decision.
One check is not enough: A robots.txt drifts across migrations and deployments. This control is meant to be repeated.
Allowing crawlers is the starting point. Citeme then checks whether the engines actually cite you, across ten engines, every week.
Frequently asked questions
An AI crawler is a bot operated by an AI company that reads web pages so a model can use them. They fall into three groups that people routinely confuse. Training crawlers such as GPTBot, ClaudeBot and Google-Extended collect text that may be used to train future models. Search crawlers such as OAI-SearchBot, PerplexityBot and Claude-SearchBot build the index an assistant queries when it answers. Live crawlers such as ChatGPT-User and Perplexity-User fetch a page in the moment, because a user just asked something that requires it. The distinction matters because the three groups have different consequences. Blocking a training crawler affects models that do not exist yet. Blocking a search or live crawler removes you from answers being generated today.