Free tool

AI crawler access checker

Enter your website and see, crawler by crawler, whether your robots.txt lets the AI engines in, which rule decides it, and what each crawler is actually used for.

Free, no sign-up. Results for the same URL are cached for an hour.

Why this matters

Every AI answer engine fetches pages with a named crawler, and most of them respect robots.txt. A single "Disallow: /" aimed at the wrong user agent, or a well-meant block on a training bot that also carries the search bot's rules, removes a site from AI answers without any error showing up anywhere. Many sites added such blocks in 2023 and 2024 and never revisited them.

The checker reads the live file, applies it the way the crawlers do, and separates the crawlers that feed answers from the crawlers that feed training.

How to use it

  1. Enter the domain to check the home page path, or a full URL to check a specific path such as /blog/.
  2. Read the verdict column. "Blocked" shows the exact rule and the User-agent group that produced it.
  3. Fix the file: remove the block, or add an Allow line for the search crawlers you want, and re-run.
  4. Check the other signals: a declared sitemap, a Content-Signal line and an llms.txt all help engines understand the site.

Training crawlers versus search crawlers

The vendors split their crawlers by purpose, and the split is what makes a sensible policy possible:

  • Search and grounding crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot) index pages so they can be cited in answers. Block these and you disappear from the answers.
  • User-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User) load a page when a user asks about it. Block these and the engine cannot read your page even when a user points at it.
  • Training crawlers (GPTBot, ClaudeBot, CCBot, Bytespider, meta-externalagent, Applebot-Extended, Google-Extended) collect text for model training. Blocking them is a policy choice that does not remove you from search-style answers.

What it does not check

  • Firewall and bot-manager rules. robots.txt can allow a crawler that your CDN then blocks.
  • Meta robots tags and X-Robots-Tag headers on individual pages.
  • Whether a crawler ignores robots.txt. Some do; the file is a request, not enforcement.

FAQ

Which AI crawlers does the checker test?

OpenAI (GPTBot, OAI-SearchBot, ChatGPT-User), Anthropic (ClaudeBot, Claude-SearchBot, Claude-User), Perplexity (PerplexityBot, Perplexity-User), Google (Googlebot, Google-Extended), Microsoft (Bingbot), Apple (Applebot-Extended), Amazon (Amazonbot), Meta (meta-externalagent), ByteDance (Bytespider) and Common Crawl (CCBot).

What is the difference between GPTBot and OAI-SearchBot?

GPTBot collects pages for training OpenAI's models. OAI-SearchBot indexes pages so they can appear as sources in ChatGPT search. Blocking GPTBot does not stop your pages from showing up in ChatGPT answers; blocking OAI-SearchBot does.

Does blocking Google-Extended remove me from AI Overviews?

No. Google-Extended only controls whether your content is used for Gemini training and grounding. AI Overviews and AI Mode use the normal Googlebot crawl, so the only way to stay out of them is to stay out of Google Search.

Should I block AI crawlers?

If you want to be cited in AI answers, allow at least the search and user-triggered crawlers (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User). Blocking the training crawlers (GPTBot, ClaudeBot, CCBot) is a separate decision about model training and does not affect citations in the same way.

My robots.txt allows everything but a crawler still cannot reach the site. Why?

robots.txt is only a request. A CDN, firewall or bot-management rule can block a crawler by user agent or IP regardless of the file. Check those rules, and check your server logs for 403 responses to the crawler's user agent.

What does the Content-Signal line mean?

Content-Signal is a newer robots.txt directive, promoted by Cloudflare, that states how content may be used: for search, as AI input, or for AI training. Not every crawler honors it yet, but it is a clear statement of intent next to your Allow and Disallow rules.

How does the checker decide which rule applies?

It follows the Robots Exclusion Protocol: the group whose User-agent token best matches the crawler name applies (falling back to *), the longest matching Allow or Disallow path wins, and Allow wins a tie. That is also how Google and OpenAI document their behavior.

More free tools

Being crawlable is step one. Being cited is the goal.

LeapScope tracks whether the engines you let in actually mention and cite you for the prompts your buyers ask, every day.