...
Contact Us Contact Us

AI Crawler Access Checker

Reads the site’s robots.txt and llms.txt, then checks the homepage for AI directives and user-agent blocking. Paste a full page URL to test that page.

What this AI crawler checker does

This AI crawler checker tells you which AI bots can read your website and which are blocked. Enter a domain and it reads your robots.txt, then shows the result for 32 crawlers from OpenAI, Anthropic, Google, Perplexity, Apple, Meta, Amazon, ByteDance, Common Crawl and others. For every crawler you see the exact rule and line number that decides it, so you know what to change, not just what is wrong.

The crawlers are split by what they do, because blocking them has very different effects:

  • AI search and answers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot): these decide whether ChatGPT, Claude, Perplexity and Siri can find and cite your pages. Block them and you disappear from those answers.
  • AI training (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider and more): these collect pages that may be used to train models. Blocking them does not affect your search rankings.
  • Fetches a user asked for (ChatGPT-User, Claude-User, Perplexity-User): an assistant opening your page because someone pasted the link or asked about it.

How to use it

  1. Check a site. Type your domain, for example example.com, and press Check access. To test one page, paste its full URL instead.
  2. Read the summary. The three tiles show how many search, training and user-fetch crawlers are allowed. The panel below lists your robots.txt, whether you have an llms.txt, any noindex, nosnippet or noai tags on the homepage, and a firewall test.
  3. Look at the notes. The checker flags common mistakes in plain English: a crawler with its own group that silently ignores your User-agent: * rules, blocking ChatGPT search while allowing OpenAI training, retired names such as anthropic-ai, and robots.txt files that are really HTML pages.
  4. Test a path. Type any path, such as /blog/ or /wp-admin/, to see which crawlers may open it. Click a line number to jump to that line in your file.
  5. Fix it. Press Fix it: build rules. Choose a policy, adjust single crawlers, and copy a complete robots.txt that keeps your existing rules and sitemaps. Press Test it to check the new file before you upload it.

No live site yet, or behind a login? Use the Paste robots.txt tab. Paste or upload the file and the results update as you type, without anything leaving your browser.

What the firewall test shows

robots.txt is a request, not a lock. Many sites also block AI bots at the server or CDN level, sometimes without the owner knowing, because a security plugin or a Cloudflare setting did it. The checker requests your homepage three times: as a normal browser, as GPTBot and as ClaudeBot. If the crawler versions get a 403 error or a challenge page while the browser gets through, something is blocking AI crawlers by name, whatever your robots.txt says.

The test can only see blocking based on the user-agent name. Rules that check the crawler’s real IP address, such as Cloudflare’s verified-bot blocking, can’t be detected from outside, so a clean result doesn’t rule them out.

What the results can’t tell you

Honest limits, so you don’t read more into a green result than it means:

  • A crawler that is “allowed” may still never visit. Being allowed makes you eligible; it doesn’t make anyone crawl you.
  • Not every crawler obeys robots.txt. The Obeys robots.txt column says so where the company itself says it may not (Perplexity-User, ChatGPT-User, meta-externalfetcher) or where it has been widely reported (Bytespider).
  • Blocking a training crawler today does not remove pages it already collected.
  • Googlebot feeds both Google Search and AI Overviews. robots.txt can’t separate the two; only nosnippet or noindex keeps a page out of AI Overviews.

Frequently asked questions

How do I check if my site blocks GPTBot?

Enter your domain above and look at the GPTBot row in the AI training group. It shows Allowed or Blocked and the rule that decides it. Also check the Firewall test line: if GPTBot gets an error there, your server or CDN blocks it even if robots.txt allows it.

Will blocking AI crawlers hurt my Google rankings?

No, as long as you only block AI crawlers. Google-Extended controls Gemini training and grounding and has no effect on Google Search. Never block Googlebot itself unless you want to leave Google.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot collects pages that may be used to train OpenAI models. OAI-SearchBot indexes pages so they can appear and be cited in ChatGPT search. You can block the first and allow the second, which is what most publishers do.

Does blocking ClaudeBot also block Claude search?

No. Anthropic uses three separate names: ClaudeBot for training, Claude-SearchBot for search and Claude-User for pages a user asks Claude to open. Each needs its own rule.

My robots.txt blocks a crawler but it still visits. Why?

Either the crawler doesn’t follow robots.txt, someone is using its name without being it, or it fetched a page a user asked for, which some companies treat differently. For crawlers that matter to you, add a firewall or CDN rule as well.

What happens if I don’t have a robots.txt?

Every crawler may read every page. That’s fine if you want to be visible to AI tools. If your robots.txt returns a server error instead, compliant crawlers do the opposite and stay away from the whole site, so fix that first.

Is my data stored?

Pasted files never leave your browser. For the Check a site tab, our server fetches the public robots.txt, llms.txt and homepage of the domain you enter and keeps that result for six hours so repeat checks are instant. Nothing about you is stored.

Where can I learn more?

Read our guide How to Block AI Crawlers with robots.txt (and When Not To). Use the robots.txt Generator for general crawler rules.

Seraphinite AcceleratorOptimized by Seraphinite Accelerator
Turns on site high speed to be attractive for people and search engines.