Most site owners who try to block AI crawlers want one thing: stop AI companies from training on their content. Many end up doing something else by accident, like disappearing from ChatGPT search, or blocking nothing at all because they used a bot name that was retired two years ago. This guide explains which crawlers exist, what each one controls, the rules to copy, and the parts robots.txt simply can’t do.
Before you change anything, see where you stand: the free AI Crawler Access Checker reads your robots.txt and shows which AI bots are allowed or blocked, and why.
Why “block AI crawlers” is really three decisions
AI companies no longer use one bot each. Since late 2024 the big ones split their crawlers by purpose, and each name needs its own line in robots.txt:
| Company | Training | Search and citations | When a user asks |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | (none) | PerplexityBot | Perplexity-User |
| Google-Extended | Googlebot | ||
| Apple | Applebot-Extended | Applebot | |
| Meta | meta-externalagent | meta-externalfetcher |
So the real questions are:
- Training: may AI companies use your pages to build future models? Blocking this has no effect on search traffic.
- AI search: may ChatGPT, Claude, Perplexity and Siri find and cite your pages? Blocking this removes you from their answers, and from the visitors those citations send.
- User fetches: may an assistant open your page when someone pastes your link? Blocking this mostly annoys your own readers.
For most blogs, shops and tool sites the sensible answer is no, yes, yes: block training, keep search. If you sell the content itself (paywalled journalism, research, stock photos), blocking everything is a reasonable business choice. Just make it on purpose.
The robots.txt rules to copy
This block stops the main training crawlers and leaves search and user fetches alone. Add it at the end of your robots.txt, below your existing rules (new to the file format? Start with how to write a robots.txt file):
# Block AI training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
User-agent: Bytespider
User-agent: Amazonbot
User-agent: cohere-training-data-crawler
Disallow: / Stacking several User-agent lines above one Disallow is valid: they form one group, and every crawler named in it follows the rule.
To block every AI crawler, including search, add the search and user-fetch names to the same group: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User and meta-externalfetcher. Don’t add Googlebot or Bingbot unless you want to leave Google and Bing altogether.
On WordPress you don’t need FTP. Yoast SEO has a robots.txt editor under Yoast SEO → Tools → File editor, and Rank Math has one under General Settings → Edit robots.txt. If neither shows up, WordPress is generating a virtual robots.txt and you can create a real file in your site’s root folder through your host’s file manager.
Three mistakes that quietly break the rules
1. A crawler’s own group cancels your general rules
This one catches a lot of people. A crawler follows only the most specific group that names it. If your file says
User-agent: *
Disallow: /wp-admin/
Disallow: /checkout/
User-agent: OAI-SearchBot
Allow: / then OAI-SearchBot ignores the first group completely and may crawl /wp-admin/ and /checkout/. If you give a crawler its own group, repeat the Disallow lines you still want inside it. The checker warns about this, and its rule builder copies your private paths into each group for you.
2. Retired names
Plenty of copy-paste lists still block anthropic-ai and Claude-Web. Anthropic no longer uses either, so those lines control nothing, and sites that rely on them are not blocking ClaudeBot at all.
3. Blocking search while allowing training
We regularly see files that block OAI-SearchBot or PerplexityBot “to be safe” but never mention GPTBot. The result is the worst of both: out of ChatGPT search results, still available for training.
What robots.txt can’t do
This is the part most guides skip. robots.txt is a public request. It works because reputable companies choose to honour it, and it has real limits:
- It isn’t enforced. OpenAI, Anthropic, Google and Apple say their crawlers follow it. Bytespider has repeatedly been reported to ignore it, and any scraper can pretend to be a normal browser.
- User fetches are a grey zone. Perplexity says Perplexity-User generally ignores robots.txt, OpenAI says robots.txt rules may not apply to ChatGPT-User, and Meta says the same of meta-externalfetcher. Their argument is that a person, not a crawler, asked for the page.
- It isn’t retroactive. Blocking GPTBot today doesn’t remove pages already collected or models already trained. Common Crawl archives go back more than a decade.
- It can’t separate Google Search from AI Overviews. Both use Googlebot. Google-Extended only covers Gemini training and grounding. The only way to keep a page out of AI Overviews is
nosnippet(ornoindex), which also removes its snippet from normal results. For most sites that trade isn’t worth it. - noai and noimageai tags are wishful. These meta tags are not part of any standard and none of the major AI companies has said it follows them. Adding them does no harm, but don’t count on them.
If blocking really matters to you, add a network-level block as well. Cloudflare can block known AI crawlers for all plans, and it checks the real IP addresses of verified bots, which robots.txt can’t. Security plugins such as Wordfence can block by user agent. Be aware these tools cut both ways: Cloudflare now blocks AI crawlers by default on many new domains, so a site can shut out AI search bots its owner meant to allow while its robots.txt still says “allowed”. The checker’s firewall test compares the answer your homepage gives a browser with the one it gives GPTBot and ClaudeBot, which catches exactly this.
Content Signals and llms.txt: worth adding?
Cloudflare’s Content Signals Policy adds a line such as Content-Signal: search=yes, ai-input=yes, ai-train=no to robots.txt. It states how you want content used after it’s fetched: for search, as input to AI answers, or for training. It’s a clear statement of your wishes and costs nothing, but no crawler is obliged to follow it, and major AI companies haven’t said publicly that they do. Add it if you like; don’t treat it as protection.
llms.txt goes the other way: instead of keeping AI out, it points assistants to your best pages. Adoption by AI companies is still unclear, so treat it as cheap and optional rather than as a ranking lever.
How to check that it worked
- Open
yoursite.com/robots.txtin a private browser window. If you see your old file, clear your cache plugin or CDN cache. If you see a web page instead of plain text, your robots.txt isn’t set up correctly. - Run the AI Crawler Access Checker. Training rows should say Blocked, search rows Allowed, and the notes should be free of warnings.
- Test a private path, such as
/wp-admin/, in the path box to make sure the crawlers you allowed still can’t open it. - Give it a day. OpenAI and Perplexity say changes take about 24 hours to reach their crawlers, and others re-read robots.txt on their own schedule.
- If you have server logs, look for the crawler names after a week. Hits on blocked paths from a name that claims to obey robots.txt usually mean someone is faking that name.
The short version
To block AI crawlers from training while staying visible in AI search, block GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot and Bytespider, and leave OAI-SearchBot, Claude-SearchBot and PerplexityBot alone. Don’t use retired names, don’t forget that a crawler’s own group overrides your general rules, and don’t expect robots.txt to stop anyone who chooses to ignore it. If that matters, back it with a firewall rule.
Check your site now: the AI Crawler Access Checker shows every AI crawler’s access in seconds and builds corrected rules you can paste straight into your robots.txt.
Keep reading
More from SEO & Website Tools.



Tools for this job
Free, in your browser, no sign-up.