Is your site blocking AI crawlers? How to check in five minutes
Check whether your robots.txt or firewall blocks GPTBot, OAI-SearchBot, ClaudeBot or Google-Extended, and decide which AI crawlers to allow.
You can spend months making your site the best answer on the internet and still be invisible in ChatGPT, Claude or Perplexity for one boring reason: a line in a text file told their crawlers to stay out. Sometimes a developer added it years ago. Sometimes a security setting did it for you without asking. Either way, the first step is figuring out what is actually there.
Here is how to check whether you block AI crawlers, what each crawler does, and how to decide which ones to let in.
Why AI crawlers are not all the same
The first thing to know is that most AI companies run more than one crawler, and each one has a different job. Blocking one does not automatically block the others.
OpenAI's documentation on its bots lists three:
- OAI-SearchBot finds pages to show in ChatGPT's search results.
- GPTBot collects content that may be used to train OpenAI's foundation models.
- ChatGPT-User visits a page when a person in ChatGPT asks about it. OpenAI says it is not used to crawl the web automatically. OpenAI notes that because a person starts these visits, robots.txt rules may not apply to ChatGPT-User.
OpenAI is explicit that the settings are independent. A site can allow OAI-SearchBot so it can appear in ChatGPT search while disallowing GPTBot, which signals that its content should not be used for training. OpenAI also notes it can take about 24 hours for a robots.txt change to show up in its search results.
Anthropic splits its crawlers the same way. Its help article on web crawling describes ClaudeBot (content that may help train its models), Claude-User (pages fetched when a person asks Claude a question) and Claude-SearchBot (indexing to improve search results). Anthropic warns that disabling Claude-SearchBot may reduce your site's visibility and accuracy in user search results.
Google works differently. Google-Extended is not a separate crawler at all. Google's list of common crawlers describes it as a robots.txt token that controls whether content Google crawls can be used to train future Gemini models and to ground Gemini's answers. Google says it has no effect on Google Search. So blocking Google-Extended does not remove you from search results, and allowing it does not change your rankings.
How to check your robots.txt in five minutes
- Open your browser and go to yourdomain.com/robots.txt.
- Look for any group that starts with
User-agent:followed by one of these names: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User or Google-Extended. - Under each one, read the next lines.
Disallow: /means that crawler is blocked from the whole site.Disallow: /blog/blocks only that folder. - Check the catch-all group,
User-agent: *. ADisallow: /there blocks every crawler that does not have its own group, AI crawlers included. - If the file does not exist, crawlers treat the site as open.
One thing robots.txt does not do: keep pages out of Google. Google's robots.txt introduction says the file is mainly for managing crawler traffic, and that it is not a mechanism for keeping a page out of Google. For that, Google points to a noindex rule or password protection.
The block you cannot see in robots.txt
A clean robots.txt does not guarantee access. Many sites sit behind a CDN or firewall with bot settings that block crawlers before they ever reach the file.
This became much more common in 2025. On July 1, 2025, Cloudflare announced it was changing the default to block AI crawlers unless they pay creators for content. If your site runs on Cloudflare, or on any host with "bot protection" switched on, open the security or bot settings and look for AI crawler rules. Your robots.txt can say "welcome" while your firewall says "go away."
Should you block AI crawlers?
There is no single right answer, but the decision gets easier once you separate the jobs:
- Search and user crawlers (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User) are how your pages get found and quoted when someone asks an AI assistant a question. If you want customers to discover you there, blocking these works against you.
- Training crawlers (GPTBot, ClaudeBot) and the Google-Extended token are about whether your content helps train future models. Some publishers block these on principle or for licensing reasons. For most small and mid-size businesses, the bigger risk is being absent from AI answers, not being included in training data.
A common middle path is to allow the search and user crawlers, then decide on training separately.
What to do next
Check robots.txt. Then check your firewall or CDN bot settings. Then ask ChatGPT, Claude and Perplexity the questions your customers ask and see whether your site shows up. If it does not, crawler access is the first thing to rule out, and the check takes five minutes.
Sources
Checked October 3, 2026. Platforms change their guidance. The linked pages are the final word.






