Cloudflare's AI Fight Mode is on by default. Old robots.txt rules from 2023 still contain blocks added when AI scraping was a hot topic.
The problem
This week we worked with a client whose website was completely invisible to ChatGPT and Perplexity, despite strong rankings in classic Google search. The cause was not their content, schema, or authority signals. Their CDN had Cloudflare's AI Fight Mode enabled, blocking GPTBot, ClaudeBot, PerplexityBot and Google-Extended at the edge. AI crawlers could not even read the homepage. We turned the setting off, and within hours the AI tools could see the site. That fix took about thirty minutes. The visibility loss it had caused was probably six to nine months long.
This is not an isolated case. Cloudflare ships AI Fight Mode on by default to a large percentage of their sites. WordPress security plugins occasionally add over-zealous AI bot blocks. Old robots.txt files from 2023 still contain blocks.
The distinction nobody talks about
There are two completely different reasons to block AI bots, and they get conflated all the time.
Reason 1: stop AI training on your content. If you have proprietary content, you may not want OpenAI or Anthropic to ingest it into the next training cycle. That is a legitimate intellectual property concern.
Reason 2: stop AI tools from reading your site at query time. When a user asks ChatGPT what are the best vintage lighting shops in Bristol, ChatGPT runs a real-time web search, fetches a handful of pages, reads them, and synthesises an answer. If you block that, you are just invisible at the moment a real customer is asking about you.
The new draft standard splits them as ai-train=no (do not use my content for training) versus ai-input=yes (do read my content to answer real-time questions). Few sites use either yet, and most blanket-block both, which is the worst of both worlds.
The 30-minute audit
Layer 1: robots.txt
Type your domain followed by /robots.txt and read the file. The user agents to look for are:
- GPTBot (OpenAI training crawler)
- ChatGPT-User (OpenAI live-search agent, used at query time)
- OAI-SearchBot (newer OpenAI search-specific crawler)
- ClaudeBot (Anthropic crawler)
- anthropic-ai (older Anthropic identifier)
- PerplexityBot (Perplexity training/index crawler)
- Perplexity-User (Perplexity live agent at query time)
- Google-Extended (Google's AI training crawler, separate from Googlebot)
- CCBot (Common Crawl, used by many AI training pipelines)
If any appear under Disallow: /, you are blocking them site-wide. The pragmatic recommendation: allow the live-query agents (ChatGPT-User, Perplexity-User, OAI-SearchBot) so AI tools can read your site when a customer is asking. Make a separate, intentional decision on the training crawlers.
Layer 2: Cloudflare or your CDN
If your site sits behind Cloudflare, check two settings. Security, Bots, AI Bots: a managed list of AI bots that can be blocked en masse. Switch it to Allow if your goal is AI visibility. AI Fight Mode or Block AI scrapers: a higher-level toggle. On a different CDN (Sucuri, Akamai, Fastly), the equivalent settings are usually under Bot management or WAF rules.
Layer 3: meta tags and HTTP headers
View source on your homepage and search for any meta name=robots tag with values like noai or noimageai, or any X-Robots-Tag headers doing the same. These typically only get set deliberately, but they can sneak in via plugins.
llms.txt: the proactive complement
Consider adding an llms.txt file to your site root. It is the emerging convention for telling AI tools how to interpret your site, similar in spirit to robots.txt but in plain English. The file lives at yourdomain.com/llms.txt. Not all AI tools read it yet, but a growing number do, and adding one costs nothing. Writing it forces you to articulate, in plain language, exactly what you want AI tools to know about your business.
The rule of thumb
- Allow live-query AI agents at the robots.txt, CDN, and meta-tag layers: ChatGPT-User, Perplexity-User, OAI-SearchBot.
- Make a deliberate choice on training crawlers (GPTBot, ClaudeBot, Google-Extended). Both being in or out are valid, but make it a decision, not a default.
- Add an llms.txt if you have any structured services or content worth flagging.
- Re-audit annually, because the bot landscape and the meta-tag standards change every six to twelve months.
The audit above is something most marketing managers can do themselves with thirty minutes and a coffee. If you want a second pair of eyes, get in touch and we will run it with you on a screen-share. The fix is rarely complicated. The visibility unlock can be enormous.
By Chris McDowell, founder of BrisTechTonic. We do the technical SEO conversations that nobody else seems to be having yet, including the boring ones about CDN settings.
