SolCrys Logo

Technical Readiness

Search, Agent, or Training? How to Set AI Crawler Access in 2026 Without Blocking Your Own Citations

Allow the search crawlers and user-triggered fetchers of every engine you want to be cited by, decide AI training separately, and never block a search engine's main crawler to stop training. Since Sept. 15, 2026, choosing Block for Training in Cloudflare also stops Googlebot, Bingbot, and Applebot, while the new Disallow AI Training setting refuses training and keeps search. Google-Extended does not control AI Overviews; Googlebot and Search Console's generative AI control do. After any change to robots.txt or CDN rules, re-run the same prompts on each engine and compare which of your URLs are still cited.

Updated

Questions this guide answers

  • Should I block AI crawlers?
  • Does blocking AI training on Cloudflare block Googlebot?
  • What is the difference between OAI-SearchBot, GPTBot, and ChatGPT-User?
  • Which bots must I allow to be cited in AI answers?
  • Does blocking Google-Extended remove my site from AI Overviews?
  • What do Cloudflare's Search, Agent, and Training bot classes mean?
  • How do I see AI agent visits on my site?

Direct answer

Allow the search crawlers and user-triggered fetchers of every engine you want to be cited by, make the AI-training decision separately with training-only controls, and never block a search engine's main crawler to stop training. Since Sept. 15, 2026, selecting Block for Training in Cloudflare also stops Googlebot, Bingbot, and Applebot, the crawlers behind Google Search and its AI Overviews, Bing and Copilot, and Apple's Siri and Spotlight results. Cloudflare's new Disallow AI Training setting is the one that refuses training and keeps search.

Then verify: a robots.txt edit or CDN toggle never shows up in a rankings report, so after any change, re-run the same prompts on each engine and compare which of your URLs are still cited.

What changed: 5 Cloudflare moves since July 2025

The biggest change to crawler access in the past year happened at the CDN, not in robots.txt. Cloudflare, which says more than 20% of web domains sit behind it, went from a single "Block AI bots" switch to 3 behavior classes, and the rules it shipped on Sept. 15, 2026 differ from the ones it announced in July. Coverage written in July and August describes the announced plan, not the shipped one.

DateWhat Cloudflare didWhy it matters for citations
July 1, 2025Began asking every new domain at sign-up whether to allow AI crawlers, starting from a block by default. Launched pay per crawl in private beta, which can answer a crawler with HTTP 402 and a price.Sites onboarded since then may block AI crawlers nobody on the marketing team chose to block.
Aug. 4, 2025De-listed Perplexity as a verified bot after reporting undeclared crawling. Perplexity disputed the account.Check how your zone actually treats PerplexityBot and Perplexity-User.
Sept. 24, 2025Published the Content Signals Policy (search, ai-input, ai-train) and served search=yes, ai-train=no on the more than 3.8 million domains using its managed robots.txt.Signals are preferences, not blocks. The ai-input signal, not search, is the one that covers grounding for AI answers.
July 1, 2026Split AI traffic into Search, Agent, and Training classes for every plan, including Free, and announced new defaults for Sept. 15.Agent includes chat fetch bots such as ChatGPT-User, not only browser automation.
Sept. 15, 2026Added Disallow AI Training, extended Block to mixed-use crawlers including Googlebot, Bingbot, and Applebot, and announced that the legacy "Block AI bots" switch and managed robots.txt will be deprecated in favor of the new controls and Bot Preference Sync.The option labeled Block now removes search crawlers too.

The Sept. 15, 2026 rules, precisely

In July, Cloudflare said that from Sept. 15 "multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training." What shipped is narrower, and it depends on which Training option you pick. Search, Training, and Agent are each set per domain to Allow, Block, or Block on pages with ads; Training also gets Disallow AI Training.

  • Disallow AI Training publishes a no-training preference in robots.txt through Bot Preference Sync. In Cloudflare's words: "Accountable mixed-use crawlers remain allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI." Apple, Google, and Microsoft are the operators Cloudflare labeled Accountable for their mixed-use crawlers.
  • Block and Block on pages with ads now reach mixed-use crawlers: they "apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training."
  • Existing settings were migrated, not broken. A legacy "Block AI bots" setting became Search Allow, Training Disallow AI Training, and Agent Block on pages with ads, and earlier Training Block selections became Disallow AI Training. Cloudflare's advice to existing customers: "Nothing, in almost every case."
  • New domains are offered a preset at onboarding, which the owner can change. A site that monetizes with ads is offered Search allowed, Training set to Disallow AI Training, and Agent blocked on pages with ads. Every other site is offered all 3 allowed.
  • Bing is the exception. Until Microsoft supports a robots.txt no-training preference, "targeted for early 2027," Disallow AI Training does not convey a no-training choice to Bing. Cloudflare points site owners to Bing's NOARCHIVE tag instead.

Where the risk moved

The migration protected existing settings, so the risk now sits with the next person who edits them. "Block" reads like the strict version of "Disallow," and since Sept. 15 it is the version that also turns away Googlebot. The likely failure on a B2B site: a security or IT owner acts on a sensible instruction ("stop AI from training on our docs") without a view of what the site's AI citations depend on.

Most site owners want search, and a minority restrict training: fewer than 1% of Cloudflare sites block search bots, while 17% use some mechanism to block training. Meanwhile 52% of crawler requests on Cloudflare's network were for AI training in June 2026, up from 22% in spring 2025, so the question of restricting training will keep coming back.

Search, agent, training: which class gets you cited

Citations come from the first 2 classes. Training shapes what a model knows when it answers without searching; it doesn't produce a cited link to your page.

  • Search builds an index an engine answers from later: OAI-SearchBot, Claude-SearchBot, PerplexityBot, and the search half of Googlebot, Bingbot, and Applebot. Blocking this class takes a site out of that engine's cited answers.
  • Agent fetches a page in real time for a person. Cloudflare's definition covers "chat fetch bots (e.g., ChatGPT-User) and browser-use agents (e.g., Gemini or Claude driving Chrome)." Claude-User and Perplexity-User do the same job, and Perplexity says Perplexity-User may "include a link to the page in its response." An Agent block can cost a citation too.
  • Training feeds model weights: GPTBot, ClaudeBot, and the training use of mixed-use crawlers, which Google and Apple let you refuse with the Google-Extended and Applebot-Extended tokens.

Every major AI crawler: what it does and what blocking it costs

Checked against each company's crawler documentation on Sept. 26, 2026. User-agent names and behavior change; re-check the linked docs before you edit a rule.

BotCompanyPurposeWhat blocking it costs you in AI answersRecommendation
OAI-SearchBotOpenAISearch: surfaces sites in ChatGPT searchOpenAI: blocked sites "will not be shown in ChatGPT search answers," except as navigational linksAllow, and allow OpenAI's published IP ranges at the CDN
GPTBotOpenAITraining for OpenAI's foundation modelsNothing in ChatGPT search; OpenAI treats it as an independent settingYour call on training
ChatGPT-UserOpenAIVisits a page during a user's request; not used for search inclusion; robots.txt "may not apply"ChatGPT can't open the page mid-conversationAllow; if you must block, do it at the CDN
Claude-SearchBotAnthropicSearch indexing for Claude's search resultsAnthropic: "may reduce your site's visibility and accuracy in user search results"Allow
Claude-UserAnthropicFetches pages when a Claude user asks a questionAnthropic: "may reduce your site's visibility for user-directed web search"Allow
ClaudeBotAnthropicTrainingFuture content excluded from training; Cloudflare says blocking training-only crawlers "does not affect search"Your call on training
PerplexityBotPerplexitySearch: surfaces and links sites in Perplexity; not used for foundation-model trainingOut of Perplexity's search resultsAllow
Perplexity-UserPerplexityFetches a page for a user's question and may link it; "generally ignores robots.txt rules"Perplexity can't read or link the page for that answerAllow; the CDN is the only effective block
GooglebotGoogleGoogle Search, including AI Overviews and AI ModeOut of Google Search and its AI features; Gemini app grounding draws on the Search index too, though Google documents Google-Extended as that controlAlways allow; limit AI features with Search Console's control or preview controls instead
Google-ExtendedGoogleRobots.txt token, not a separate crawler: Gemini training plus grounding in the Gemini app and Vertex AINo effect on Google Search, AI Overviews, or AI Mode; removes your content from Gemini app groundingAllow if Gemini citations matter; blocking it is Google's training opt-out
Google-AgentGoogleAgents on Google infrastructure acting on a user's requestThe agent can't complete a task on your siteAllow on pages agents need
ApplebotAppleSearch in Spotlight, Siri, and Safari, and context for Siri and Search answers that link sourcesOut of Apple's search features and those answersAllow; use nosnippet to opt pages out of Siri and Search AI answers
Applebot-ExtendedAppleRobots.txt token for foundation-model trainingNo effect on Apple searchYour call on training
BingbotMicrosoftBing index, which Copilot and Bing's AI summaries ground inOut of Bing and Copilot citations; OpenAI's ChatGPT search help says it partners with other search providers and links Microsoft's privacy statementAlways allow; NOARCHIVE removes a page from Copilot answers and training together

Reading the table: 4 rules

Each rule guards against a mistake that can cost citations without anyone noticing.

1. Stop training with training controls, never with a search-crawler block

Use GPTBot, ClaudeBot, Google-Extended, and Applebot-Extended rules in robots.txt, or Disallow AI Training on Cloudflare. Microsoft is the exception today: Bing documents NOARCHIVE as keeping a page out of both Copilot answers and training, and NOCACHE as limiting both to the URL, title, and snippet, so until the robots.txt preference arrives there is no documented Bing setting that refuses training and keeps full Copilot citations.

2. Google-Extended is not the AI Overviews switch

Google says "robots.txt directives for Googlebot is the control" for how a site is crawled for Search, AI features included, and that Google-Extended "does not impact a site's inclusion in Google Search." Since Aug. 31, 2026, every site can use Search Console's Search generative AI control to exclude its links and content from AI Overviews, AI Mode, and generative features in Discover; Google says it "doesn't affect AI training." Google-Extended does govern grounding in the Gemini app, so blocking it has a citation cost there. For Google's AI surfaces, AEO remains an operating layer on top of SEO foundations; Google's own guide calls optimizing for them "still SEO."

3. User-triggered fetchers mostly skip robots.txt

OpenAI says robots.txt "may not apply" to ChatGPT-User, Perplexity says Perplexity-User "generally ignores robots.txt rules," and Google says its user-triggered fetchers "generally ignore robots.txt rules." Anthropic says its bots, Claude-User included, honor robots.txt. Where robots.txt doesn't apply, the CDN is the only control, in both directions: it is how you block these fetchers, and it is what silently blocks them when nobody meant to.

4. Robots.txt asks; the CDN decides

Google's list of SEO basics for AI features now includes "Ensuring that crawling is allowed in robots.txt, and by any CDN or hosting infrastructure." OpenAI's ChatGPT search help says eligibility means allowing OAI-SearchBot and confirming "that the website host or content delivery network allows traffic from OpenAI's published searchbot IP addresses." A permissive robots.txt behind a restrictive firewall rule still blocks the crawler, and nothing in robots.txt will tell you so.

Browser agents fetch as the user

Much of the Agent class doesn't look like a crawler at all. Claude in Chrome became generally available on every paid Claude plan on Aug. 26, 2026, able to act without per-step approval while a safety classifier checks each action. OpenAI's Atlas browser stopped working on Aug. 9, 2026, and OpenAI moved that agentic browsing into the ChatGPT desktop app and a ChatGPT Chrome extension. Google documents Google-Agent for agents on its own infrastructure that "navigate the web and perform actions upon user request."

Extension-based agents work inside the buyer's own browser. Reporting OpenAI's guidance, 9to5Mac noted that OpenAI recommends its Chrome extension "when a task requires an existing Chrome profile, signed-in session, open tabs, or other extensions," and that Anthropic's extension "can work with existing logged-in sessions." That visit is your buyer, delegated. Robots.txt was never written for it, and a Cloudflare Agent block, which counts browser-use agents, can turn it away mid-task. Whether the agent succeeds depends on the page: see your website's third audience. And an agent only arrives if an answer engine chose you first, the argument in identity and capability.

How to see agent visits

Start with server or CDN logs. Declared fetchers identify themselves by user agent (ChatGPT-User, Claude-User, Perplexity-User, Google-Agent), and OpenAI, Perplexity, and Google publish the IP ranges to check them against, because user-agent strings can be spoofed. Cloudflare's AI Crawl Control reports AI crawler traffic by user agent on the Free plan. Some agents sign their requests: Cloudflare reported in August 2025 that ChatGPT agent uses the proposed Web Bot Auth standard, and Google says it is experimenting with Web Bot Auth for Google-Agent. An extension acting in a signed-in browser may be indistinguishable from the user, so no log gives a complete count.

llms.txt and Content Signals: preferences, not access

Neither file grants or denies access. Google's guide says you don't need to create "new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." Asked on Jan. 20, 2026 whether llms.txt files on Google's own properties were an endorsement, John Mueller answered: "I'm tempted to say something snarky since this has come up so often, but to be direct, no." None of the crawler documentation from OpenAI, Anthropic, Perplexity, Google, or Apple cited here says its crawler reads llms.txt. Google's Lighthouse tool does check for it: version 13.3 added an llms.txt audit to an experimental Agentic Browsing category in May 2026, according to Search Engine Journal, as a readiness check for browser agents rather than a search signal. Our longer argument is in why llms.txt is not a strategy.

Content Signals are different: they state what a crawler may do after fetching, and Cloudflare says plainly that "Some companies might simply ignore them." If you use them and want to be cited, the signal to leave alone is ai-input, which covers "retrieval augmented generation, grounding, or other real-time taking of content for generative AI search answers." Cloudflare's own definition of search excludes AI summaries, so search=yes with ai-input=no reads as a request to stay out of AI answers.

A configuration that keeps your citations

For a B2B brand that wants to be cited and doesn't sell ad inventory, this is a sound default. Adjust the training lines to your own policy; leave the rest alone unless you mean to leave an engine.

  • Robots.txt: allow OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot, Bingbot, and Applebot. Decide GPTBot, ClaudeBot, Google-Extended, and Applebot-Extended one by one, remembering that Google-Extended also covers Gemini app grounding. If robots.txt doesn't mention Applebot but does mention Googlebot, Apple says Applebot follows the Googlebot rules.
  • Cloudflare: Search set to Allow; Agent set to Allow; Training set to Allow or Disallow AI Training, and set to Block only if you intend to leave Google, Bing, and Apple search.
  • Firewall and bot rules: check custom rules and bot-fighting features that block by user agent or network, and confirm that OpenAI's and Perplexity's published IP ranges get through.
  • Search Console: leave the Search generative AI control on Include unless you intend to leave AI Overviews and AI Mode.
  • Bing: keep NOARCHIVE off pages you want cited in Copilot.
  • Change log: record the date of every robots.txt, CDN, and firewall change, so the next step has a before and an after.

Verify: did the change cost you citations?

Only an outcome check answers that, and it has to be engine by engine, because each engine reads a different set of controls. Before a change, capture which of your URLs each engine cites on the prompts you track. After it, re-run the same prompts on the same engines and compare. A drop confined to one engine points first to that engine's crawler; a drop on every engine points first to the CDN.

Allow for lag. OpenAI says a robots.txt change takes about 24 hours to reach ChatGPT search, and Perplexity says up to 24 hours. Google says recrawling can take "several days to several months," and a Search Console control change generally takes a few days. The free platform reports cover part of the picture: Bing Webmaster Tools' AI Performance report, introduced in public preview in February 2026, counts citations in Copilot and Bing's AI summaries, and Search Console's generative AI performance report shows impressions in Google's AI features. Read any before-and-after as an observation, not proof, since engines change what they cite on their own; how many runs you need covers the noise.

How SolCrys fits

SolCrys does not read or change your CDN, firewall, or bot-management configuration, and it can't tell you whether Cloudflare lets the real Googlebot through. It helps in 3 narrower ways.

  • The Content Audit's AI Bot Check. For a live page, the audit reads robots.txt and requests the page as 7 named crawlers (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Perplexity-User, Googlebot, and Bingbot), reporting each as OK, Blocked, or Degraded. Those requests come from SolCrys's servers, not from OpenAI's or Google's published IP ranges, so a CDN that verifies crawlers by IP can treat the real crawler differently. The check does not cover Claude-SearchBot, Claude-User, ChatGPT-User, or Applebot. Treat it as a robots.txt and origin-server check, not a CDN audit.
  • Verify on the prompts you already track. Your frozen prompt set keeps running on the same engines (ChatGPT, Google AI Overviews, Gemini, Perplexity, and Claude, depending on plan), so after a bot-policy change you can compare which of your URLs each engine cites, before and after. That is the Verify step of Measure → Diagnose → Execute → Verify, and it answers the question logs can't: whether a CDN toggle cost you ChatGPT or Perplexity citations. SolCrys does not measure Copilot or Siri; for Copilot, use Bing Webmaster Tools.
  • A deliberate agent surface. We run a remote MCP server with OAuth, so assistants such as Claude, ChatGPT, and Cursor read a workspace through typed tools, and write only through narrow opt-in permissions, instead of driving our dashboard through a browser. Most marketing sites don't need one. The point is that agent access worth having is designed and permissioned, not left to whatever gets past the bot rules.

Next step

To see which of your pages ChatGPT cites today, Start Free: the free workspace runs 10 tracked prompts on ChatGPT and 1 Content Audit (free, no credit card, email verification). To verify a crawler-policy change across engines, talk to us.

Sources

FAQ

Should I block AI crawlers?

Block training if your policy requires it, but don't block the crawlers that put your pages in AI answers. For a brand that wants to be cited, allow the search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot, Applebot) and the user-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User), and handle training with training-only controls such as GPTBot, ClaudeBot, Google-Extended, and Applebot-Extended rules in robots.txt, or Cloudflare's Disallow AI Training setting. Fewer than 1% of Cloudflare sites block search bots; 17% restrict training.

Does blocking AI training on Cloudflare block Googlebot?

It depends on the setting. Since Sept. 15, 2026, setting Training to Block, or Block on pages with ads, applies to mixed-use crawlers; Cloudflare says Block "will stop Applebot, Bingbot, and Googlebot from reaching your site," search included. Setting Training to Disallow AI Training refuses training while Googlebot, Bingbot, and Applebot keep crawling for search. Existing "Block AI bots" and Training Block settings were migrated to Disallow AI Training on Sept. 15, so the risk is a future edit, not the migration.

What is the difference between OAI-SearchBot, GPTBot, and ChatGPT-User?

OAI-SearchBot is OpenAI's search crawler; OpenAI says sites that block it "will not be shown in ChatGPT search answers," except as navigational links. GPTBot collects content that may be used to train OpenAI's foundation models, and blocking it doesn't affect search. ChatGPT-User visits a page during a user's request, isn't used to decide search inclusion, and, because a user initiated the visit, robots.txt rules "may not apply." Allow OAI-SearchBot and ChatGPT-User if you want ChatGPT citations; decide GPTBot on training policy alone.

Does blocking Google-Extended remove my site from AI Overviews?

No. Google-Extended is a robots.txt token that controls whether Google may use your content to train Gemini models and to ground answers in the Gemini app and Vertex AI; Google says it "does not impact a site's inclusion in Google Search." AI Overviews and AI Mode follow Googlebot. To leave those features, use the Search generative AI control in Search Console (available to all sites since Aug. 31, 2026), or limit what they show with preview controls such as nosnippet. Blocking Google-Extended can cost you citations in the Gemini app.

Which bots must I allow to be cited in AI answers?

By engine: OAI-SearchBot and ChatGPT-User for ChatGPT; Claude-SearchBot and Claude-User for Claude; PerplexityBot and Perplexity-User for Perplexity; Googlebot for AI Overviews and AI Mode, plus Google-Extended if you want Gemini app grounding; Bingbot for Copilot; and Applebot for Siri and Spotlight. Allowing them in robots.txt isn't enough if your CDN or firewall blocks them, and OpenAI explicitly asks sites to let its published search-bot IP addresses through.

Do AI agents follow robots.txt?

Some do. Anthropic says its bots, including Claude-User, honor robots.txt. OpenAI says robots.txt "may not apply" to ChatGPT-User, Perplexity says Perplexity-User "generally ignores robots.txt rules," and Google says the same of its user-triggered fetchers, including Google-Agent. Browser extensions such as Claude in Chrome act inside the user's own browser session, where robots.txt was never the control. For agents, the CDN and the page itself are the real controls.

How do I see AI agent visits on my site?

Check server or CDN logs for declared user agents such as ChatGPT-User, Claude-User, Perplexity-User, and Google-Agent, and verify them against the IP ranges OpenAI, Perplexity, and Google publish, since user-agent strings can be spoofed. Cloudflare's AI Crawl Control reports AI crawler traffic by user agent on its Free plan, and some agents sign requests with the proposed Web Bot Auth standard. Extension-based agents working in a signed-in browser can look like ordinary users, so treat any count as a floor.

Does llms.txt control AI crawler access?

No. llms.txt is a proposed summary file, not an access control; robots.txt and your CDN decide access. Google says Google Search doesn't use llms.txt, and none of the crawler documentation from OpenAI, Anthropic, Perplexity, Google, or Apple says its crawler reads it. Google says publishing one neither helps nor harms visibility in Google Search, and it won't unblock a crawler your firewall stops or block one your robots.txt allows.

Free ChatGPT visibility check

See where AI answers skip your brand — then fix it, free

Start a free workspace with your domain: 10 buyer-intent prompts through ChatGPT show where you are mentioned, cited, or skipped, and who gets recommended instead. A free content audit in the same workspace hands you the first fix to ship.

Start Free

Free · No credit card · About 5 minutes

Related guides

Strategy & Positioning

Why llms.txt Is Not a Strategy

llms.txt is a proposed standard for AI-friendly content delivery, but it is neither widely adopted by major AI engines nor a substitute for AEO fundamentals. This essay explains what llms.txt does, what it does not, and why brands should focus on the unsexy basics.

Strategy & Positioning

Your Website Has a Third Audience Now — Agents

AI agents are a third audience for your website — alongside humans and search crawlers. Eason Wang on what changes for product UX when agents see what humans can't, and miss what humans take for granted.

How SolCrys Works

What a SolCrys Content Audit Looks Like

What a SolCrys Content Audit checks on one URL, how the score and recoverable points work, and how a fix is promoted, published, and re-audited.