A&A INSIGHTS
Before you block AI crawlers: decide search, agent and training separately
Allowing or blocking AI crawlers is not one switch. Split it by behaviour — search, agent, training — and by whether the URL is a discovery or a handover surface.
日本語で読む
THE STARTING POINT
Do not decide AI crawler access with a single switch. Split it by three behaviours — fetching to answer questions later, acting in real time on a person's behalf, and fetching for training — and then by whether the URL is a surface you need to be found on or one you use to hand work to a customer. A blanket block at the CDN lands asymmetrically: robots.txt is only a request, so absorption into models is not reliably stopped, while the CDN block is enforcement, so the path by which people find you closes for certain.
"Should I block AI bots?" is not the unit of the decision
A&A perspective
Whether to allow or block AI crawlers is not a single switch. There are two units to decide on. The first is the bot's behaviour: is it fetching in order to answer questions later, is it acting in real time on behalf of a person to get something done, or is it fetching in order to train a model? The second is the area of your own site: is this URL a surface you need to be found on, or a surface you use to hand work over to a customer? Cross those two before you touch any setting. This is A&A's own position.
From the sources
In its own documentation Cloudflare writes that AI crawlers and agents interact with your site for very different reasons, and that you may want to treat those reasons differently. It then states that rather than relying on a single "AI bot" label, it classifies bots by behaviour — what a bot does on your site — so that you can allow the behaviour that helps your business and block the behaviour that harms it.
From the sources
On the same page Cloudflare says it lets all customers manage three AI-related use cases directly, and defines each one. Search collects or indexes your content so it can answer questions about it later. Agent is automated activity acting in real time on a person's behalf to get something done, with chat fetch bots and browser-use agents given as the examples. Training crawls your content to train or fine-tune a model, permanently absorbing your data into the model.
A blanket block lands asymmetrically: it misses what you wanted stopped and closes what you wanted open
From the sources
In "AI features and your website", Google Search Central writes that AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search. It then points to a separate control: to limit AI training and grounding in some of Google's other systems, read more about Google-Extended.
From the sources
The same page says that no specific optimization is required for these AI features, and then lists, as examples, the SEO fundamentals that remain worthwhile. The first item in that list is ensuring that crawling is allowed in robots.txt, and by any CDN or hosting infrastructure.
From the sources
Cloudflare's AI Crawl Control is described as giving you visibility into which AI services are accessing your content, and as letting you set allow or block rules for individual crawlers. Its feature list includes monitoring robots.txt compliance: track which crawlers follow your directives and create enforcement rules. The product is described as available on all plans and as working automatically with zero configuration.
A&A perspective
The publisher of the search engine is itself naming your CDN as a place where crawling can be switched off. Start there, put that next to the compliance tracker, and the asymmetry appears. A robots.txt directive is a request, not enforcement. That is precisely why a compliance tracker can exist as a product feature at all. The inference that tracking is needed because some crawlers do not comply is ours, not a statement Cloudflare makes. A block at the CDN, by contrast, is enforcement. So the blanket, label-based switch fails to reliably stop the thing you wanted stopped — absorption into models — while reliably closing the thing you did not want closed: the path by which people and AI answers find you. You meant to fail safe, and the direction of the failure is inverted. That is the shape of the problem.
| Area of the site | Do you need to be found here? | Default handling |
|---|---|---|
| Public blog articles | Yes — in practice the entrance for inquiries | Allow search and agent. Decide training on commercial grounds |
| Services and how pricing is approached | Yes — content you want AI answers to draw on | Allow search and agent. Decide training on commercial grounds |
| Write-ups of past work | Yes | Allow search and agent. Whether you may publish it is a contract question first |
| Per-customer shared URLs for proposals and quotes | No | Deny everything. Close it with authentication and non-public settings anyway |
| Deliverables and working files | No | Deny everything. An authentication boundary that sits earlier than bot classification |
| Admin interface and API | No | Deny everything, regardless of bot type |
An allowlist built from bot names starts ageing the day you write it
From the sources
Where it explains classification by behaviour, Cloudflare states plainly that a single bot can have more than one behaviour.
From the sources
Google has the same shape. Google-Extended is presented as the control for limiting AI training and grounding in some of Google's other systems, and it is separate from the Googlebot robots.txt directives that govern crawling for Search itself. One provider, several fetching paths, different purposes.
A&A perspective
What follows is that a decision to allow or block a given crawler name also silently decides the other use cases that same crawler is carrying out. This is why an allowlist of bot names starts ageing the day you write it. Providers add new crawlers for new purposes, and existing crawlers acquire additional purposes. What you can actually maintain is not a list of names but a position on each behaviour: fetching in order to answer questions later is allowed; fetching in real time on a person's behalf is allowed; fetching for training is handled this way. Decide that, and an unfamiliar name arriving next month has somewhere to land. Try to maintain the list of names instead and you have added a monthly chore — one that will always be late. This is A&A's own position.
Before the behaviours, take inventory of your own URLs

A&A perspective
A position on behaviours does not yet produce a setting. The same rule — allow fetching for search — gives opposite answers for a public description of your services and for a shared URL holding one customer's quote. What has to come first is an inventory that splits your own domain into surfaces you need to be found on and surfaces you use to hand work over.
A&A perspective
In a solo or small service business the two usually live on the same domain. Public articles and service descriptions are the first kind, and in practice they are where inquiries come from. Shared proposal links, delivery file locations and dashboards for work in progress are the second kind, and you do not want anyone to find them at all. For that second group the behaviour discussion is unnecessary: deny everything is the right answer. More to the point, that group belongs behind authentication and non-public settings, which sit earlier in the chain than any bot classification. Do not use AI crawler settings as a substitute for a login.
Hypothetical example
As a hypothetical example, take inventory of the domain of a one-person development shop and you tend to get six areas: public blog articles; the description of services and how pricing is approached; write-ups of past work; per-customer shared URLs for proposals and quotes; the location holding deliverables and working files; and the admin interface plus API. This is not a real customer's configuration — it is a hypothetical used to show the shape of the decision. The first three are discovery surfaces, the last three are handover surfaces, and only the first three have a per-behaviour decision to make at all.
Write the default for each area on one sheet
A&A perspective
The inventory fits on one sheet with three columns: the area, whether you need to be found there, and the default handling. The table below turns the hypothetical from the previous section into that sheet. It is A&A's own arrangement: the behaviour names follow Cloudflare's classification, but the way the areas are cut and the defaults are chosen are not in the sources.
From the sources
What it means to set the search column to allow can be checked against Google's documentation. To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. The page also states that there are no additional technical requirements. A page that is not indexed is not a candidate for that entrance in the first place.
A&A perspective
The training column, by contrast, can be decided on commercial grounds and kept separate from discovery. Denying it does not close the entrance described above, as long as search and agent fetching remain allowed. Allowing it does not buy citations or referrals either — neither our own records nor the sources read here support that. Whether to permit training is properly a decision about terms: how you feel about what you wrote being permanently absorbed into a model. Mix it into the discovery question and both decisions get muddy.
Allowing access does not buy you a citation
From the sources
On the same page Google writes that there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary. It goes further: you don't need to create new machine readable files, AI text files, or markup to appear in these features, and there is also no special schema.org structured data that you need to add.
From the sources
The same page also states that just because a page meets all requirements, best practices, and complies with the policies, doesn't mean that Google will crawl, index, or serve its content: indexing and serving isn't guaranteed.
A&A perspective
Two things follow. First, adjusting crawler settings is not a tactic for getting cited. It is hygiene that keeps you from removing your own precondition for being found. Blur that distinction and you will score the month you fixed the setting as a failure because inquiries did not rise. Second, do not count having placed a machine-readable file as having handled AI. For Google Search at least, the provider itself writes that it is unnecessary. Other engines may treat it differently, but we have not verified that, and this article does not lean on it as evidence.
Do not count fetches as traffic or as demand
From the sources
AI Crawl Control is described as letting you monitor the dashboard to see crawler activity and request patterns, so you can see which AI services access your content. It also lists a function for gaining insight into how AI crawlers are interacting with your pages.
From the sources
The same page lists pay per crawl among the monetization options it describes, and marks it explicitly as private beta.
A&A perspective
Once you can see the fetches, you will want to count them as a result. Keep them apart. A fetch count is how many times a bot came to collect something. Search impressions and clicks are a different number, and inquiries are a third. There are perfectly ordinary stretches where fetches rise and inquiries do not, and the reverse happens too. The cost-of-serving question and the discovery question belong in separate ledgers as well. The smaller the business, the stronger the pull to substitute a newly visible number for a result. While the monetization mechanism is not generally available, we think the sensible reading is that fetch counts are a cost-side indicator. That is A&A's own position, not a claim made by the sources.
Where this split does not apply, and the next step
From the sources
Both Cloudflare pages cited here are the vendor's own description of its own product. The page explaining bot classification shows a last-updated date of July 1, 2026, and the AI Crawl Control overview shows August 14, 2026; both were read on September 30, 2026. Neither is independent third-party verification.
A&A perspective
There are three situations where this split does not help. First, if your CDN is not Cloudflare, you do not have the same control surface. The idea of splitting by behaviour transfers; the specific settings do not. Second, if your site has no discovery surface at all — you publish a company outline and work arrives through referrals and existing customers — then almost nothing in this article applies, and closing everything is fine. Third, if regulation or a contract restricts reuse of the content, that constraint comes first and is not traded off against discovery.
A&A perspective
The next step is the inventory, not the settings screen. Split your domain's URLs into discovery surfaces and handover surfaces, and check first that the handover surfaces are behind authentication. Then, for the discovery surfaces only, allow search and agent fetching and decide training on commercial grounds. For where this decision sits in the flow from winning customers through to retention, see "AI-native GTM: a practical guide for solo founders and small teams". If you cannot tell whether the line you have drawn matches the shape of your business, that is the kind of question suited to a direct consultation.
Allowing or blocking AI crawlers is not a single switch. Split it by behaviour and by area of the site before deciding. What makes a blanket block dangerous is not that it is blunt but that it is asymmetric: robots.txt is only a request, so absorption into models is not reliably stopped, while a block at the CDN is enforcement, so the path by which people and AI answers find you closes for certain. Take inventory of your URLs first, close the handover surfaces with authentication, and decide the three behaviours separately for the discovery surfaces alone.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Bots (Cloudflare bot solutions documentation)
Cloudflare · last updated 2026-07-01
Accessed 2026-09-30 - AI Crawl Control overview
Cloudflare · last updated 2026-08-14
Accessed 2026-09-30 - AI features and your website
Google Search Central · last updated 2025-12-10
Accessed 2026-09-30
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-30