On 15 September 2026 Cloudflare — the service my sites’ traffic passes through on its way to a reader, and quite possibly yours; an earlier article shows how to tell — retired the control labelled Block AI bots and left three rows in its place: Search, Agent, Training. I have screenshots of that dialog from the 10th and from the 16th, and they are not of the same screen. What I did next is the part I could find nobody else published: I turned the new switches on and read what they wrote into my site.
Because that is what these switches do. They do not only turn programs away at the door. They write a list of names into robots.txt — the public file every site keeps at one address, where it says which programs it would rather not have reading it; it asks, it cannot refuse — and the screen never shows you the list.
The first thing to do takes a minute. Open your domain followed by /robots.txt. If the file
opens with # BEGIN Cloudflare Managed content — or # BEGIN Cloudflare Bot Preference Sync, the
name Cloudflare’s August announcement gives the same block — a row is publishing a list in your
name, and a heading inside, # Training crawlers or # AI Agents, says which. If it does not,
nothing is being published for you; my four sites read that way on the 16th.
And the question most readers arrive with. On the three rows, probably nothing: Cloudflare’s table of old settings to new maps a site whose old control was off to Allow on all three, and Allow is what my four read. Two things that table does not settle. If your old control was set to block, the table sends it to Disallow AI Training on Training and Block on pages with ads on Agent — but one of my own sites read Block only on pages with ads on the 10th and reads Allow on all three now, so I cannot tell you the table is what happened. And if you had the old robots.txt switch on, something did change whatever your rows say: what it published is gone, and no published table has a row for it. Both are below.
The control that was gone by the 16th
The dialog is headed Configure AI bot policies. I saw it on two of my sites on 10 September and on all four on the 16th.
| 10 September | 16 September | |
|---|---|---|
| The old control | 10 SeptemberBlock AI bots, labelled [Deprecating on September 15] — reading Block only on pages with ads on one site and Block on all pages on another, same account |
16 Septembergone from all four sites |
| The switch that writes to robots.txt | 10 SeptemberInstruct AI bots to not scrape content, on | 16 SeptemberEnable Bot Preference Sync, on |
| kovalseo.com/robots.txt | 10 SeptemberCloudflare’s block, nine names, above my own lines | 16 Septembermy own lines, nothing above |
Neither reading was mine. I had never opened that control on any site; the two settings were simply there when I first looked, opposite each other in one account, and the screen shows what a setting is, not who set it. The first blocked nothing — that site carries no ads. The second was shutting a site to every AI search bot I sent, so I set it to Do not block that day. The three rows were there on both dates and read the same on both, Allow (do not block); by the 16th they are all that is left, and beside each Cloudflare prints Recommended — on Training too. Below them sits the new switch, on: one line says it writes your settings into the top of your robots.txt.
Cloudflare announced this on the day, in a blog post and a press release. The post’s answer, under “What you need to do”: “Nothing, in almost every case. Your current settings carry over on their own.” What the post does not print — what nothing on the screen prints — is what goes into the file when a row is not on Allow. So I found out.
Thirty-two names under the Training row
On 17 September I set the Training row on kovalseo.com to Disallow AI Training — the setting Cloudflare’s post names for the purpose, “To stop training and keep search, use Disallow AI Training.” — and fetched my file minutes later.
The block was back, under the same marker, and it was not the block I had seen. On the 10th it
had asked nine names to stay out. Now it asks thirty-two — thirty-three User-agent: lines if you
count, one of them * — under a heading of its own, # Training crawlers. Eight of the old nine
are still there; the ninth, CloudflareBrowserRenderingCrawler, is gone. Twenty-four had never been
in my file:
Diffbot omgili anthropic-ai Claude-Web
cohere-ai MistralAI-Training GoogleOther Baiduspider
PetalBot AwarioSmartBot AwarioRssBot Google-CloudVertexBot
QualifiedBot Cotoyogi ICC-Crawler atlassian-bot
FishBot BorderxBot NavuBot SemrushBot-SWA
WARDBot magpie-crawler KimiBot CitibotSiteCrawler
Four names stand out to a site owner. GoogleOther — Google’s own page on its bots calls it “the generic crawler that may be used by various product teams for fetching publicly accessible content from sites”. PetalBot — a name I had met the day before on another of my sites, among the few Cloudflare showed answered with a page: seven in a day, alongside Googlebot and BingBot. Baiduspider and SemrushBot-SWA — names I read as a search engine’s and an SEO tool’s, and I say read because I found no page from Cloudflare, Baidu or Semrush to check them against, and will not say what they do on the strength of their names.
Cloudflare’s account of what a training block costs is two sentences. Its post says the Block settings “now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot”, which Disallow AI Training spares for search; then: “Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI — blocking those does not affect search.” The first holds: none of the three is in the file. The second is where the four names above ought to be and are not: they are none of those four companies’, and the post says nothing of them. Choose this row to keep your pages out of AI training and you are also asking whatever PetalBot and Baiduspider are to stay away, under a row that reads “Crawlers that crawl content to train AI models.”
Fifteen more under the Agent row
With Training still set, I moved the Agent row off Allow too, and fetched again. The file gained a
second heading, # AI Agents, with fifteen more names — forty-seven in all:
Amzn-User ChatGPT-User ChathiveCrawler Claude-User
FireCrawl FirecrawlAgent Google-Agent Google-GeminiNotebook
Google-NotebookLM Instapaper Kimi-User MistralAI-User
Perplexity-User Retool meta-externalfetcher
The screen describes the Agent row as “Bots that pull information from your site to provide
responses to user questions.” Six of the fifteen names end in -User, and for the three you are
likeliest to recognise, the maker’s own page says what the program is. OpenAI’s page on its
bots, of ChatGPT-User: “When users ask ChatGPT or a
CustomGPT a question, it may visit a web page with a ChatGPT-User agent.” Anthropic’s page says
the same of
Claude-User,
Perplexity’s of Perplexity-User.
Google files Google-Agent and its notebook
fetcher
among what it calls user-triggered fetchers, “initiated by users to perform a fetching function
within a Google product” — and prints Google-NotebookLM, which Cloudflare’s list carries beside
its replacement Google-GeminiNotebook, as “Former agent (supported until August 2026)”. Meta’s
page says
meta-externalfetcher
“fetches individual links at a user’s request”. So the user question in the row’s description
is, some of the time, your own visitor’s: a person asks an assistant to open your page, and this
row asks the assistant not to. What the wording does not carry is the list, and on it sit
Instapaper, Retool and FireCrawl. I recognise the first
two as products, and found no page from any of the three saying what a program of that name does
on a site, so I will not guess in print.
Four of the five makers say the line may not bind them
I wrote above that robots.txt asks and cannot refuse. For this row, the makers say so themselves. OpenAI, in the paragraph I quoted: “Because these actions are initiated by a user, robots.txt rules may not apply.” Perplexity, of Perplexity-User: “Since a user requested the fetch, this fetcher generally ignores robots.txt rules.” Google, of all its user-triggered fetchers: “Because the fetch was requested by a user, these fetchers generally ignore robots.txt rules.” Meta, of meta-externalfetcher: “Accordingly, this crawler may bypass robots.txt rules.”
Four of the five makers, then, publish in writing that a line in this file may not bind the fetch a person asked for. The fifth, Anthropic, claims no such exemption: “Anthropic’s Bots respect ‘do not crawl’ signals by honoring industry standard directives in robots.txt”, and disabling Claude-User “prevents our system from retrieving your content in response to a user query”.
That is why the list matters when the file only asks. The row does two things, and I read one: Cloudflare’s documentation says a block turns the programs it recognises away at the door, and I did not probe the door. The file, the part published in your name, says less than it appears to: for Claude-User, by its maker’s word, the line is the refusal; for the other four, a request their makers have said in advance they may not honour.
The limits. One site, one plan, one fetch of each state; the Agent row was set while Training was still on, so I credit the second heading to Agent because it arrived when that row changed, not because I tested the rows apart. Both rows went back to Allow the same day, and the file to my own lines. Whether the list is the same on your site, your plan, or next month, I cannot say; on mine, that day, these were the names.
Read your own file
The same address as at the top. Everything between either # BEGIN Cloudflare line and its
# END is what a row currently publishes. After any change to a row, fetch the file again rather
than trusting the screen: on mine the block appeared within minutes, and it is the file that
programs read.
The file is where you read a block, never where you change it: those lines are not in your own robots.txt to edit — Cloudflare puts them in front of whatever your site serves, from the row’s setting. A list you did not mean to publish: set the row back to Allow, and fetch the file again.
Then the rows: Security → Settings → Bot traffic → Configure AI bot policies. Read the sentence under each; it says more than the label. If the login is not yours, this is the message:
Copy and send
Subject: Cloudflare’s AI bot settings for my site
Please open Cloudflare for my site — Security → Settings → Bot traffic → Configure AI bot policies — and tell me what the Search, Agent and Training rows each read, and what the sentence under each row says. Please change nothing yet.
Then open my robots.txt. If it holds a block between a line beginning “# BEGIN Cloudflare” and the matching “# END Cloudflare” line, paste me the whole block, headings included. I will decide what to do once I have read it.
The documentation, the day after
Cloudflare announced the change; its developer documentation — the pages that explain each setting — did not move. I had read them on 12 September and fetched them again on the 16th, between 12:22 and 12:30 UTC: not one had changed. The page for the old control still reads “Last updated Jul 1, 2026”, still describes 15 September in the future tense, and still sends the reader to “Security Settings > Block AI bots”, a screen that by then existed on none of my four sites. That can change the day Cloudflare edits a page; but for that day at least, a switch live on all four of my dashboards was on no documentation page a person would open to look it up.
One example there disproves itself. The page for the old robots.txt
feature,
“Last updated Aug 3, 2026”, still teaches it as current and says that with the feature on,
Cloudflare prepends its content to yours, “resulting in what you can view at”
https://www.crawlstop.com/robots.txt. I opened that address on the 16th. It served:
User-agent: *
Disallow: /lp
Disallow: /feedback
Disallow: /langtest
Sitemap: https://www.crawlstop.com/sitemap.xml
That is, line for line, the page’s own “Feature not enabled” example (the served file adds one blank line). Four days earlier the same address had served a Cloudflare block of nine names. The page sends you to a link to see the feature on; the link shows it off.
And one setting left without a row. On the 10th, kovalseo.com had the old robots.txt
switch on, and the file said so: nine names asked out, and a line reading ai-train=no. On the
16th the switch no longer exists, Enable Bot Preference Sync is on in its place, the Training
row reads Allow, and the file carries nothing of what it carried. Cloudflare’s
post says “Customers who enabled
Managed Robots.txt will migrate to the new system” and prints two tables of what migrates to what;
neither has a row for this switch. What became of my selection, nothing published says.
Seven predictions, and what the screens say
Seven pages predicted this change before it came, the newest revised on 8 September; everything published since the date restates the announcement, and nobody I could find had looked at a screen or a file. Wise Media alone said the announcement “does not state that untouched existing zones get new blocks”, and that is what my four sites show: every row on Allow, no block in any file, free plan among them.
The same page’s headline is the prediction four of the seven made:
Wise Media’s headline, 24 August
If You Ever Clicked “Block AI Bots” on Cloudflare, September 15 Will Deindex You From Google
To deindex a site is to drop its pages from Google’s results, the reason many readers of that headline are here. No site of mine was left blocking on the old control, so I cannot tell you whether it happened. What Cloudflare published: its migration table sends an old Block not to the new Block but to Disallow AI Training, the setting built to keep Googlebot reading. Unconfirmed by anything I saw, and contradicted by the table.
Two threads on Cloudflare’s own community forum come closest. My tools cannot open that forum, so I opened both in a browser. One, posted 23 August, reports a managed block still served with the feature switched off and every row on Allow; the other, 31 August, the old block still in the file twenty minutes after the old control was set to Do not block. Eighty-one views and thirty-three; each closed automatically after fifteen days, with no reply visible in either. Both predate 15 September, so neither reports the change — only its neighbour, a block that stayed after the setting told it to go.
Allow is the reading, and a block is a list
All four of my sites read Allow on all three rows, which Cloudflare marks Recommended. Whether you should block is not my case to make here — the previous article measured that on one site. What this one adds is missing from every page I read: a block on either row is not a switch. It is a list — thirty-two names on Training, forty-seven with Agent added — written into a public file in your name by a service whose documentation, the day after, did not yet describe the switch; the screen never shows the list; and four of the five makers I checked on the Agent half of it have said in advance that it may not bind them. If you reach for the switch anyway, reach for the file next. It is one address, and now you know what to look for.
The part worth remembering
The rows are worth opening, and the sentence under each is worth reading. Allow is what my four sites read and what Cloudflare recommends on every one. If you set a block anyway, know what you are doing: not flipping a switch but publishing a list of names in a public file — one the screen never shows, and one the documentation, the day after, did not yet describe. Read the file. It is one address, and the list is the whole of the setting.
Koval SEO Console reads the public pages of your site — the same ones Google sees — and explains each finding in plain words, with the evidence beside it, then answers fixed, not fixed, or couldn't confirm when you check again.
Check my site