Open your robots.txt — your domain, then /robots.txt — and it reads like a rule. Every line is an
order: User-agent, then a name; Disallow, then a place. It sits at an address every program
knows, and it is public. So the belief that it keeps programs out is not foolish; it is what the
file looks like.
I went looking for anything that makes a program obey it. The standard that defines the file says nothing does. The companies whose programs read it say so on their own pages — the one that sells bot-blocking among them. My own servers showed me what does turn one away, and it was never the file. And the one court ruling now quoted on the question, from a United States federal district court in New York, says the file is not a lock — about one claim, under one American law, and nothing else. That last part is the half the quoting leaves out, and why this article reports and does not advise: what the documents say, I can show you; what the law of your country makes of a program that read past your file is a question for a lawyer there.
The standard says it in its introduction
The file is defined by RFC 9309, an internet standard of September 2022. Its introduction says what the rules are, and the next sentence what they are not:
RFC 9309, the standard that defines robots.txt, in its introduction — crawlers are the programs that read websites, URIs the addresses they read
This document specifies the rules originally defined by the “Robots Exclusion Protocol” [ROBOTSTXT] that crawlers are requested to honor when accessing URIs.
These rules are not a form of access authorization.
Requested to honor — in the sentence that defines the rules. The standard’s security section adds a consequence for the owner:
RFC 9309, under Security Considerations
The Robots Exclusion Protocol is not a substitute for valid content security measures. Listing paths in the robots.txt file exposes them publicly and thus makes the paths discoverable.
The companies say it on theirs
Google, on its page introducing robots.txt:
Google’s introduction to robots.txt
The instructions in robots.txt files cannot enforce crawler behavior to your site; it’s up to the crawler to obey them.
Cloudflare, which sells a product that turns programs away at the door, says in the blog post announcing a feature that writes into its customers’ robots.txt:
Cloudflare’s blog, July 2025
But robots.txt is an honor system. Nothing forces bots to follow it.
So whether a program obeys is a promise each company makes, or does not, on its own page. Anthropic’s page on its bots says they “respect ‘do not crawl’ signals by honoring industry standard directives in robots.txt”. Google’s page on its bots says of its main ones, “They always obey robots.txt rules when crawling automatically.” OpenAI’s page describing its bots contains no such sentence.
One official document I could find has the companies commit to the file, and it is European: the Copyright Chapter of the EU’s Code of Practice for general-purpose AI models, July 2025, whose signatories — OpenAI, Google and Anthropic among them — commit “to employ web-crawlers that read and follow instructions expressed in accordance with the Robot Exclusion Protocol (robots.txt)”. The Commission’s own page for the code calls it “a voluntary tool”.
What turned a program away on my sites was never the file
Before the court, my own servers. On 16 September 2026 I asked each of three sites of mine for a page that does not exist, under the name GPTBot — OpenAI’s program that collects pages for training. The day before, the block Cloudflare had been writing into my files had gone, so they said nothing about GPTBot. All three answered 403, refused, on my host’s refusal page, not Cloudflare’s — on a refusal I kept rather than set: on 10 September I opened my hosting company’s own list of refused programs for the first time, found GPTBot already on it, and left it there. The file did not say so. The door did.
Two answers of the opposite kind. On kovalseo.com on 10 September the file told ClaudeBot, Anthropic’s training program, to stay out, in lines Cloudflare had written above mine, and the host let it in — that one is in the article of 14 September. And on kovalweb.com, before 10 September, the file said nothing about GPTBot and the host refused it, on a setting nobody had chosen. Neither party broke a rule written in that file, because the file writes none. Every answer came from a door, and no door on my sites reads the file: a program that reads it stops itself, or does not.
You can tell which of the two answered on your site. Have whoever holds your hosting login ask the
site for a page that does not exist, giving a bot’s name, and read the code that comes back. Page
not found means the request got in: no door stopped that name, and what your file asks of it is a
request that program alone will honour or not. 403 means a door refused the name before any page
was looked for, and your file had no part in it. On ordinary hosting that door is a list your
hosting company may keep of programs it refuses by name — a setting in its panel, not a line in
your file, possibly set before you ever looked. Which door refused can sometimes be read off the
refusal itself: on my sites, while the service in front of them was the one turning names away,
the page came back saying Your request was blocked., and once that service stopped, the same
request was refused on my host’s own page instead. That is one company’s wording, not a rule. If
the refusal you get names a company, ask that company first; otherwise the person to ask is
whoever you pay for hosting, and the 14 September article has the message to send.
One limit: a request wearing GPTBot’s name is not GPTBot, and a door that checks where a visitor comes from could answer OpenAI’s own program differently. The test settles only which of the two, file or door, answered.
What the court said, in plain words first
The part of the American law the file was tested against is about breaking locks. The judge said a note is not a lock. About the file, that is all he said.
The rest is paperwork, here because the sentence going round — a court said ignoring robots.txt is legal — is not one the judge wrote. Two of the paperwork’s words first. A claim is one separately pleaded count in a lawsuit, standing or falling on its own. A technological measure and its circumvention are the law’s words for a lock and for breaking past it, not for walking past a sign.
In December 2025 a United States federal district court in New York — one judge, Sidney H. Stein — ruled on a motion in a case brought by Ziff Davis, the publisher of CNET, PCMag, Mashable and IGN, against OpenAI; the case is In re OpenAI, Inc. Copyright Infringement Litigation, and Ziff Davis’s suit is one member of it. Ziff Davis had pleaded that its robots.txt files told OpenAI’s GPTBot to stay out, that GPTBot came in anyway, and that this was “circumvention” of a “technological measure” under an American copyright law, the Digital Millennium Copyright Act. The opinion of 15 December 2025 — Document 300 in case 1:25-cv-04315, read from the filed PDF — answers in the sentence now quoted everywhere:
The opinion of 15 December 2025 — a United States federal district court in New York
Robots.txt files instructing web crawlers to refrain from scraping certain content do not “effectively control” access to that content any more than a sign requesting that visitors “keep off the grass” effectively controls access to a lawn.
And the second ground, on the same page — the whole sentence, because the quoting drops its last clause:
The same opinion, a paragraph later
At most, Ziff Davis alleges that OpenAI disregarded the instructions that were contained in robots.txt files. This is not “circumvention” under the DMCA, and Ziff Davis’s claim fails for this reason as well.
Ziff Davis’s claim. One claim. OpenAI had asked the court to throw out six claims; the robots.txt sentence decided one, under one section of one American law. How each claim stood after this motion, from the opinion’s conclusion and the list it was aimed at — the names are the law’s, not mine; the right-hand column is the part that matters:
| Claim | On this motion |
|---|---|
| Circumvention under the DMCA — the robots.txt claim | dismissed |
| Unjust enrichment | dismissed |
| Trademark dilution | dismissed in part, kept in part |
| Contributory copyright infringement | went forward |
| Removing copyright management information | went forward |
| Distributing works with that information removed | went forward |
| Copyright infringement by training | not before the court on this motion |
| Copyright infringement by the model’s outputs | not before the court on this motion |
| Dilution and injury to business reputation under Delaware state law | not before the court on this motion |
The stage matters too: a motion to dismiss, where the judge assumes everything the complaint says is true — “All factual allegations in the First Amended Complaint (“FAC”) are assumed to be true for the purposes of OpenAI’s motion to dismiss” — and asks only whether, even so, the claim can stand. No evidence had been heard. Whether reading past the file was lawful under the case’s other claims — training and outputs among them, never before the judge that day — was left open, and remains so. One judge, one district, binding no other court; as of 24 September 2026 no appeal appears on the docket, and the case is still before the same court. An earlier federal district court, in Pennsylvania in 2007, had gone the other way on the narrow point — and said, in words the 2025 opinion quotes, that its finding “should not be interpreted as a finding that a robots.txt file universally qualifies as a technological measure that controls access to copyrighted works under the DMCA.”
Three days later, the same judge again
Ziff Davis asked to try again — a new complaint with more on how OpenAI had treated the file. On 18 December 2025 the same judge refused, in a four-page order — Document 983 in the multidistrict docket, entered as Document 302 in Ziff Davis’s own case: the same order, docketed twice — because nothing in it would change the answer. Its reasoning is the plainest account of the file I read anywhere; the PSAC is the proposed new complaint:
The order of 18 December 2025 — the same United States federal district court
On the PSAC’s telling, robots.txt files prevent a bot from accessing Ziff Davis’s works only if the creator of the bot affirmatively elects to configure the bot to read and comply with robots.txt files before the bot scrapes the website. If the bot creator does not take this affirmative step of choosing to heed the requests embodied in Ziff Davis’s robots.txt files, the robots.txt files have no effect on the bot’s ability to access Ziff Davis’s works. The allegation that all “reputable” bot operators configure their bots to comply with robots.txt files by default (id. ¶ 122) does not change the precatory nature of the requests embodied in robots.txt files.
The argument refused is the one every site owner would make — that reputable companies configure their bots to read the file before they crawl. Yes, said the court; that is the point. A file that works only when the other side chooses is precatory, a lawyer’s word for a request. The same one-claim limit holds: the order says what the file is not under that one law, and nothing more.
The losing side’s own file
Ziff Davis’s robots.txt files today carry something the court never saw. On 24 September 2026 I fetched cnet.com/robots.txt — the same header sits on Mashable’s and IGN’s; PCMag’s would not open for me — and above its rules, this:
# Ziff Davis content is made available for your non-commercial use subject to our
# Terms of Use here: https://www.ziffdavis.com/terms-of-use
# Use of any robot, crawler, or other tool to scrape, harvest, extract, or retrieve any content on
# this website using automated means is prohibited without written permission from Ziff Davis.
# Prohibited uses include but are not limited to:
# (1) text and data mining under Art. 4 of the EU Directive on Copyright in the Digital Single
# Market;
Terms of use, a prohibition, and a reference to a European directive, as comments inside the file an American court had, nine months earlier, likened to a keep-off-the-grass sign. When those lines were added I could not establish, and they are not what the court read: the complaint described directives, not comments.
The directive they name is the EU’s Directive 2019/790, whose Article 4 lets anyone make copies of lawfully accessible works for “text and data mining” — its term for programs reading text in bulk to find patterns — on one condition, in Article 4(3):
Directive (EU) 2019/790, Article 4(3)
The exception or limitation provided for in paragraph 1 shall apply on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online.
That is all Article 4(3) says about how: one example, “machine-readable means”. The directive as a whole names no file, format or protocol; the word robots.txt does not occur in it. Whether a comment block, or any line in any file, is “an appropriate manner”, nothing I read says, and I cannot.
Cloudflare’s own files carry one too
Until 15 September 2026 Cloudflare wrote a paragraph of that kind into its customers’ robots.txt files, mine among them; on the 16th it was gone from mine. Two of Cloudflare’s own properties still serve it. developers.cloudflare.com/robots.txt, fetched 24 September 2026, opens “As a condition of accessing this website, you agree to abide by the following content signals:”, defines three signals, and then:
# ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS
# RESERVATIONS OF RIGHTS UNDER ARTICLE 4 OF THE EUROPEAN
# UNION DIRECTIVE 2019/790 ON COPYRIGHT AND RELATED RIGHTS
# IN THE DIGITAL SINGLE MARKET.
User-agent: *
Content-Signal: ai-train=yes, search=yes, ai-input=yes
The paragraph says any restriction the signals express is a reservation of rights under European law; the line under it restricts nothing — yes to training, search and answers. Its opening line is contract language, in a file whose own standard calls its contents requests. Nothing I read says any court, regulator or company has given that sentence effect — not that it has none; only that nobody I can cite has said it has any.
What is left for your file
It is worth having: Google’s main crawlers and Anthropic’s bots say in writing that they read it, and a request those visitors say they honour is not nothing. It is worth reading, for two reasons. Every path you list in it is public, as the standard warns. And more than one party can write into it: a service in front of my sites has, my host has a switch that would, and the earlier article on a request Google obeyed shows a file I inherited with a line I never chose.
What it cannot do is make anyone obey — by the standard’s word, by Google’s, by Cloudflare’s, and by the American court’s on that one question. The things that are not requests live elsewhere: the door that answered 403 on my sites, a password, a paywall (the EU code above keeps paywalls under a different commitment in the same chapter as robots.txt: measures not to be circumvented, not reservations of rights). Whether any law where you are makes a company answer for reading past your file, I cannot say. That question goes to a lawyer in your country — I am not one — with the file, and the ruling’s actual reach, in hand.
The part worth remembering
The file asks. The standard says so, the companies say so, and one American court said so about one claim — and about nothing else. Read yours for what it is: a public list of requests, kept by parties who may not all be you, that a well-behaved program honours and any other program can read straight past. The things that are not requests do not live in that file.
Koval SEO Console reads the public pages of your site — the same ones Google sees — and explains each finding in plain words, with the evidence beside it, then answers fixed, not fixed, or couldn't confirm when you check again.
Check my site