One of the rows in Search Console’s Why pages aren’t indexed table reads Blocked due to access forbidden (403). Under it sits the report’s one-size line — the same sentence it hangs under every reason, so ignore it. What is worth reading is Google’s definition of this reason, because it makes an assumption about you.
“Forbidden” is a hard word, and a number in brackets looks like a fault code. Google’s own help page for the report does not soften it. This is the opening of its definition, uncut — where user agent is Google’s word for whatever program is asking for the page, Googlebot is the name of Google’s own reading program, and credentials means a login and password:
Google’s help page, defining the reason
HTTP 403 means that the user agent provided credentials, but was not granted access. However, Googlebot never provides credentials, so your server is returning this error incorrectly. The page will not be indexed.
Incorrectly. Read that sentence once more, slowly, because it rests on a quiet assumption: that the page was meant to be open to everyone, so refusing Google must be a mistake. Sometimes it is exactly that — a page you want found, and something on your site is turning Google away. And sometimes the page was never meant to be open — a practice copy of your site, an admin door, a members-only area — and the refusal is the whole point. Two very different readers arrive at this row. The job of this article is to tell you, fast, which one you are, and then give each of you a path. It is one question: is the listed address one a customer should be able to open? If yes, something is refusing Google on a page that should welcome it, and If it’s a page you want found is your section. If it is a practice copy, an admin door or a members-only page, nothing is broken, and your path starts at If it’s something you locked on purpose.
What a 403 actually is
Every time a browser asks your site for a page — and Google’s reading program asks the same way — the site answers with a three-digit code that says how it went. 200 means here it is. 404 means no such page. 403 means something different from both: the request was understood, the page may well exist, and the answer is no. Not lost. Not broken. Refused. Something in front of your pages — the server itself, a security plugin, a firewall — looked at who was asking and turned them away.
Its nearest neighbour in the same report is “Blocked due to unauthorized request (401)”, and the contrast is useful. The same help page defines that one as “The page was blocked to Googlebot by a request for authorization (401 response)” — a locked door that asks for a key, the login box you have met on a hundred sites. A 403 is a door that has already decided. No key is asked for; the answer is no.
One of Google’s own analysts does not treat the number as an alarm. On the Search Central blog, Gary Illyes counts “forbidden” among the codes in this family he calls “pretty benign” and adds: “They don’t suggest anything wrong going on with the server itself.” What Google does with a 403 is written down too, in its page on HTTP status codes:
Google’s documentation, on every code in the 400s
In the case of Google Search, Google doesn’t index URLs that return a 4xx status code, and URLs that are already indexed and return a 4xx status code are removed from the index.
So the row’s consequence is exactly what it says: this address is not listed, and if it was, it is being taken out. Whether that is bad news depends entirely on which address.
One question sorts you
Open the row and read the listed addresses. Ask that one question about each: is this an address a customer should be able to open?
- Fix this A page you want found — your home page, services, products, an article. Nothing on your site should be refusing Google here. Real, and worth today.
- Fix this Every page at once — the same fault, with a hint attached: it points at one rule sitting in front of the whole site, not at the pages themselves.
-
Leave it
A practice copy of your site — an address with
dev.,staging.ortest.in front of your domain name, where the next version gets built. Locked on purpose. - Leave it Admin and login addresses — pages whose whole purpose is to refuse anyone who is not you.
- Leave it Members-only pages — content that is supposed to open only for people who are signed in.
The first two kinds: the next section is yours. The last three: nothing is broken — go straight to If it’s something you locked on purpose, because the standard advice for this row is the one thing you should not follow. If it is a mix, take the real pages through the next section and let the locked ones sit.
If it’s a page you want found
Owners who bring this row to Google’s help forum very often open with the same observation, and one thread from 2025 puts it in eleven words:
An owner, as reported in Google’s help forum
The page is accessible when I open it in a browser.
That is the signature of the fault, and it is the first thing to check. Open the listed address in a private window of your browser — Google’s definition of the 401 sibling suggests exactly this: “You can verify this error by visiting the page in incognito mode.” If it opens for you while Search Console reports a refusal, the block is aimed at Google, and the rest of this section is yours. If it does not open for you either, this is not a rule aimed at Google — your site is refusing everyone, and that is a same-day call to your host, not a Search Console question. The message further down this section, browsers get in fine, would not be true for you; do not send it.
Your site lets people in and refuses Google, so the refusal is aimed — something is looking at who is asking. In the reports owners have posted, that something is usually one of three: a security plugin, a firewall — Cloudflare, a service many sites put in front of their pages to decide who gets in, comes up often — or a setting at your host. One owner traced it to Cloudflare’s firewall: “I had blocked bots in the Firewall section and deleted what I had setup.” Another to a security filter on the host: “So the problem was that Mod security filter blocked different IP’s including Google Bot.” A Product Expert in the same forum — one of the volunteer regulars Google badges there, not Google staff — summed it up for an owner who had asked to be told plainly what to do: “The issue was that something at your end (so your server, your firewall) was refusing the googlebot access.”
Google’s advice sits in the second half of the help page’s definition. It is conditional — note the opening if:
Google’s help page, on the remedy
If you do want Googlebot to index this page, you should either admitting non-signed-in users or explicitly allow Googlebot requests without authentication (though you should verify its identity).
In plain words: let Googlebot in — but make sure it is Googlebot. That last bracket is doing more work than it looks. Anyone can knock on your site claiming to be Google; the name a request gives is just text. Google’s Googlebot page says so:
Google’s Googlebot documentation
Before you decide to block Googlebot, be aware that the HTTP user-agent request header used by Googlebot is often spoofed by other crawlers. It’s important to verify that a problematic request actually comes from Google.
Which is why the fix is neither “let in anyone who says Google” nor “switch the protection off”. In one forum thread where turning Cloudflare off had been suggested, a later reply put it in capitals: “Disabling cloudflare is a VERY BAD IDEA.” The fix is narrower: find the rule that is catching Google, and let the real Googlebot through — Google publishes how to check that a request “really is from Google”. Where to look: in Cloudflare, that is the Events tab of its security Analytics page — Cloudflare’s own guide to it says it lets you “review mitigated requests”, which in plain words is a list of what it blocked and why; a security plugin usually has a blocked-requests or live-traffic log that answers the same question. If a Google address or the name Googlebot appears in that log, you have found the rule. Finding it and fixing it is a job for whoever runs your security plugin, your Cloudflare account or your hosting. If that is not you, the message to send is short: Googlebot — Google’s own program — gets a 403 from my site; browsers get in fine; please find the rule refusing it.
The second test is Search Console’s own, and it is how you know the fix took. Paste the address into URL Inspection at the top and press TEST LIVE URL. That reads the page as Google, this second, not as Google last saw it. Once the rule is fixed, the Page fetch line in the result reads Successful. One reported wrinkle, if the live test still fails after your fix: an owner found their Cloudflare rule was catching Google’s inspection tool specifically, and had to add an exception for the name “Google-InspectionTool”. That is your whole path. The next two sections are the other reader’s — jump to How to know it’s settled.
If it’s something you locked on purpose
Now my own row — the version that Google’s definition of this reason never imagined.
On one of my sites, the Page indexing report carries this reason with 1 affected page. The
issue page says First detected: 7/25/26, and its Examples table lists a single address —
a dev. subdomain of that site, next to a Last crawled date of Jul 18, 2026, the day
Google last read it. I am not going to print the address here, for a reason this article is about
to make obvious.
That subdomain was a practice copy, and by then nobody was using it. It sits behind Cloudflare, a service that decides who gets in, and answers a plain visit with a 403; it did so when I checked on 17 August 2026, and it is supposed to. Nothing on the main site links to it. Nothing is broken — the lock is doing its job. And there it sits, in a table headed “Why pages aren’t indexed”, with Google’s help page telling me my server is returning the error “incorrectly”.
Which leaves the question I asked, and the one the usual guides to this row leave alone: how did Google find it?
How the name became public record
Google does not say. Its own description of how it finds addresses names three sources — pages it already knows, links from those pages, and sitemaps you hand it. The first two need a link, and I had published none; the third needs you to give Google a list, and a locked practice copy is the last thing anyone lists. Its URL Inspection help admits there are others: a page can be flagged “URL might be known from other sources that are currently not reported”. And its Googlebot page is blunt: “It’s almost impossible to keep a site secret by not publishing links to it.”
So I went looking for how the name became public, and found where it had — somewhere that has
nothing to do with Google’s search. When the practice copy went up, it got a certificate — the
thing behind the padlock in your browser’s address bar, the reason a site is https:// and not
http://. A certificate names the addresses it covers. And certificates, it turns out, are public
by design.
crt.sh is a free search over the public certificate records. I searched it for
%. followed by that site’s domain — the % stands for anything in front of — and the table came back
with four distinct names across its rows. Two were expected: the domain itself and its www.
twin, first certificate in June 2023. The third was the practice copy — a certificate naming that
exact subdomain — first issued on 4 July 2026. Two weeks before Google’s date of last reading.
Three weeks before the row appeared. (The fourth I will come back to.)
Here is why that record exists, in plain words. Every padlock issued since spring 2018 has been written into a public ledger anyone can search — Chrome’s own page for site operators puts the rule as “all certificates issued after 30 April 2018 are expected to be disclosed via Certificate Transparency”, and spells out what disclosure means: “the domains a certificate are for will be included in the Certificate Transparency log”. Certificate Transparency is the name of the ledger system, and its own site describes the ledgers this way — append-only meaning entries can be added but never taken out:
The Certificate Transparency project, on its logs
Certificate logs are append-only ledgers of certificates. Because they’re distributed and independent, anyone can query them to see what certificates have been included and when.
So the moment that practice copy got its padlock, its name was in a public, permanent record that anyone can search. That is not a leak or a bug; it is the design, and security testers use it on purpose. The OWASP testing guide — a handbook for people who test websites for weaknesses — gives the same crt.sh search as a standard step and says what it turns up:
The OWASP Web Security Testing Guide, on searching certificate logs
The results may list subdomains such as dev.example.com, staging.example.com, or other hostnames that are not directly referenced from the primary site.
Now the careful part. Google does not say how it found the address, and I am not going to claim it read these ledgers — I don’t know that. What I can say is narrower, and enough: from 4 July the name was public record; on 18 July Google read the address; on 25 July Search Console marked it first detected. No link from the site was needed, and I had published none.
Which is also why the address is not printed above. A name in those ledgers is already public, so withholding it protects nothing — but there is no reason to hand it to a reader, and an article that tells you your practice copies are exposed should not go advertising one. Deciding not to print it also pointed at the option I had not considered: I was done with that copy. A practice copy you have finished with does not need a better lock. It needs taking down.
The fourth name in the table is the practical lesson, and it is something to tell whoever sets up
your next practice copy. Two days after the named certificate, another was issued for *. plus
the domain — a wildcard, one certificate covering any name in front of the domain without
listing them, which Chrome’s
page says “can
limit how much of the domain is disclosed”. Had the practice copy only ever been covered by the
wildcard, its name would not be in any ledger; it was the certificate naming it that published it.
The next copy can be covered without being named.
The wrong fix, named
The standard advice for this row is “allow Googlebot”. For a page you want found, that is right. For a practice copy, it is the one move that makes things worse: you would be publishing a half-built site. The row is not asking you to open the door. It is only telling you Google knows where the door is.
Two right moves, and Google’s own help forum has both on record. The first is to leave the 403 alone. Asked how to get a staging site out of Google, one Product Expert answered this:
A Product Expert in Google’s help forum, as reported
Make the staging site return a 4xx status (403 is ideal!) Currently it returns a 500.
Another, answering an owner whose staging site returned 403 to everyone but a few allowed addresses, said the same: if Google meets a clear-cut 403 when it tries the address, it will drop the page from its list, and there is no need to let it in. The follow-up in the same thread is the detail people miss: Google should be allowed to reach the address and see the 403, because seeing it is what makes Google drop the page — and that takes time. Don’t hide the locked address behind a “keep out” note in robots.txt — let Google reach the door and read the refusal. Google’s own FAQ, higher up the same help page that called the 403 incorrect, says the same about robots.txt, and names the second right move:
Google’s help page, FAQ: “Why is my page in the index? I don’t want it indexed.”
If you want your page to be blocked from Google Search results, you can either require some kind of login for the page, or you can use a noindex directive on the page. Using a robots.txt rule is not recommended for blocking a page, and will actually prevent noindex from being seen by Google.
The words to hold on to are “require some kind of login.” Password-protect the practice copy, and Google and strangers meet the same closed door. The second Product Expert above recommends exactly that as often the simplest route for a staging site: put a password on it, and programs like Google’s cannot get in, because they do not know it. If yours already answers 403 to everyone, you have the first move and are done. The second is for a practice copy that is still open.
Notice the gentle contradiction, because it is the whole reason this row confuses people. In one place on Google’s page, the definition says your server is returning the 403 incorrectly. Elsewhere on the same page, the FAQ blesses a login wall as a proper way to keep a page out of Search. Both are true. The definition assumes the page was meant to be public; the FAQ is written for the case where it wasn’t. Only you know which page you have — and that is the one question this article asked you at the start.
How to know it’s settled
If it was a fault, the proof is the live test: URL Inspection, TEST LIVE URL, and a Page fetch that reads Successful where the rule used to refuse. Then, on the issue’s page, the strip that asks Done fixing? offers VALIDATE FIX; press it, and Google starts re-reading the affected pages at its own pace, which can run for weeks. The live test is your proof, today. The row itself clears when Google’s own reading program next tries the address and is let in — and, as the lock case below shows, Google’s help page says it will probably keep trying an address that turned it away, so you are not waiting for a visit that might never come, only for one at Google’s pace.
If it was a lock, there is nothing to validate, because there is nothing to fix. A 403 is the state you want. So don’t press VALIDATE FIX on this row — it would ask Google to confirm the door is open, and it isn’t; the answer would come back Failed for a lock that did its job. The row will keep showing the address for a while. Google says the same about addresses that keep answering with an error, under the neighbouring “Not found (404)” reason: it will probably keep trying, less and less often, and there is no way to make it forget.
So a row that stays put is not a lock that failed — and it is not the same kind of leftover as a redirect row, either. A redirect row can be a memory: an address Google filed once and may not have looked at since. This row is the opposite. Google says it will probably keep coming back to a refused address, less and less often, and each visit meets the same no. Put that beside the rule on its status-code page — an address that was ever listed and now returns a 4xx is taken out — and the row stops looking like a stale record. It is the receipt: the address is not in Google’s list, and every knock the lock turns away is what you built.
And one action if you have — or ever had — a practice copy, or a developer set things up for you:
check your own domain the way I checked mine. Open crt.sh — a free service that
is sometimes down for a while; if it is, try again later — type %. followed by
your domain name into the search box, and read the Matching Identities column. Expect many
rows — each renewal is one, and most repeat. You are reading for distinct names, not counting
rows; mine were four names spread over more than a dozen lines. Every name in front of your domain
that has been given a padlock of its own since 2018 should be there — including the ones you set
up years ago and the ones a developer made without mentioning it. If a name surprises you, you
have just found out that it is public record, and you can decide today whether that address should
still be up, and whether it is locked. The name itself is already public and stays so — the
wildcard lesson is for the next copy you build. That is more than this Search Console row will
ever tell you.
The row itself was describing something real: your site refused Google. It just never asked whether the refusal was meant. That was always your question, and now you can answer it.
The part worth remembering
"Forbidden" is a decision, not a breakage. The report is describing something real — your site refused Google — but only you know whether it was meant to. On a page customers should open, find the rule and let the real Googlebot through. On a page you locked, keep the lock: the row will linger, because Google keeps trying addresses it knows, and knowing an address was never the same as being let in.
Koval SEO Console reads the public pages of your site — the same ones Google sees — and explains each finding in plain words, with the evidence it read beside it — then answers fixed, not fixed, or couldn't confirm when you check again.
Check my site