Security & devFreeNo signup
Blocking a URL in robots.txt does not keep it out of Google. Google's own documentation is explicit: it can't index the content of pages which are disallowed for crawling, but it may still index the URL and show it in search results without a snippet
. The file above controls crawling. Nothing in it controls indexing, and the two are separate systems that the same document is routinely expected to do both jobs for.
Last updated 2 October 2026
If empty, rules apply to User-agent: *
What this file can and cannot do. robots.txt asks well-behaved crawlers not to fetch certain paths. It is not a security control, it is not a way to hide a page from search results, and it is published at a public URL that anyone can read. RFC 9309 states the first point in one sentence: These rules are not a form of access authorization.
Anything that must not be reached needs authentication, not a Disallow line.
Disallow controls crawling. It has no effect on indexing. Google's robots.txt documentation says it in one sentence: Google can't index the content of pages which are disallowed for crawling, but it may still index the URL and show it in search results without a snippet
.
That is how a URL you deliberately blocked turns up in search results as a bare link with no description. Google never fetched the page. It found the address somewhere else, usually a link from another site, and the blocked page is a page it is allowed to list but not to read.
The fix people reach for next is a noindex rule, and it collides with the first one. Google: For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.
A crawler that is forbidden to fetch the page cannot see the noindex instruction inside it. Blocking the page is what prevents the deindexing from working.
So the two jobs need two different tools, and they cannot be stacked:
noindex on the page. Let the crawler in so it can read the instruction to leave.The Robots Exclusion Protocol spent twenty-five years as an informal convention that everyone implemented slightly differently. It stopped being one in September 2022, when it was published as RFC 9309 on the IETF Standards Track. Most pages about robots.txt still describe it as a convention or a gentlemen's agreement. It has MUST and SHOULD requirements and an ABNF grammar.
What follows is what the standard requires, checked against the output this page actually produces.
This is the single most useful clause to know, because it is the opposite of how firewall rules and .htaccess files behave. RFC 9309: The most specific match found MUST be used. The most specific match is the match that has the most octets. Duplicate rules in a group MAY be deduplicated. If an "allow" rule and a "disallow" rule are equivalent, then the "allow" rule SHOULD be used.
The preset above produces exactly this shape when you add Disallow paths to Allow all crawlers:
User-agent: * Allow: / Disallow: /admin/ Disallow: /cart/
It looks self-contradictory and it is not. For the URL /admin/settings, Allow: / matches one octet and Disallow: /admin/ matches seven, so the disallow wins and the path is blocked. For /about, only Allow: / matches, so it is crawled. Moving the Allow line to the bottom would change nothing. Google describes the same behaviour and adds the tie-break from the other direction: When matching robots.txt rules to URLs, crawlers use the most specific rule based on the length of the rule path. In case of conflicting rules ... Google uses the least restrictive rule.
One consequence worth holding on to: a longer Disallow cannot be overridden by a shorter Allow, so blocking a folder but allowing one file inside it only works if the Allow line is the longer string. Disallow: /assets/ with Allow: /assets/logo.png works, because the second is longer. The reverse never does.
RFC 9309 is strict about what may appear after User-agent:, and the grammar is tighter than almost any generator enforces: The product token MUST contain only uppercase and lowercase letters ("a-z" and "A-Z"), underscores ("_"), and hyphens ("-").
The ABNF agrees, defining the identifier as 1*(%x2D / %x41-5A / %x5F / %x61-7A), which is hyphen, A to Z, underscore, a to z. No digits. No dots. No slashes or spaces.
The trap is that the name a crawler sends in its HTTP header is not the name you write here. The RFC gives the example directly: a crawler sending User-Agent: Mozilla/5.0 (compatible; ExampleBot/0.1; https://www.example.com/bot.html) is matched by the robots.txt line user-agent: ExampleBot, and the spec notes that the product token (ExampleBot) is a substring of the User-Agent HTTP header
.
The Extra user-agents box checks that. Paste a full header into it and no group is written for it: the page names what it rejected and tells you to write the short token instead. Until 4 October 2026 the line went out verbatim, digits, slashes, parentheses and all, so you got a group no conforming parser would match and rules that looked applied while being inert. If every token you give is invalid the file comes out with no group at all, rather than falling back to * and quietly applying your rules to every crawler.
Two further rules save work. Matching MUST use case-insensitive matching
, so googlebot and Googlebot are the same group and there is no need to write both. And if the same token appears twice, the matching groups' rules MUST be combined into one group
rather than the second replacing the first.
Set the preset to Custom rules, type two crawler names, and this page writes one group per name:
User-agent: Googlebot Disallow: /admin/ User-agent: Bingbot Disallow: /admin/
That file restricts Googlebot and Bingbot. It restricts nothing else at all. RFC 9309: If no matching group exists, crawlers MUST obey the group with a user-agent line with the "*" value, if present.
There is no such group here, and the rule for a crawler with no applicable group is that If no match is found amongst the rules in a group for a matching user-agent or there are no rules in the group, the URI is allowed.
If you want a named exception, write the * group as well and let the named group override it for that one crawler. A file that names bots without a * fallback is usually a file whose author believed naming the big two covered the field.
RFC 9309 lists exactly three special characters crawlers MUST support.
# Designates a line comment.Everything after it on the line is ignored. A path containing a fragment,
Disallow: /page#section, is therefore read as Disallow: /page, which blocks far more than intended. Fragments never reach a server anyway, so they do not belong in robots.txt at all.$ Designates the end of the match pattern.
Disallow: /*.pdf$ blocks paths ending in .pdf. Without the $ it would also block /report.pdf.html.* Designates 0 or more instances of any character.Note that matching is a prefix match to begin with, so
Disallow: /admin already covers /administrator. The trailing slash in /admin/ is what narrows it to the folder.If you need to block a path that genuinely contains one of these characters, the RFC says to percent-encode it: If crawlers match special characters verbatim in the URI, crawlers SHOULD use "%" encoding
, giving /path/file-with-a-%2A.html for a literal asterisk.
RFC 9309: Octets in the URI and robots.txt paths outside the range of the ASCII coded character set, and those in the reserved range defined by [RFC3986], MUST be percent-encoded as defined by [RFC3986] prior to comparison.
The RFC's own table gives /foo/bar/ followed by U+E38384 as encoding to /foo/bar/%E3%83%84.
Type /café/ into the Disallow box and you now get Disallow: /caf%C3%A9/, which is what the standard wants. A space comes out as %20. The number sign matters most of the three, because it starts a comment: /page#section used to be emitted verbatim and parsed as Disallow: /page, blocking everything under /page rather than one fragment, and it now comes out as /page%23section.
Three things deliberately survive the encoder, because encoding them would break what they mean. * and $ are the wildcards robots.txt itself defines, so /*.pdf$ is passed through. The reserved sub-delimiters are legal in a path unencoded, and the RFC says a percent-encoded ASCII octet MUST be unencoded prior to comparison, unless it is a reserved character
, so encoding them would change what matches. And an existing escape is left alone: /caf%C3%A9/ typed in stays itself rather than becoming /caf%25C3%25A9/, which would match nothing.
Paths are also case sensitive, which trips up anyone coming from Windows or from a case-insensitive CMS. Google: The path value must start with / to designate the root and the value is case-sensitive.
Disallow: /Admin/ does not block /admin/.
Three operational limits, none of which the output above shows you.
The parsing limit MUST be at least 500 kibibytes. Google enforces exactly that:
Google enforces a robots.txt file size limit of 500 kibibytes (KiB). Content which is after the maximum file size is ignored.A rule past that cut-off does nothing, and nothing warns you.
Crawlers SHOULD NOT use the cached version for more than 24 hours, so a change can take a day to take effect even for a crawler that visits constantly.
apply only to the host, protocol, and port number where the robots.txt file is hosted. The robots.txt on
https://example.com does not govern http://example.com, https://www.example.com or a subdomain. Each needs its own file at its own root.This is the clause worth knowing before your next outage. RFC 9309 splits fetch failures into two cases that point in opposite directions.
For 4xx: If a server status code indicates that the robots.txt file is unavailable to the crawler, then the crawler MAY access any resources on the server.
No robots.txt means no restrictions, which is why a site with nothing to block does not need the file at all.
For 5xx: If the robots.txt file is unreachable due to server or network errors, this means the robots.txt file is undefined and the crawler MUST assume complete disallow.
So a robots.txt returning 500 during an incident asks every conforming crawler to stop fetching the entire site, and the RFC allows that to persist: if the file stays undefined for a reasonably long period of time (for example, 30 days)
, crawlers may then fall back to treating it as unavailable. Redirects are handled separately, with crawlers expected to follow at least five consecutive redirects, even across authorities
.
Google's robots.txt documentation names the fields it reads: user-agent, allow, disallow and sitemap. On everything else it is blunt: other fields such as crawl-delay aren't supported
. RFC 9309 does not define Crawl-delay either; the ABNF comment simply invites implementers to define additional lines you need
.
The checkbox above is therefore for other crawlers, several of which do honour it, and the label calling it non-standard is accurate. It will not slow Googlebot down. Crawl rate for Google is a Search Console matter, not a robots.txt one.
Sitemap is the useful non-group field. It takes an absolute URL, it is not tied to any user-agent group, and Google notes You can specify multiple sitemap fields, with no limit to the number of sitemaps you can include.
One caution about the Block all crawlers preset: adding a sitemap to a file that says Disallow: / produces a document that hands over a map and then forbids every address on it, including the sitemap's own URL. Only /robots.txt itself is exempt, because The /robots.txt URI is implicitly allowed.
It validates the two things that silently produce a file which does not do what it says. Paths are percent-encoded, including the # that would otherwise truncate a rule, and user-agent names are checked against the grammar in RFC 9309 so an invalid product token is reported rather than written into a group that never matches. It still does not check that a path exists, or that blocking it is wise.
Both path boxes are read under Allow all crawlers and under Custom rules. Under Block all crawlers neither is, because that preset forbids every path, and the page says so in the status line rather than dropping what you typed in silence.
It does not know your site. It cannot tell whether /admin/ exists, whether blocking it will strand a stylesheet the page needs to render, or whether the URL you are blocking is one that already ranks. Blocking a path that Google currently crawls does not remove it from the index; it freezes what Google last knew and removes the snippet.
And it cannot test the result. Paste the output into a validator or Search Console's robots.txt report before you rely on it, particularly if you used a wildcard.
* group and the rule that an unmatched URI is allowed; the most-octets precedence rule and the allow-wins tie-break; the percent-encoding requirement and its examples; the three special characters; the 500 kibibyte parsing limit; the 24-hour caching limit; the five-redirect rule; and the opposite treatment of 4xx and 5xx responses.Every quotation above was read on 2 October 2026 from the document named, and the RFC quotations were checked against the plain-text copy at rfc-editor.org rather than against a summary. The output samples are the literal output of the generator on this page, produced by running its own code against those inputs. Part of the QuikUtil tools collection; the Meta Tag Generator writes the noindex rule this page cannot.
No, and this is the most expensive misunderstanding in the file. Google: it can't index the content of pages which are disallowed for crawling, but it may still index the URL and show it in search results without a snippet
. To keep a page out of results you need a noindex rule, and Google is equally clear that For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.
Blocking and deindexing are mutually exclusive instructions. Pick one.
The longer one, not the first one. RFC 9309: The most specific match found MUST be used. The most specific match is the match that has the most octets.
So Allow: / followed by Disallow: /admin/ does block /admin/, because /admin/ is seven octets against one. Order in the file is irrelevant. When an allow and a disallow are exactly the same length, the "allow" rule SHOULD be used
.
Letters, underscores and hyphens, and nothing else. RFC 9309: The product token MUST contain only uppercase and lowercase letters ("a-z" and "A-Z"), underscores ("_"), and hyphens ("-").
That rules out digits, dots, slashes and spaces, so pasting a browser-style user agent string into the Extra user-agents box produces a line no conforming parser will match. The name you want is the short product token, such as Googlebot, not the full header it appears inside.
Almost always yes. RFC 9309: If no matching group exists, crawlers MUST obey the group with a user-agent line with the "*" value, if present.
If there is no such group either, nothing applies, and If no match is found amongst the rules in a group for a matching user-agent or there are no rules in the group, the URI is allowed.
A file listing only Googlebot and Bingbot leaves every other crawler completely unrestricted.
Because # starts a comment. RFC 9309 lists it as a special character crawlers MUST support, one that Designates a line comment.
So Disallow: /page#section is read as Disallow: /page plus a comment, which blocks every URL beginning /page. This page does not escape it for you.
They do now, because the page encodes them. RFC 9309 requires that Octets in the URI and robots.txt paths outside the range of the ASCII coded character set, and those in the reserved range defined by [RFC3986], MUST be percent-encoded as defined by [RFC3986] prior to comparison.
So /café/ has to be written /caf%C3%A9/ and /my docs/ has to be /my%20docs/. The page applies that encoding for you, and it also encodes a space as %20 and a number sign as %23. Until 4 October 2026 it passed paths through unchanged apart from adding a leading slash, so the encoding was yours to do and a number sign silently truncated the rule.
Not for Google. Its robots.txt documentation says Google supports user-agent, allow, disallow and sitemap, and that other fields such as crawl-delay aren't supported
. RFC 9309 does not define it either. Some other crawlers honour it, which is why the checkbox exists and why it is labelled non-standard.
The two error families mean opposite things, which catches people out during an outage. RFC 9309: for 4xx, If a server status code indicates that the robots.txt file is unavailable to the crawler, then the crawler MAY access any resources on the server.
For 5xx, If the robots.txt file is unreachable due to server or network errors, this means the robots.txt file is undefined and the crawler MUST assume complete disallow.
A 404 means crawl everything. A 500 means crawl nothing.
Yes, since 4 October 2026. Both path boxes are read under Allow all crawlers and under Custom rules. Until that date Allow paths was read only under Custom rules, so under Allow all crawlers, which is how this page loads, anything typed there was discarded while the Disallow box beside it was honoured, and nothing said so. Under Block all crawlers neither box is used, because that preset forbids every path, and the page now tells you that rather than ignoring what you typed.