“We blocked /admin in robots.txt — why does it show up on Google?” Because robots.txt blocks crawling, not indexing: Google lists the URL (title-only, no description) without ever fetching it. Teams discover this during security reviews, and the fix is a different tool entirely. This guide draws the bright line between robots.txt, noindex, authentication and removal — with the Disallow-matching rules that decide real cases.
Part of the SEO publishing guide. Build rules in the robots.txt generator; map crawling in the sitemap generator.
The bright line (memorize this table)
| Goal | Tool | Why |
|---|---|---|
| Save crawl budget | robots.txt Disallow | Stops fetching (facets, filters, staging) |
| Keep out of index | noindex meta + allow crawl | Google must fetch to see noindex |
| Keep secret | Authentication | Neither robots nor noindex is security |
| Remove urgently | Removals tool + noindex | Temporary hide, then permanent fix |
The classic error is combining Disallow with noindex on the same URL: blocked crawlers never see the noindex tag, so the URL lingers indexed indefinitely. To deindex, allow crawling of the noindexed URL — counterintuitive, correct, and the fix for most “blocked but indexed” mysteries.
Disallow matching: longest rule wins
- Longest match applies:
Allow: /admin/publicbeatsDisallow: /adminfor that path. Order in the file does not matter — specificity does. $anchors ends:Disallow: /*.pdf$blocks PDFs only, not/pdf-guidepages.*wildcards spans:Disallow: /*?sort=kills faceted crawl traps while keeping clean category URLs.- Crawl-delay is advisory: respected by some crawlers (shared-host relief at
Crawl-delay: 5), ignored by Googlebot — use Search Console crawl settings instead.
Validate every ruleset in a tester before deploying — one misplaced wildcard has deindexed entire blogs. The 500KB robots cap rarely binds, but bloated files signal undisciplined crawling that sitemaps should instead organize (see sitemap splitting).
Three real cases (admin, staging, facets)
Admin login indexed: remove the Disallow, add noindex, let Google recrawl, then re-evaluate — plus authentication, because login pages deserve locks, not hints. Staging clone indexed: noindex + password-protect staging permanently; relying on Disallow alone leaks titles. Faceted filters eating budget: Disallow parameter variants (?color=, ?sort=), canonical clean versions, submit only canonicals in sitemaps. Bing vs Google diverge on edge directives — test both Search Console and Bing Webmaster before declaring victory.
General guidance only. Robots.txt is a crawling courtesy with security-adjacent consequences — audit quarterly, never assume.