Robots.txt Explained: What It Does, How to Write One, and 5 Mistakes That Can Kill Your Rankings
One small text file decides which parts of your website Google's crawler is allowed to visit. Get it right and your robots.txt file keeps crawlers focused on the pages that matter; get it wrong and you can accidentally block your entire site from search overnight. This guide explains exactly what robots.txt does, how to write one correctly, and the five mistakes I see site owners make most often.
What is a robots.txt file, and where does it live?
A robots.txt file is a plain-text file that tells search engine crawlers which URLs they may and may not access on your site. It lives in one fixed location: the root of your domain, so for www.example.com it must be reachable at www.example.com/robots.txt. Google is explicit about this: a robots.txt file follows the Robots Exclusion Standard, consists of one or more rules, and each rule blocks or allows access for a crawler to a specified path on the domain where the file is hosted. A site has exactly one robots.txt, and it is case-sensitive — Robots.txt will not do.
The most important thing to understand is the default: unless you say otherwise, everything on your site is allowed for crawling. Robots.txt works by exclusion. You list the paths you want to exclude, and everything else is implicitly allowed. An empty file — or no file at all — means "crawl everything".
Does robots.txt stop a page from appearing in Google's search results?
No — and this is the single most misunderstood fact in technical SEO. Blocking a page in robots.txt only asks crawlers not to visit it. It does not remove the page from Google's index. If other pages on the web link to a blocked page, Google can still index it without ever crawling it — often with the page title replaced by the message that the page is blocked by robots.txt.
Google's own documentation is blunt about this: do not use robots.txt to hide web pages from Google Search results. If you want a page out of the index, use a noindex directive instead — via a meta robots tag or an X-Robots-Tag HTTP header — which works precisely because it lets Google crawl the page and read the instruction. Blocking the crawl makes noindex unreadable, so the two tools fight each other. That combination — blocking a page in robots.txt while also adding noindex to it — is one of the most common technical SEO errors in existence, and it silently leaves pages in the index for months.
Robots.txt is also not security. Malicious scrapers ignore it completely. Never rely on it to protect genuinely sensitive content — use authentication for that.
How do I write a robots.txt file that actually works?
The syntax is small enough to learn in five minutes. A robots.txt file is made of rules, and each rule uses a handful of directives. Google's Search Central documentation lists the building blocks:
-
User-agent: identifies which crawler the following rules apply to.
User-agent: Googlebottargets Google's main crawler;User-agent: *matches every crawler. Google's documentation notes that crawlers match the user-agent line against their own, case-insensitively, so a group matching both*andGooglebotwill be read by Googlebot too. -
Disallow: blocks crawling of a path.
Disallow: /admin/blocks everything under/admin/. An empty value —Disallow:with nothing after it — means nothing is disallowed, i.e. everything is allowed. - Allow: carves an exception inside a blocked path, e.g. allowing one public page inside a disallowed directory.
-
Sitemap: points crawlers at your XML sitemap, e.g.
Sitemap: https://www.example.com/sitemap.xml. Google supports sitemap entries in robots.txt as a hint, but listing a sitemap there does not replace submitting it in Search Console.
Google's crawler also understands limited pattern matching: * matches any sequence of characters and $ marks the end of a URL, so Disallow: /*.pdf$ would block every URL on the site ending in .pdf. Full regular expressions are not supported.
Here is a sane, real-world starting point for a typical small website:
User-agent: *
Disallow: /admin/
Disallow: /wp-admin/
Disallow: /cart/
Disallow: /checkout/
Disallow: /search?
Sitemap: https://www.example.com/sitemap.xml
What it does: every crawler is blocked from the admin area, the shopping cart and checkout, and internal search result pages (query-string URLs that create infinite duplicate pages), while everything else stays crawlable. The sitemap line tells crawlers where the important URLs are. If you would rather hand-write this than build it from memory, Toolxz's free robots.txt generator builds a correctly formatted file from checkboxes — no signup needed — which is exactly the sort of small task a tool like this exists for.
What are the 5 most common robots.txt mistakes?
1. Blocking the whole site by accident
User-agent: *
Disallow: /
These two lines tell every crawler to stay away from everything. It is the correct configuration for a staging site — and a catastrophe on a live one. This usually happens when someone copies a robots.txt from a development environment during launch and forgets to flip it. If your traffic fell to zero right after a deploy, check your robots.txt before you do anything else.
2. Using robots.txt to "remove" pages from Google
As covered above, blocking ≠ noindex. A blocked page with external links can sit in Google's index indefinitely, and a page you blocked and tagged noindex may never be recrawled so Google never sees the noindex. For pages you want gone from search, allow the crawl and use noindex, or remove the page and let it return a 404/410.
3. Guessing at case and syntax
Robots.txt rules are case-sensitive: Disallow: /Admin/ does not block /admin/. Other quiet breakages include putting the file anywhere other than the domain root (Google ignores a robots.txt at /pages/robots.txt), splitting one site's rules across multiple files on the same host, and assuming unsupported directives like Crawl-delay are honoured — Google's parser does not officially support Crawl-delay, so don't rely on it.
4. Blocking CSS, JS, and image files
Google's crawler needs your CSS and JavaScript to render pages the way users see them. Blocking /wp-content/ or /static/ wholesale — common in copy-pasted templates — can make your pages render as broken skeletons during Google's rendering, which is not the impression you want the ranking systems to work with. Block specific sensitive paths, not the assets folder.
5. Writing it once and never checking it again
Sites change: directories get renamed, sitemaps move, developers push hotfixes. A rule that was harmless last year can quietly block your most important section today. Make robots.txt review part of every site redesign and major deploy, and whenever traffic drops suddenly for no obvious reason, look at this file first — it takes thirty seconds and has found more "mysterious" SEO disasters than any audit tool.
How do I test and monitor my robots.txt file?
Start with the simplest check of all: open yourdomain.com/robots.txt in a browser or run curl yourdomain.com/robots.txt and read what comes back. You would be surprised how many "SEO emergencies" are a visibly broken file.
Then verify what Google actually sees. In Google Search Console, open Settings → robots.txt: Google shows the robots.txt file it fetched, when it last crawled it, and any warnings or syntax problems it found. This report is the source of truth, because what Google has cached can differ from what you just uploaded — Google refetches robots.txt periodically, so changes are not instant. If you are debugging why a URL is not being crawled, this report and the URL Inspection tool together tell you whether robots.txt is the reason.
For a fuller picture of how crawlers reach your pages, I also recommend learning to read the Performance report in Search Console — I wrote a walkthrough of exactly that here, and the pattern-finding habit from that article applies to diagnosing crawl problems just as well.
FAQ
Does every website need a robots.txt file?
No. Small, fully linked sites under a few hundred pages usually do fine without one. You need it when you have sections crawlers shouldn't waste time on: admin panels, internal search results, cart and checkout flows, staging areas, or API endpoints.
Will blocking a page in robots.txt remove it from Google?
No. It only stops crawlers from visiting the page. The page can still be indexed via links from elsewhere. Use noindex for removal.
Can I have different rules for Googlebot and Bingbot?
Yes. Add a separate group of rules under User-agent: Googlebot and another under User-agent: Bingbot, plus a User-agent: * fallback group for everything else.
How long does a robots.txt change take to affect Google?
Google caches robots.txt and refetches it periodically (typically at least daily for active sites), so allow a day or two for changes to take effect. There is no instant "flush".
Does the Sitemap line in robots.txt submit my sitemap?
It helps crawlers discover it, but Google's documentation recommends also submitting the sitemap directly in Search Console under Sitemaps, so you get reporting on how many URLs were discovered and indexed.
Founder of Toolxz (toolxz.com) — 45+ free browser-based tools. I write about practical AI tooling and developer workflows.
Top comments (0)