Robots.txt generator and tester

Build a robots.txt file with Allow and Disallow rules per user-agent plus sitemaps, and test whether a crawler may fetch a URL using the RFC 9309 matching rules.

robots.txt tells crawlers which URLs they may crawl. To remove or hide pages from search results (indexing), use noindex or the search engine removal tools instead. It is not access control or security either: crawlers follow it voluntarily, and anyone can read the file.

Presets

The AI crawler names are examples. Names change over time and obeying robots.txt is voluntary for each operator. Check each company’s documentation for current names.

Editing mode

Paths start with /. * matches any characters and a final $ marks the end of the URL (for example /*.pdf$). An empty Disallow allows everything.

URL tester

Checks, with the RFC 9309 rules, whether a crawler that obeys this robots.txt may crawl the URL.

How to use

  1. Pick a preset, or enter user-agents and Allow/Disallow rules for each group. Add sitemap URLs if you have them.
  2. The robots.txt text appears below. Choose "Edit text directly" to paste an existing file and check or fix it.
  3. In the URL tester, enter a crawler name and a URL or path to see whether it is allowed and which rule decided.
  4. Copy the text or download robots.txt and place it at the root of your site (https://example.com/robots.txt).

Notes and limits

  • robots.txt controls crawling, not indexing. A blocked URL can still appear in search results if other pages link to it. To keep a page out of results, use noindex (crawlers must be able to read the page to see it) or the search engine removal tools.
  • It is not access control or security. Crawlers follow it voluntarily, and anyone can read the file, so listing secret paths there only advertises them.
  • The AI crawler preset is an example. Crawler names change, and the list does not cover every AI service.
  • The tester follows RFC 9309. Some crawlers add their own behavior, such as support for Crawl-delay.

How it works

Group selection: the crawler name (product token) is compared with each User-agent line without regard to case, and all matching groups are combined. If no group names the crawler, the * group is used; if there is none, everything is allowed.

Rule selection: among the rules that match the URL path (including the query) from its start, the one with the longest path in bytes wins. If an Allow and a Disallow of the same length match, Allow wins. With no matching rule the URL is allowed, and /robots.txt is always allowed.

* matches any sequence of characters and a final $ matches the end of the URL. Non-ASCII characters and their %XX percent-encoded forms are treated as equal; paths are case-sensitive.

FAQ

Can robots.txt remove a page from search results?

No. It only stops crawling. To remove a page, add noindex, let crawlers read the page, and use the search engine’s removal tool.

Can I stop AI companies from training on my site?

You can disallow known AI crawlers by name, but compliance is voluntary for each operator and names change. It is not a guarantee.

Is Crawl-delay supported?

It is not part of RFC 9309 and Google ignores it. Some crawlers use it, so you can write it in text mode, but the tool flags it.

Is anything I enter sent to a server?

No. Building and testing happen in your browser. Nothing is sent or stored.

Last updated: