How to use the Robots.txt Generator
- Choose the default rule — most sites use "Allow all".
- List paths to block in "Disallow", one per line (e.g. /admin/).
- Pick an AI crawler policy and enter your sitemap URL.
- Press "Download" and upload robots.txt to your site root (https://your-domain/robots.txt).
Examples
Block the admin area and AI training
Disallow /admin/, /cart | AI: block training crawlers | Sitemap
User-agent: *
Disallow: /admin/
Disallow: /cart
# Block AI training crawlers
User-agent: GPTBot
User-agent: Google-Extended
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Bytespider
Disallow: /
Sitemap: https://example.com/sitemap.xml
Before you block AI crawlers
AI crawlers come in two kinds: training crawlers (GPTBot, Google-Extended, ClaudeBot, CCBot) collect content to train models, while search and user-triggered crawlers (OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-SearchBot) fetch pages when someone asks a question and cite the source.
If you want your site cited in ChatGPT search, Perplexity and similar AI answers, block only training crawlers and keep search crawlers allowed. Blocking Google-Extended does not affect Google Search rankings.
Specs & key facts
| Standard | RFC 9309 (Robots Exclusion Protocol) |
|---|---|
| Options | Default rule, Disallow, Allow, Crawl-delay, Sitemap |
| AI training crawlers | GPTBot, Google-Extended, ClaudeBot, CCBot, Applebot-Extended, meta-externalagent, Bytespider |
| AI search crawlers | OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User |
| Location | Site root, e.g. https://example.com/robots.txt |
| Price | Free, no sign-up |
FAQ
Does disallowing a page keep it out of Google?
Not necessarily. robots.txt only blocks crawling; if other sites link to the page, Google may still list the URL. To keep a page out of results, use noindex and allow crawling so Google can see it.
Does Crawl-delay work?
Bing and some crawlers respect it; Googlebot does not. Most sites do not need it.
Do all bots obey robots.txt?
Major search engines and AI companies do, but robots.txt is a public request, not access control. Protect sensitive data with authentication.
Can paths use wildcards?
Yes. * matches any characters and $ marks the end of the URL — Disallow: /*.pdf$ blocks every PDF.