如何使用robots.txt 產生器
- 選擇預設規則:一般網站選「允許全部」。
- 在「禁止抓取的路徑」每行輸入一個路徑,例如 /admin/。
- 選擇 AI 爬蟲政策,並填入 Sitemap 網址。
- 按「下載」取得 robots.txt,上傳到網站根目錄(https://你的網域/robots.txt)。
範例
一般網站,封鎖後台並禁止 AI 訓練
禁止:/admin/、/cart|AI:封鎖訓練用爬蟲|Sitemap
User-agent: *
Disallow: /admin/
Disallow: /cart
# Block AI training crawlers
User-agent: GPTBot
User-agent: Google-Extended
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Bytespider
Disallow: /
Sitemap: https://example.com/sitemap.xml
封鎖 AI 爬蟲前要考慮什麼
AI 爬蟲分成兩類:訓練用(GPTBot、Google-Extended、ClaudeBot、CCBot)會把內容拿去訓練模型;搜尋與即時讀取用(OAI-SearchBot、ChatGPT-User、PerplexityBot、Claude-SearchBot)則是在使用者提問時讀取網頁並附上來源連結。
如果你希望網站出現在 ChatGPT、Perplexity 等 AI 搜尋的引用來源中,建議只封鎖訓練用爬蟲、不要封鎖搜尋用爬蟲。封鎖 Google-Extended 不會影響 Google 搜尋排名。
規格與重點
| 依循標準 | RFC 9309(Robots Exclusion Protocol) |
|---|---|
| 可設定 | 預設規則、Disallow、Allow、Crawl-delay、Sitemap |
| AI 訓練爬蟲 | GPTBot、Google-Extended、ClaudeBot、CCBot、Applebot-Extended、meta-externalagent、Bytespider |
| AI 搜尋爬蟲 | OAI-SearchBot、ChatGPT-User、Claude-SearchBot、Claude-User、PerplexityBot、Perplexity-User |
| 放置位置 | 網站根目錄,例如 https://example.com/robots.txt |
| 費用 | 免費、免註冊 |
常見問題
用 robots.txt 禁止的頁面就不會出現在 Google 嗎?
不一定。robots.txt 只禁止「抓取」,如果其他網站連到該頁,Google 仍可能只顯示網址。要確保不被收錄,請改用 noindex(並允許抓取讓 Google 看得到 noindex)。
Crawl-delay 有用嗎?
Bing 與部分爬蟲會遵守,Googlebot 不支援。一般網站不需要設定。
所有爬蟲都會遵守 robots.txt 嗎?
主要的搜尋引擎與 AI 公司爬蟲都會遵守,但 robots.txt 只是公開的請求,惡意爬蟲可能忽略,敏感資料請用登入或防火牆保護。
路徑可以用萬用字元嗎?
可以。* 代表任意字元、$ 代表網址結尾,例如 Disallow: /*.pdf$ 會封鎖所有 PDF。