選項
預設規則
robots.txt
輸入內容後會自動處理

如何使用robots.txt 產生器

  1. 選擇預設規則:一般網站選「允許全部」。
  2. 在「禁止抓取的路徑」每行輸入一個路徑,例如 /admin/。
  3. 選擇 AI 爬蟲政策,並填入 Sitemap 網址。
  4. 按「下載」取得 robots.txt,上傳到網站根目錄(https://你的網域/robots.txt)。

範例

一般網站,封鎖後台並禁止 AI 訓練

輸入
禁止:/admin/、/cart|AI:封鎖訓練用爬蟲|Sitemap
輸出
User-agent: *
Disallow: /admin/
Disallow: /cart

# Block AI training crawlers
User-agent: GPTBot
User-agent: Google-Extended
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Bytespider
Disallow: /

Sitemap: https://example.com/sitemap.xml

封鎖 AI 爬蟲前要考慮什麼

AI 爬蟲分成兩類:訓練用(GPTBot、Google-Extended、ClaudeBot、CCBot)會把內容拿去訓練模型;搜尋與即時讀取用(OAI-SearchBot、ChatGPT-User、PerplexityBot、Claude-SearchBot)則是在使用者提問時讀取網頁並附上來源連結。

如果你希望網站出現在 ChatGPT、Perplexity 等 AI 搜尋的引用來源中,建議只封鎖訓練用爬蟲、不要封鎖搜尋用爬蟲。封鎖 Google-Extended 不會影響 Google 搜尋排名。

規格與重點

依循標準RFC 9309(Robots Exclusion Protocol)
可設定預設規則、Disallow、Allow、Crawl-delay、Sitemap
AI 訓練爬蟲GPTBot、Google-Extended、ClaudeBot、CCBot、Applebot-Extended、meta-externalagent、Bytespider
AI 搜尋爬蟲OAI-SearchBot、ChatGPT-User、Claude-SearchBot、Claude-User、PerplexityBot、Perplexity-User
放置位置網站根目錄,例如 https://example.com/robots.txt
費用免費、免註冊

常見問題

用 robots.txt 禁止的頁面就不會出現在 Google 嗎?

不一定。robots.txt 只禁止「抓取」,如果其他網站連到該頁,Google 仍可能只顯示網址。要確保不被收錄,請改用 noindex(並允許抓取讓 Google 看得到 noindex)。

Crawl-delay 有用嗎?

Bing 與部分爬蟲會遵守,Googlebot 不支援。一般網站不需要設定。

所有爬蟲都會遵守 robots.txt 嗎?

主要的搜尋引擎與 AI 公司爬蟲都會遵守,但 robots.txt 只是公開的請求,惡意爬蟲可能忽略,敏感資料請用登入或防火牆保護。

路徑可以用萬用字元嗎?

可以。* 代表任意字元、$ 代表網址結尾,例如 Disallow: /*.pdf$ 會封鎖所有 PDF。