AI 助手通过两种方式抓取网页:一类收集内容用于训练模型,另一类在用户查询时实时抓取页面并引用。你可以在 robots.txt 中分别对这两类做不同处理,但前提是了解每个 user‑agent 的功能。
常见 AI 爬虫与作用
- GPTBot / ClaudeBot / CCBot / Google‑Extended / Applebot‑Extended:收集内容用于训练。若不想让内容被用于未来模型,可在 robots.txt 中Disallow: /。
- OAI‑SearchBot / ChatGPT‑User / Claude‑SearchBot / Claude‑User / PerplexityBot / Perplexity‑User:在用户查询时抓取页面并引用。若想让页面出现在 AI 回答中,应允许这些爬虫访问。
推荐 robots.txt 示例
```
# 训练爬虫:拒绝
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
# 搜索与回答爬虫:允许
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
```
保持对 Googlebot 与
* 的原有规则不变。此配置既能让你的网站在 AI 搜索中可见,又能避免被用于训练。若有付费或私有内容,可进一步细化路径。测试方法可用 Python 的 urllib.robotparser 或在线工具检查。
