# Author-it robots.txt # Last updated: May 2026 User-agent: * Allow: / Disallow: /search Disallow: /cdn-cgi/ # Sitemap Sitemap: https://www.author-it.com/sitemap.xml # ----------------------------------------------- # Major search engines # ----------------------------------------------- User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: Slurp Allow: / User-agent: DuckDuckBot Allow: / User-agent: Baiduspider Allow: / User-agent: YandexBot Allow: / # ----------------------------------------------- # AI training and LLM crawlers # ----------------------------------------------- # OpenAI — ChatGPT browsing + training User-agent: GPTBot Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / # Google — Gemini / AI Overviews training User-agent: Google-Extended Allow: / # Anthropic — Claude training User-agent: ClaudeBot Allow: / # Perplexity User-agent: PerplexityBot Allow: / # Common Crawl (feeds many open LLMs incl. LLaMA, Mistral) User-agent: CCBot Allow: / # Meta — Llama training User-agent: FacebookBot Allow: / # Apple — Applebot-Extended covers AI/ML training User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / # Amazon — Alexa + Bedrock User-agent: Amazonbot Allow: / # Cohere User-agent: cohere-ai Allow: / # Diffbot (knowledge graph, feeds AI apps) User-agent: Diffbot Allow: / # You.com User-agent: YouBot Allow: / # Bytedance / Bytespider (feeds Doubao and TikTok AI) User-agent: Bytespider Allow: / # Timpi (emerging EU-based AI search) User-agent: Timpibot Allow: / Sitemap: https://www.author-it.com/sitemap.xml