# ========================================================================== # The Green Line โ€” robots.txt # Allows search engines, blocks AI training crawlers. # Ref: PRD ยง AI Metering / Content Protection # ========================================================================== # Search engines โ€” welcome User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: * Allow: / # Sitemap (generated by Yoast / wp-sitemap) Sitemap: https://thegreeline.to/wp-sitemap.xml # -------------------------------------------------------------------------- # AI / LLM training crawlers โ€” blocked # These bots scrape content to train large language models. # We block them to protect our editorial content and copyright. # -------------------------------------------------------------------------- # OpenAI User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Disallow: / # Google AI (Gemini training, not search indexing) User-agent: Google-Extended Disallow: / # Common Crawl (used by many AI labs) User-agent: CCBot Disallow: / # Anthropic User-agent: anthropic-ai Disallow: / User-agent: Claude-Web Disallow: / # Meta / Facebook AI User-agent: FacebookBot Disallow: / # Apple AI User-agent: Applebot-Extended Disallow: / # Perplexity User-agent: PerplexityBot Disallow: / # Cohere User-agent: cohere-ai Disallow: / # ByteDance / TikTok User-agent: Bytespider Disallow: / # Amazon User-agent: Amazonbot Disallow: / # Scrapy-based generic crawlers User-agent: Scrapy Disallow: /