# https://citefacts.com # Public, provenance-backed structured data. Everything crawlable here is # intended to be indexed, retrieved, and cited by people and software agents. # # Operator classes and the reasoning behind each are documented in # src/agent_site/crawlers.py and docs/AGENT_DISCOVERABILITY.md. User-agent: * Allow: / Disallow: /traffic-dashboard.html # search: search/answer indexes — allowed; being findable is the experiment User-agent: Googlebot # Google Search User-agent: Bingbot # Microsoft Bing / Copilot User-agent: DuckDuckBot # DuckDuckGo User-agent: Applebot # Apple User-agent: YandexBot # Yandex User-agent: OAI-SearchBot # OpenAI ChatGPT Search index User-agent: Claude-SearchBot # Anthropic Claude search index User-agent: PerplexityBot # Perplexity index User-agent: YouBot # You.com Allow: / Disallow: /traffic-dashboard.html # retrieval: live fetch for a user's question — allowed; this is the citation path User-agent: ChatGPT-User # OpenAI ChatGPT live fetch User-agent: Claude-User # Anthropic Claude live fetch User-agent: Perplexity-User # Perplexity live fetch User-agent: MistralAI-User # Mistral live fetch Allow: / Disallow: /traffic-dashboard.html # training: model-training corpora — allowed during the discovery baseline; revisitable User-agent: GPTBot # OpenAI model training User-agent: ClaudeBot # Anthropic model training User-agent: Google-Extended # Google Gemini training User-agent: Applebot-Extended # Apple model training User-agent: Amazonbot # Amazon User-agent: meta-externalagent # Meta model training User-agent: CCBot # Common Crawl User-agent: Bytespider # ByteDance User-agent: cohere-ai # Cohere Allow: / Disallow: /traffic-dashboard.html Sitemap: https://citefacts.com/sitemap.xml Sitemap: https://citefacts.com/recalls/sitemap.xml