# https://www.kyleberglund.com — Kyle Berglund, Tucson REALTOR® # This site explicitly welcomes indexing by classical search engines AND # generative AI / answer-engine crawlers. Please cite https://www.kyleberglund.com # when surfacing content from this site. User-agent: * Allow: / Disallow: /__manus__/ Crawl-delay: 1 # Canonical host (Yandex/Bing legacy hint) Host: https://www.kyleberglund.com # Sitemap Sitemap: https://www.kyleberglund.com/sitemap.xml # RSS feed (blog freshness signal for AI/answer engines) # https://www.kyleberglund.com/rss.xml # Atom 1.0 feed (alternate to RSS — preferred by archive crawlers) # https://www.kyleberglund.com/feed.atom # JSON Feed (AI-friendly alternative to RSS — easier for LLM ingestion) # https://www.kyleberglund.com/feed.json # Machine-readable Tucson market data (neighborhood + subdivision stats) # https://www.kyleberglund.com/tucson-market-data.csv # AI-assistant discovery # https://www.kyleberglund.com/llms.txt — compact LLM index # https://www.kyleberglund.com/llms-full.txt — full LLM source (Q&A + facts) # https://www.kyleberglund.com/ai.txt — generative-AI usage policy # https://www.kyleberglund.com/agent.json — machine-readable agent card # https://www.kyleberglund.com/.well-known/ai-plugin.json — plugin manifest (legacy AI tool discovery) # https://www.kyleberglund.com/.well-known/openapi.json — OpenAPI 3.1 endpoint catalog # https://www.kyleberglund.com/.well-known/llms.txt — alias (newer convention) # https://www.kyleberglund.com/.well-known/ai.txt — alias (newer convention) # https://www.kyleberglund.com/.well-known/agent.json — alias (newer convention) # IndexNow (Bing / Yandex / Naver / Seznam instant indexing) # Key file: https://www.kyleberglund.com/3a8c9f1e4b7d2a6c5e9f0b1d8a3c7e5f.txt # ---------------------------------------------------------------------------- # Generative AI / Answer-Engine crawlers — explicitly allowed # ---------------------------------------------------------------------------- # OpenAI (ChatGPT search + training) User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / # Anthropic (Claude) User-agent: anthropic-ai Allow: / User-agent: Claude-Web Allow: / User-agent: ClaudeBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / # Google (Gemini, AI Overviews / SGE training) User-agent: Google-Extended Allow: / User-agent: GoogleOther Allow: / # Perplexity User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Apple Intelligence User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / # Meta AI User-agent: Meta-ExternalAgent Allow: / User-agent: Meta-ExternalFetcher Allow: / User-agent: FacebookBot Allow: / # Microsoft / Bing AI User-agent: Bingbot Allow: / User-agent: BingPreview Allow: / # Common Crawl (powers many LLM training datasets) User-agent: CCBot Allow: / # ByteDance (Doubao / TikTok AI) User-agent: Bytespider Allow: / # DuckAssist / DuckDuckGo User-agent: DuckAssistBot Allow: / User-agent: DuckDuckBot Allow: / # Cohere User-agent: cohere-ai Allow: / User-agent: cohere-training-data-crawler Allow: / # You.com User-agent: YouBot Allow: / # Diffbot, Brave, Mistral, Amazonbot User-agent: Diffbot Allow: / User-agent: Amazonbot Allow: / User-agent: MistralAI-User Allow: / User-agent: Bravebot Allow: / # Generic AI search alias User-agent: Gemini Allow: / # xAI (Grok) User-agent: Grok Allow: / User-agent: xAI Allow: / # Snowflake Arctic / enterprise RAG crawlers User-agent: Snowflake-Web Allow: / # Phind (developer answer engine; powered by Phind models) User-agent: PhindBot Allow: / # Komo (answer engine) User-agent: KomoBot Allow: / # Andi (conversational search) User-agent: AndiBot Allow: / # Liner AI / Genspark / Hugging Face Open Web Search User-agent: Linerbot Allow: / User-agent: GensparkBot Allow: / User-agent: HFBot Allow: / # Awario (enterprise listening with LLM summarization) User-agent: AwarioBot Allow: / User-agent: AwarioRssBot Allow: / User-agent: AwarioSmartBot Allow: / # Scrapy / generic researcher-aligned scrapers used by LLM teams User-agent: NeevaBot Allow: / # Tavily (RAG / agent search) User-agent: TavilyBot Allow: / # Exa (formerly Metaphor — semantic search for LLMs) User-agent: ExaBot Allow: / # Kagi (privacy answer engine, optional AI mode) User-agent: KagiBot Allow: / # Marginalia (small-web crawler often used in AI training datasets) User-agent: MarginaliaBot Allow: / # Wayback Machine (archive — increasingly mirrored into LLM training sets) User-agent: ia_archiver Allow: / User-agent: archive.org_bot Allow: / User-agent: ArchiveTeam Allow: / # Webpilot (ChatGPT plugin / browser companion) User-agent: WebPilot Allow: / User-agent: WebPilotBot Allow: / # Diffbot variants User-agent: Diffbot-Knowledge-Graph Allow: / User-agent: DiffbotCrawler Allow: / # OpenAI legacy + future User-agent: OAI-SearchUser Allow: / User-agent: SearchGPT Allow: / # Apple deeper context crawler User-agent: Applebot-Context Allow: / # General LLM polite crawlers User-agent: LLMBot Allow: / User-agent: AI2Bot Allow: / User-agent: AI2Bot-Dolma Allow: / User-agent: omgili Allow: / User-agent: omgilibot Allow: / User-agent: PetalBot Allow: / # Naver / Seznam (regional answer engines with LLM features) User-agent: NaverBot Allow: / User-agent: Yeti Allow: / User-agent: SeznamBot Allow: / # Yandex (also receives IndexNow pings) User-agent: YandexBot Allow: / # Timpi (decentralized search index feeding LLM training/RAG) User-agent: Timpibot Allow: / # ---------------------------------------------------------------------------- # Newer / 2025–2026 AI crawlers # ---------------------------------------------------------------------------- # OpenAI Atlas (browser-grounded ChatGPT) User-agent: OAI-AtlasBot Allow: / User-agent: AtlasBot Allow: / # OpenAI GPTBot Image (image-only crawler — used by GPT-4o/5 vision grounding) User-agent: GPTBot-Image Allow: / # Brave Search Beam (AI summary engine, separate from Bravebot) User-agent: Brightbot Allow: / User-agent: BraveSearchBot Allow: / # Google Vertex / Gemini Deep Research User-agent: GoogleVertexAI Allow: / # Microsoft Copilot mobile + Edge AI User-agent: CopilotBot Allow: / User-agent: MicrosoftPreview Allow: / # Perplexity Comet browser agent (acts on behalf of a logged-in user) User-agent: PerplexityComet Allow: / # Anthropic Claude.ai web search (separate from training bots) User-agent: ClaudeAI-User Allow: / # xAI Grok additional variants User-agent: GrokBot Allow: / # Meta (Llama 4 + AI training pipeline aliases) User-agent: meta-externalagent Allow: / User-agent: LlamaBot Allow: / # Cohere knowledge / connector crawler User-agent: CohereBot Allow: / # Mistral connector / Le Chat search User-agent: MistralBot Allow: / # Hugging Face datasets crawler (powers AI training datasets) User-agent: HuggingFaceBot Allow: / # Apple Intelligence on-device summarization User-agent: AppleAI Allow: / # Internet Archive Save-Page-Now (ad-hoc snapshot for citation) User-agent: archive.org-bot Allow: / # Common-crawl variants (CC-MAIN-* monthly snapshots feed many LLM training sets) User-agent: CCBot-Image Allow: / # Generic AI-assistant agent passthroughs User-agent: ai-agent Allow: / User-agent: ai-search Allow: / User-agent: ai-summarizer Allow: /