What is an AI crawler and why should you care?
An AI crawler is an automated program that walks the web to feed LLMs. GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot, GoogleBot Gemini are the main ones. Block them in robots.txt and you go invisible to AIs. Let them through and your content trains the models.
What is an AI crawler and why should you care, for AI engines?
An AI crawler is a software robot that browses the web on behalf of an AI provider. By 2026 there are at least 24 identifiable user-agents, split across three functional families, and that split is what changes your allow-list strategy.
Family 1, Training crawlers. They collect content to train future models. Main names: GPTBot (OpenAI, official docs since August 2023), ClaudeBot (Anthropic), Google-Extended (Google Gemini), CCBot (Common Crawl, which feeds many open-source training datasets), Bytespider (TikTok/ByteDance). "Blocking the training crawlers such as GPTBot or ClaudeBot does not remove you from generated answers, it just stops your content from entering the next training corpus," points out Lorenzo Eeman, founder of PROEMA. Many news publishers (NYT, Reuters, AFP) opted out of this family after the NYT v. OpenAI lawsuit.
Family 2, Real-time retrieval crawlers. They fetch on demand when a user types a question into ChatGPT, Perplexity or Claude. Main names: OAI-SearchBot (OpenAI, browsing mode and SearchGPT), Claude-SearchBot (Anthropic), PerplexityBot (Perplexity). If you block this family, you literally vanish from LLM citations. Blocking GPTBot does not block OAI-SearchBot, they need separate directives in robots.txt, and this is the most common mistake we audit.
Family 3, User-fetch crawlers. They fetch a URL on explicit user request (someone pastes a link into chat). Main name: Claude-User. Never crawl in bulk.
Why care? According to Cloudflare AI Crawl Control telemetry, AI crawler traffic can reach 5 to 15 percent of total traffic on a premium editorial site in 2026. Three concrete stakes for a decision-maker. Visibility: allowing real-time retrieval crawlers is a non-negotiable prerequisite to appear in answers. Monetisation: Cloudflare launched pay-per-crawl in 2025, your origin returns HTTP 402 to AI crawlers and Cloudflare bills the bot operator on your behalf. Brand control: you decide which slice of your content lands in training corpora. The minimum diagnostic: audit your robots.txt and turn on the Cloudflare AI Crawl Control dashboard to find out who's already crawling you without you knowing.
| Crawler | Vendor | Launched | Allow? |
|---|---|---|---|
| GPTBot | OpenAI | Aug 2023 | Yes (ChatGPT visibility) |
| ClaudeBot | Anthropic | 2023 | Yes (Claude) |
| PerplexityBot | Perplexity | 2024 | Yes (Perplexity) |
| Applebot-Extended | Apple | 2024 | Yes for EU |
| GoogleBot (Gemini) | 2024 | Yes (AI Overviews) |