15 AI crawlers and robots.txt tokens from 6 operators, taken from each operator's own documentation and checked against their live IP feeds on 8 August 2026. They do three different jobs, and blocking the wrong one is how sites disappear from AI answers while believing they are protecting their content.
The three jobs are training, search indexing, and fetching a single page because a human asked. GPTBot trains; OAI-SearchBot answers. Blocking GPTBot keeps your writing out of the next model and leaves you visible in ChatGPT today. Blocking OAI-SearchBot is what removes you from the answers. Anthropic and Perplexity split the same way, and a blanket rule against "AI bots" catches both halves.
This page is the companion to our August 2026 crawler audit, which fetched 54 marketing sites as GPTBot and found nine that served an AI crawler nothing usable — none of it visible in robots.txt. The adoption figures below come from that study's corpus of 52 robots.txt files.
Training. Collects pages that may become training data for a future model. Blocking these keeps your writing out of the next model's weights. It does not remove you from answers, because modern assistants retrieve live pages at question time rather than reciting what they memorised. Publishers whose product is the writing itself have a real reason to block here. For most businesses this is the category where blocking costs least.
Search indexing. Builds the index an assistant searches when someone asks it a question. Blocking these is how you disappear from AI answers. The assistant has nothing of yours to retrieve, so it recommends whoever it can read instead. This is the expensive one, and it is the one people block by accident. Almost nobody should block this category. If you want to be mentioned when someone asks an assistant about your market, these are the crawlers that decide whether you can be.
User-triggered fetch. Fetches one page, right now, because a human asked the assistant about it. Blocking these breaks the moment a customer pastes your URL into ChatGPT and asks what you do. There is no crawl budget and no index involved — just a person waiting for an answer about you. Operators disagree on whether robots.txt even applies here, so a block may not be honoured. Treat these as traffic from a human, because that is what they are.
The practical consequence: people who write "User-agent: GPTBot / Disallow: /" to protect their content usually mean to block training, and that is exactly what they get. People who write a blanket rule against every AI bot also block the search crawlers, which is how a decision meant to protect content ends up handing a category's recommendations to a competitor.
User agent strings, purpose and robots.txt behaviour below are quoted from each operator's own documentation. Where an operator does not document something, this says so rather than filling the gap. Bytespider, Meta-ExternalAgent and Amazonbot are marked as lightly sourced because their operators publish far less than the others.
Respects robots.txt: Yes · Executes JavaScript: no · IP feed: openai.com/gptbot.json (21 prefixes) · Named in 10 of 52 robots.txt files sampled
Crawls content that may be used to train OpenAI's foundation models.
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
Respects robots.txt: Yes · Executes JavaScript: no · IP feed: openai.com/searchbot.json (35 prefixes) · Named in 7 of 52 robots.txt files sampled
Builds the index behind ChatGPT's search features. Not used for model training.
User agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
Respects robots.txt: May not apply · Executes JavaScript: no · IP feed: openai.com/chatgpt-user.json (258 prefixes) · Named in 3 of 52 robots.txt files sampled
Fetches a page live when a ChatGPT user or a Custom GPT asks about it.
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
OpenAI's documentation states that because the fetch is user-initiated, "robots.txt rules may not apply".
Respects robots.txt: Only visits submitted pages · Executes JavaScript: no · No published IP feed · Named in 0 of 52 robots.txt files sampled
Checks the safety of advertiser landing pages. Only visits pages submitted as ads, and the data is not used for training.
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot
Respects robots.txt: Yes · Executes JavaScript: no · IP feed: claude.com/crawling/bots.json (20 prefixes) · Named in 9 of 52 robots.txt files sampled
Collects web content for training and developing Claude.
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)
Respects robots.txt: Yes · Executes JavaScript: no · IP feed: claude.com/crawling/bots.json (20 prefixes) · Named in 0 of 52 robots.txt files sampled
Indexes content to improve the quality of Claude's search results.
User agent: Contains "Claude-SearchBot"; Anthropic documents the token rather than a full string.
Respects robots.txt: Yes · Executes JavaScript: no · IP feed: claude.com/crawling/bots.json (20 prefixes) · Named in 0 of 52 robots.txt files sampled
Fetches a page live when a Claude user asks about it.
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +Claude-User@anthropic.com)
The one user-triggered fetcher in this table whose operator says robots.txt is honoured. OpenAI and Perplexity both say theirs may not be.
Respects robots.txt: Yes · Executes JavaScript: no · IP feed: perplexity.com/perplexitybot.json (8 prefixes) · Named in 9 of 52 robots.txt files sampled
Surfaces and links websites in Perplexity search results. Perplexity states it "is not used to crawl content for AI foundation models".
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Respects robots.txt: Generally ignored · Executes JavaScript: no · IP feed: perplexity.com/perplexity-user.json (4 prefixes) · Named in 0 of 52 robots.txt files sampled
Visits a page when a Perplexity user asks a question that needs it.
User agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
Perplexity states it "generally ignores robots.txt rules" because a user requested the fetch.
Respects robots.txt: Token only · Executes JavaScript: n/a · IP feed: developers.google.com/.../googlebot.json (315 prefixes) · Named in 7 of 52 robots.txt files sampled
Controls whether your content trains future Gemini models and grounds answers in Gemini Apps and Vertex AI.
User agent: None. This is a robots.txt token only.
Google states it "doesn't have a separate HTTP request user agent string" — crawling happens under existing Google agents and the token is used "in a control capacity". Google also states it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal".
Respects robots.txt: Token only · Executes JavaScript: n/a · No published IP feed · Named in 5 of 52 robots.txt files sampled
Controls whether content already crawled by Applebot may be used to train Apple's foundation models.
User agent: None. This is a robots.txt token only.
Respects robots.txt: Yes · Executes JavaScript: no · IP feed: index.commoncrawl.org/ccbot.json (6 prefixes) · Named in 5 of 52 robots.txt files sampled
Builds the open Common Crawl dataset, which many AI labs use as a training source. Blocking CCBot blocks a source several models draw on at once.
User agent: CCBot/2.0 (https://commoncrawl.org/faq/)
Respects robots.txt: Claimed · Executes JavaScript: no · No published IP feed · Named in 3 of 52 robots.txt files sampled
ByteDance's crawler, widely reported as gathering training data. Documentation is thinner than the operators above.
User agent: Contains "Bytespider".
Respects robots.txt: Claimed · Executes JavaScript: no · No published IP feed · Named in 3 of 52 robots.txt files sampled
Meta's crawler for training data and product indexing. A companion token, meta-externalfetcher, covers user-triggered fetches.
User agent: Contains "meta-externalagent".
Respects robots.txt: Claimed · Executes JavaScript: no · No published IP feed · Named in 1 of 52 robots.txt files sampled
Amazon's crawler, used to improve Alexa and Amazon search answers.
User agent: Contains "Amazonbot".
A crawler reads exactly one group in your robots.txt: the most specific one that names it, or the wildcard group if none does. Groups do not merge, and they do not inherit.
This does not do what it looks like: a wildcard group disallowing /private/, followed by "User-agent: GPTBot" with "Allow: /". GPTBot now obeys only its own group. Because that group says nothing about /private/, GPTBot is free to crawl it — the wildcard rule the author thought applied to everyone no longer applies to GPTBot at all.
The version that works repeats the rule inside every group that names a crawler. A crawler reads exactly one group: the most specific one naming it, or the wildcard group if none names it. Groups do not merge and they do not inherit. The moment you name a crawler, you have taken responsibility for every rule it needs.
Our crawler audit found this pattern in the wild. Moz groups GPTBot with rogerbot and AhrefsBot and disallows all three from /blog/ and /learn/seo/ — a deliberate, correctly written rule. Any audit that only checks the site root reports it as "allowed", which is true and completely misses that the two libraries Moz is known for are closed.
Of the 52 robots.txt files collected for the crawler audit, only 10 named an AI crawler at all. Among those that did, the most-named tokens were GPTBot (10), ClaudeBot (9), PerplexityBot (9), OAI-SearchBot (7), Google-Extended (7).
Three sites still write rules for anthropic-ai and two for Claude-Web, both retired identifiers that Anthropic no longer crawls under. Not one names Claude-SearchBot, Claude-User, Perplexity-User or OAI-AdsBot, all of which are live. Claude-SearchBot is the crawler that decides whether Claude can cite you.
That is not carelessness so much as the shape of the problem: robots.txt is a file you write once and never get a reason to open again, while the operators behind it reorganise every few months.
Anything can put GPTBot in a header — our audit sent a user agent we invented and was served pages that real crawlers were refused. Five operators publish the ranges their crawlers actually use, and all of them have converged on Google's schema: a creationTime and a list of ipv4Prefix and ipv6Prefix objects, so one parser reads every feed.
Google (Googlebot): https://developers.google.com/static/search/apis/ipranges/googlebot.json — 315 prefixes, last regenerated 2026-08-07, 1 days before we fetched it.
OpenAI (ChatGPT-User): https://openai.com/chatgpt-user.json — 258 prefixes, last regenerated 2026-08-07, 1 days before we fetched it.
Common Crawl (CCBot): https://index.commoncrawl.org/ccbot.json — 6 prefixes, last regenerated 2026-08-04, 4 days before we fetched it.
Anthropic (ClaudeBot, Claude-User, Claude-SearchBot): https://claude.com/crawling/bots.json — 20 prefixes, last regenerated 2026-05-01, 99 days before we fetched it.
OpenAI (OAI-SearchBot): https://openai.com/searchbot.json — 35 prefixes, last regenerated 2026-01-02, 218 days before we fetched it.
OpenAI (GPTBot): https://openai.com/gptbot.json — 21 prefixes, last regenerated 2025-10-30, 282 days before we fetched it.
Perplexity (Perplexity-User): https://www.perplexity.com/perplexity-user.json — 4 prefixes, last regenerated 2025-10-17, 295 days before we fetched it.
Perplexity (PerplexityBot): https://www.perplexity.com/perplexitybot.json — 8 prefixes, last regenerated 2025-02-07, 547 days before we fetched it.
The cadence has not converged at all. Google's Googlebot ranges and OpenAI's ChatGPT-User feed had both been regenerated the day before we fetched them, and Common Crawl's was four days old; Perplexity's PerplexityBot feed was 547 days old. A stale feed does not mean the ranges are wrong, but an allow-list built strictly from one may reject legitimate traffic from addresses added since — an argument for verifying a request against the feed rather than allow-listing from it.
Anthropic is the one operator publishing a single feed covering all three of its bots. That confirms a request came from Anthropic but cannot tell ClaudeBot from Claude-User, so a policy of "allow retrieval, refuse training" cannot be enforced by IP there. robots.txt is the only lever, and Anthropic says it honours it — the one user-triggered fetcher in this list whose operator makes that commitment, where OpenAI says robots.txt "may not apply" to ChatGPT-User and Perplexity says Perplexity-User "generally ignores" it.
Everything here assumes the crawler that reaches you finds something to read. In our audit of 54 marketing sites that assumption failed nine times: four sites blocked or throttled a crawler by name at the CDN, and five served a page with almost no readable HTML because it was assembled in the browser after the crawler had already left.
Getting the robots.txt right is the cheap half. Serving real HTML to a client that does not run JavaScript is the half that decides whether any of it counts.
There is a third failure mode between the two: a rule at your CDN or WAF that refuses these crawlers by name regardless of what robots.txt says. Four sites in the audit corpus did exactly that. https://altheo.ai/cdn-blocking-ai-crawlers covers how to detect it from outside, where the rule lives, and what changes when Cloudflare resets its default AI-traffic handling on 15 September 2026.
Every user agent string, stated purpose and robots.txt claim above is quoted from one of six operator pages, all read on 8 August 2026.
OpenAI — Bots and crawlers: https://developers.openai.com/api/docs/bots — covers GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot.
Anthropic — Does Anthropic crawl data from the web?: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler — covers ClaudeBot, Claude-SearchBot, Claude-User.
Perplexity — PerplexityBot documentation: https://docs.perplexity.ai/guides/bots — covers PerplexityBot, Perplexity-User.
Google Search Central — Google crawlers overview: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers — covers Google-Extended, Googlebot.
Apple Support — About Applebot: https://support.apple.com/en-us/119829 — covers Applebot, Applebot-Extended.
Common Crawl — CCBot: https://commoncrawl.org/ccbot — covers CCBot.
Prefix counts and regeneration dates are not from documentation. Each IP feed was fetched directly and parsed, because a documented feed and a live one are different things — while checking, one URL that looked like a published feed turned out to be a documentation site's catch-all returning HTML for any path. Bytespider, Meta-ExternalAgent and Amazonbot are marked as lightly sourced because their operators publish far less than the others.
The ones worth knowing by name are OpenAI's GPTBot, OAI-SearchBot and ChatGPT-User; Anthropic's ClaudeBot, Claude-SearchBot and Claude-User; Perplexity's PerplexityBot and Perplexity-User; Common Crawl's CCBot; and two robots.txt tokens that have no crawler of their own, Google-Extended and Applebot-Extended. ByteDance's Bytespider, Meta-ExternalAgent and Amazonbot also appear in logs. They do three different jobs — training, search indexing, and fetching a page because a user asked — and treating them as one category is the mistake that costs the most.
No. GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and CCBot fetch a URL and read the HTML that comes back; they do not run a browser. Googlebot does render JavaScript, which is why a site can rank in Google Search and still be a blank page to an assistant. We measured this across 54 marketing sites in August 2026: nine of them served an AI crawler nothing usable, five because the page was assembled in the browser after the crawler had already left.
GPTBot collects content that may train OpenAI's future models. OAI-SearchBot builds the index behind ChatGPT's search features and is explicitly not used for training. The practical consequence is that blocking GPTBot keeps you out of the next model's weights, while blocking OAI-SearchBot removes you from the answers ChatGPT gives today. People who write "User-agent: GPTBot / Disallow: /" to protect their content usually mean the first and get only the first — which is fine. People who block both, or who block with a wildcard, get the second as well, and that is almost always a mistake.
Not by itself. GPTBot is the training crawler; ChatGPT's live answers are assembled by searching an index built by OAI-SearchBot and by fetching pages with ChatGPT-User. Blocking GPTBot alone keeps your writing out of future training runs while leaving you visible in answers. Blocking OAI-SearchBot is what makes you disappear. This is the single most useful distinction on this page, and the reason a blanket "block all AI bots" rule tends to achieve the opposite of what the person writing it wanted.
Check the source IP against the operator's published list, because a user agent header is just a string that anything can copy. OpenAI publishes a feed per bot (openai.com/gptbot.json, searchbot.json, chatgpt-user.json), Anthropic publishes one shared feed at claude.com/crawling/bots.json covering all three of its bots, Perplexity publishes perplexitybot.json and perplexity-user.json, Common Crawl publishes index.commoncrawl.org/ccbot.json, and Google publishes its Googlebot ranges. They all use the same schema — a creationTime and a list of ipv4Prefix and ipv6Prefix entries — so one parser reads all of them.
They vary enormously, and this matters if you are building allow-lists from them. Fetched on 8 August 2026, Google's Googlebot ranges and OpenAI's ChatGPT-User feed had both been regenerated the previous day, and Common Crawl's was four days old. Anthropic's shared feed was 99 days old, OpenAI's OAI-SearchBot feed 218 days, OpenAI's GPTBot feed 282 days, and Perplexity's PerplexityBot feed 547 days — over eighteen months. A stale feed does not mean the ranges are wrong, but it does mean an allow-list built strictly from it may reject legitimate traffic from addresses added since.
The operators disagree, and it is worth knowing which is which. OpenAI states that because ChatGPT-User acts on a user's request, "robots.txt rules may not apply". Perplexity states that Perplexity-User "generally ignores robots.txt rules" for the same reason. Anthropic is the exception: it says all three of its bots, Claude-User included, honour robots.txt. So a Disallow aimed at user-triggered fetching is reliable for Claude and unreliable elsewhere — which makes sense once you notice these requests are a person waiting for an answer about your page, not a crawler working through a queue.
Quite possibly, and in a specific way. Across the 52 robots.txt files we collected for our August 2026 crawler audit, only 10 named any AI crawler at all. Among those that did, three still wrote rules for anthropic-ai and two for Claude-Web — both retired identifiers that Anthropic no longer crawls under — while not one named Claude-SearchBot, Claude-User, Perplexity-User or OAI-AdsBot, all of which are live. The pattern is a file written once when a crawler made the news and never revisited as the operators split one bot into three.
It depends entirely on which job the crawler does, which is why the answer is never a single rule. If your product is the writing itself, blocking the training crawlers — GPTBot, ClaudeBot, CCBot, and the Google-Extended and Applebot-Extended tokens — is a defensible business decision with a real cost you can reason about. Blocking the search crawlers is different: it removes you from the answers assistants give about your market, and hands those mentions to whoever is still readable. And whichever you choose, decide it deliberately in robots.txt rather than discovering later that a bot-management default in your CDN chose for you.