GUIDE № 06LAST UPDATED · 7 MIN READ

Technical

AI crawlers and robots.txt: which to allow, which to block?

AI companies run separate crawlers for separate purposes: one to train models, one to cite sources in search answers, one for when a user asks to open a link. Blocking them all in one line can quietly remove you from AI search too.

Written by: Vextoo Geo Agency Team · Last updated:

✳ Short answer

If you want to appear in AI answers, allow the search/citation crawlers: OAI-SearchBot (ChatGPT search)[1], Claude-SearchBot (Claude)[2], PerplexityBot (Perplexity)[3], Bingbot (Copilot) and Googlebot (AI Overviews). Training crawlers (GPTBot, ClaudeBot, Google-Extended[4], meta-externalagent[5], CCBot) are a separate decision; blocking them limits use of your content in model training but doesn’t directly cut visibility in search answers.

What does each crawler do?

CrawlerCompanyPurposeOur advice
OAI-SearchBotOpenAI[1]ChatGPT search resultsAllow
ChatGPT-UserOpenAIUser-requested visitsAllow
GPTBotOpenAIModel trainingYour choice
Claude-SearchBotAnthropic[2]Claude search indexAllow
Claude-UserAnthropicUser-requested visitsAllow
ClaudeBotAnthropicModel trainingYour choice
PerplexityBotPerplexity[3]Perplexity search indexAllow
Perplexity-UserPerplexityUser-requested visitsAllow
GooglebotGoogleSearch + AI Overviews / AI ModeAllow
Google-ExtendedGoogle[4]Gemini training and grounding in Gemini appsYour choice
BingbotMicrosoftBing index → CopilotAllow
meta-webindexerMeta[5]Meta AI search qualityAllow
meta-externalagentMetaTraining / direct indexingYour choice
CCBotCommon CrawlOpen web archive (training data for many models)Your choice

Sample robots.txt

A configuration open to search and citation crawlers and closed to training crawlers (adjust to your own preference):

# Search and citation crawlers: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: meta-webindexer
Allow: /

# Model training crawlers: blocked (optional)
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /

# Everything else
User-agent: *
Allow: /
Disallow: /admin/

Sitemap: https://example.com/sitemap.xml

This site’s own file is open to all AI crawlers: robots.txt — because being quoted is the point.

robots.txt is clean, but the bot may still be locked out

  • CDN / firewall: “block AI bots” settings and bot-management rules at providers like Cloudflare work independently of robots.txt. Check your dashboard.
  • JavaScript-dependent content: even when the bot gets in, text loaded by JavaScript often looks like an empty page.
  • Login walls and cookie barriers: pop-ups that fully cover content can break parsing.
  • Spoofed bot traffic: bad actors can fake these names; companies like OpenAI and Perplexity publish IP ranges for verification[1][3].

Check all of it in seconds: GEO Readiness Check.

Frequently asked questions

01Should I block GPTBot?

It’s a choice. GPTBot is for model training; according to OpenAI, ChatGPT search visibility is governed by OAI-SearchBot. If you don’t want your content used for training, block GPTBot and allow OAI-SearchBot.

02Does robots.txt block user-initiated bots?

Not always. OpenAI says robots.txt rules may not apply to ChatGPT-User; Perplexity says Perplexity-User generally ignores robots.txt. Anthropic says all three of its bots, including Claude-User, honor robots.txt.

Sources

  1. OpenAI — Overview of OpenAI crawlers · accessed 27 September 2026
  2. Anthropic — Claude crawlers and robots.txt · accessed 27 September 2026
  3. Perplexity — Perplexity Crawlers · accessed 27 September 2026
  4. Google — Google’s common crawlers (Google-Extended) · accessed 27 September 2026
  5. Meta — Meta Web Crawlers · accessed 27 September 2026

Next step

Where does your brand stand inside AI answers?

Our free preliminary scan maps your brand across 6 AI engines — we come back within 48 hours.