Verdy.

Can AI Crawlers Read Your Website? How to Check and Allow GPTBot, ClaudeBot and PerplexityBot

Updated 13 min readBy the Verdy team

Short answer

To let AI assistants read your site, allow their search crawlers and user-triggered agents in robots.txt (OAI-SearchBot and ChatGPT-User for ChatGPT, Claude-SearchBot and Claude-User for Claude, PerplexityBot and Perplexity-User for Perplexity) and make sure your CDN or firewall isn't blocking them. Training crawlers such as GPTBot, ClaudeBot and CCBot are separate, so you can refuse training without disappearing from AI search. Then confirm access in your server logs, not just in robots.txt.

Key takeaways
  • AI bots come in three kinds: training crawlers, search crawlers and user-triggered fetchers. Each company names them separately.
  • Blocking GPTBot stops OpenAI training. It doesn't remove you from ChatGPT search. Blocking OAI-SearchBot does.
  • A bot with its own user-agent group in robots.txt ignores your User-agent star rules, so one stray group can undo your restrictions.
  • Firewalls and CDNs can block AI bots even when robots.txt allows them. Cloudflare's AI controls changed on September 15, 2026.
  • Test with your server logs. A robots.txt that says yes doesn't prove bots get in.
  • llms.txt is a proposal, not a standard, and Google Search ignores it. It's still a cheap file to add.

#The three kinds of AI bots

"AI crawler" covers three different jobs, and the companies that run them now name each one separately.

  • Training crawlers collect pages that may be used to train future models. Blocking one keeps your new content out of that company's training data. Examples: GPTBot, ClaudeBot, CCBot.
  • Search crawlers build the index an assistant searches when it answers with sources. Blocking one makes it hard for that assistant to cite you. Examples: OAI-SearchBot, Claude-SearchBot, PerplexityBot.
  • User-triggered fetchers visit a page because a person asked for it right now, for example by pasting your link into a chat. Examples: ChatGPT-User, Claude-User, Perplexity-User. Several companies say robots.txt may not apply to these.

Google and Apple add a fourth thing: control tokens. Google-Extended and Applebot-Extended aren't crawlers at all. They're names you use in robots.txt to say how content already fetched by Googlebot or Applebot may be used for AI.

#The AI user agents that matter (September 2026)

Each row below comes from the operating company's own documentation, checked on September 28, 2026. The last column reports what the company says, not what we've observed.

User agentCompanyWhat it's forTypeFollows robots.txt? (per its docs)
GPTBotOpenAICrawls content that may be used to train OpenAI's foundation modelsTrainingYes
OAI-SearchBotOpenAISurfaces sites in ChatGPT's search featuresSearchYes
ChatGPT-UserOpenAIVisits pages for user actions in ChatGPT and Custom GPTsUser-triggered"Robots.txt rules may not apply"
ClaudeBotAnthropicCollects content that could contribute to model trainingTrainingYes
Claude-SearchBotAnthropicIndexes content to improve Claude's search resultsSearchYes
Claude-UserAnthropicFetches pages when a Claude user asks a questionUser-triggeredYes
PerplexityBotPerplexitySurfaces and links sites in Perplexity; not used for foundation model trainingSearchYes
Perplexity-UserPerplexityVisits pages to answer a user's questionUser-triggered"Generally ignores robots.txt rules"
Google-ExtendedGoogleControls whether Google-crawled content is used for Gemini training and groundingControl tokenIt is a robots.txt rule
Google-AgentGoogleAgents on Google infrastructure acting on a user's requestUser-triggered"Generally ignore robots.txt rules"
ApplebotAppleSearch for Spotlight, Siri and Safari; data may also train Apple's modelsSearch and trainingYes
Applebot-ExtendedAppleOpts content out of training Apple's foundation modelsControl tokenIt is a robots.txt rule
meta-externalagentMetaTraining foundation AI models or indexing content for productsTraining and indexingYes, blocking via robots.txt is documented
meta-externalfetcherMetaFetches individual links at a user's requestUser-triggered"May bypass robots.txt rules"
CCBotCommon CrawlBuilds the free, open Common Crawl web archive that anyone can useOpen datasetYes, including Crawl-delay

A few details from those docs that change real decisions:

  • Google-Extended doesn't touch Google Search. Google says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." AI Overviews and AI Mode are part of Search: to appear, a page must be indexed and eligible for a snippet, and to limit it you use nosnippet, data-nosnippet, max-snippet or noindex. Blocking Googlebot to avoid AI Overviews removes you from Google entirely.
  • Applebot-Extended doesn't touch Apple search. Apple says pages that disallow it "can still be included in search results." Also, if your robots.txt doesn't mention Applebot but does mention Googlebot, Applebot follows the Googlebot rules.
  • OpenAI may reuse one crawl. If you allow both GPTBot and OAI-SearchBot, OpenAI says it may use one crawl for both purposes.
  • Microsoft has no AI-specific robots.txt token yet. Cloudflare reports one is planned for early 2027. Microsoft's 2023 guidance: the noarchive meta tag keeps content out of its chat answers and model training, and nocache limits use to URL, title and snippet, with pages still in Bing search. Check Bing's current docs before relying on it.
  • User agents can be spoofed. The companies publish IP lists so you can verify real traffic (links in the testing section below).

#Which AI bots should you allow?

If you want to be found, allow the search crawlers and user-triggered fetchers. OpenAI says sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers, though can still appear as navigational links." Anthropic says disabling Claude-SearchBot "may reduce your site's visibility and accuracy in user search results."

Training is a real judgment call. Nobody can show you exactly how training data turns into recommendations. For a new product that needs awareness more than protection, we lean toward allowing training too. For publishers whose writing is the product, blocking training while allowing search is a sensible middle ground, and the companies above have made it possible by splitting their bots. For the wider picture of getting cited, see how to get recommended by ChatGPT and AI search.

#robots.txt examples

Grouping several User-agent lines above one set of rules is valid under the robots.txt standard (RFC 9309). If you prefer, give each bot its own block.

#Allow AI search and user requests, block training

User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
Disallow: /

User-agent: *
Allow: /

Sitemap: https://yourdomain.com/sitemap.xml

meta-externalagent covers both training and product indexing, so blocking it blocks both. And blocking Google-Extended or Applebot-Extended leaves Google and Apple search untouched.

#Allow everything

User-agent: *
Allow: /

Sitemap: https://yourdomain.com/sitemap.xml

A bot with no group of its own follows the * group, so this is all you need. Just make sure no leftover bot-specific groups sit further down the file.

#Block all AI bots

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: meta-externalfetcher
User-agent: CCBot
Disallow: /

This leaves regular Googlebot, Bingbot and Applebot search crawling alone. Remember that OpenAI, Perplexity, Google and Meta say their user-triggered fetchers may not follow robots.txt, so if you truly need them out, you'll have to block them at the firewall.

#The trap: a specific group replaces the star group

Under RFC 9309, a crawler obeys the group that names it. It only falls back to User-agent: * if no group matches. Google's documentation describes the same behaviour: only one group applies. So this file lets GPTBot into /admin/ and /drafts/:

User-agent: *
Disallow: /admin/
Disallow: /drafts/

User-agent: GPTBot
Allow: /

If you add a group for any bot, copy your Disallow lines into it. And don't use robots.txt to protect anything private: RFC 9309 says plainly that these rules "are not a form of access authorization." Use a login.

#Check your CDN and firewall too

robots.txt is a request. A firewall, CDN or bot-protection setting can block an AI crawler outright, even when robots.txt says yes. The bot gets a 403, a challenge page or a timeout, and you never find out unless you look.

#Cloudflare (changed September 15, 2026)

If you're on Cloudflare, check its AI settings. Per Cloudflare's blog posts of July 1 and September 15, 2026:

  • AI traffic is managed in three categories: Search, Training and Agent (user-directed agents such as chat fetch bots). Each can be set to Allow, Block on pages with ads, or Block. Training also gets Disallow AI Training, which publishes a no-training rule in your robots.txt and blocks training-only crawlers while keeping accountable mixed-use crawlers allowed for search.
  • New domains onboarding since September 15 are offered presets. Sites that don't show ads: everything allowed. Sites that show ads: Search allowed, Training set to Disallow AI Training, Agent blocked on pages with ads.
  • The old one-click "Block AI Bots" setting is being deprecated. Existing choices migrate: a legacy Block becomes Search allowed, Training set to Disallow AI Training, Agent blocked on pages with ads.
  • The trap: Block and Block on pages with ads now also apply to mixed-use crawlers, which Cloudflare says include Googlebot, Bingbot and Applebot. Choosing Block for Training can stop them crawling for search as well. To refuse training and keep search, pick Disallow AI Training.
  • Managed robots.txt is being replaced by Bot Preference Sync, which writes preferences into your robots.txt. Your live file may contain lines you didn't write, so read it.

Cloudflare's AI Crawl Control dashboard, available on all plans according to its docs, shows which AI crawlers visit and whether they follow your robots.txt.

#Other firewalls and hosts

If you allowlist bots in a WAF, match both the user agent and the company's published IP ranges. Perplexity's crawler docs walk through exactly this for Cloudflare WAF and AWS WAF. For opting out, prefer robots.txt: Anthropic warns that blocking its IPs "may not work correctly or persistently guarantee an opt-out, as doing so impedes our ability to read your robots.txt file." Bot challenges matter too. Anthropic says its bots won't try to bypass CAPTCHAs, so a challenge page is effectively a block.

#How to test AI crawler access

#1. Fetch robots.txt

Run curl -s https://yourdomain.com/robots.txt. You want a 200 and plain-text rules. If you get your app's HTML instead, a missing file is being answered by your frontend. Then read it as each bot would: find the group that names the bot, and if there's none, it follows *. Check subdomains separately, since Anthropic notes opt-outs apply per subdomain.

#2. Request a page as the bot

curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" https://yourdomain.com/

A 200 is a good sign. A 403, 429 or 503 suggests a block. Treat it as a hint only: firewalls that verify bots by IP may treat your spoofed request differently from the real crawler.

#3. Check your server logs

This is the real test:

grep -E "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|CCBot" access.log | awk '{print $9}' | sort | uniq -c

In the common Nginx and Apache log format, field 9 is the status code, so lots of 403s means something is blocking. To confirm a visitor is genuine, compare its IP with the published lists: openai.com/gptbot.json, openai.com/searchbot.json, openai.com/chatgpt-user.json, claude.com/crawling/bots.json, perplexity.com/perplexitybot.json, perplexity.com/perplexity-user.json and index.commoncrawl.org/ccbot.json.

#4. Check the HTML has your content

Google's JavaScript guide notes that "not all bots can run JavaScript." If curl shows an empty shell, a crawler that doesn't render may see nothing. Our guide on why a new site isn't showing up on Google covers the rendering fix.

#5. Give changes a day

OpenAI says search changes take about 24 hours, Perplexity and Meta say up to 24 hours, and Google caches robots.txt for up to 24 hours. Add this check to your website launch checklist so it runs on every launch.

Verdy's free audit checks AI crawler access in your robots.txt (GPTBot, ClaudeBot, PerplexityBot and others) and whether you have an llms.txt, and the AI-readiness check in our free tools needs no signup. Both read robots.txt, so pair them with the log check to catch firewall blocks.

#llms.txt: what it is and what it's worth

#What it is

llms.txt is a proposal by Jeremy Howard, first published on September 3, 2024 and now in a second version. It's a Markdown file at /llms.txt (or at a subpath like /docs/llms.txt) that gives AI agents a short, curated map of your site: a name, a one-paragraph summary, and lists of links to the pages that matter, ideally pointing to clean Markdown versions. Only the H1 title is required. robots.txt controls access. llms.txt describes content.

#The honest status

  • It's a proposal, not a standard. No standards body publishes it, unlike robots.txt, which the IETF published as RFC 9309.
  • Google Search ignores it. Google's AI optimization guide (updated July 2026) says: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." It adds that maintaining llms.txt for other services is "completely fine" and "will neither harm nor help" your Google visibility.
  • Google's John Mueller was blunter. In an April 2025 Reddit thread, reported by Search Engine Journal, he said that as far as he knew no AI service had said it uses llms.txt, and called it "comparable to the keywords meta tag."
  • The crawler docs don't mention it. None of the OpenAI, Anthropic or Perplexity crawler pages we checked say their bots use llms.txt to decide what to cite.
  • But it's not nothing. Chrome's Lighthouse has an experimental Agentic Browsing category with an llms.txt audit. A missing file is marked not applicable and a server error gets flagged. Lighthouse says that without one, agents "may spend more time crawling the site to understand its high-level structure." OpenAI, Anthropic, Perplexity and Google's Gemini API all publish llms.txt files for their own developer docs.

#Why add one anyway

It costs about twenty minutes and one static file, and Google says it won't hurt you in Search. Agents and tools that do fetch it get a clean summary instead of parsing your navigation. And a precise two-sentence description of your product is worth writing anyway: it's the clarity AI answers need from your homepage. Just don't expect it to move rankings or citations on its own.

A minimal example. The file opens with a single H1 title line, # Acme Invoices here, followed by this (serve it as plain text):

> Acme Invoices is invoicing software for freelancers in the EU. It creates VAT-compliant invoices, tracks payments and sends reminders.

## Product

- [Features](https://acme.example/features): what the app does
- [Pricing](https://acme.example/pricing): plans and limits

## Docs

- [Getting started](https://acme.example/docs/start.md): setup in five steps

## Optional

- [Blog](https://acme.example/blog): product updates

Verdy's llms.txt starter generator in the free tools drafts one from your site.

#Common mistakes

  • Mixing up GPTBot and OAI-SearchBot. Blocking GPTBot doesn't remove you from ChatGPT search. Blocking OAI-SearchBot does.
  • A bot-specific group that drops your rules. That bot ignores everything under User-agent: *.
  • Picking Block for Training on Cloudflare. It can now stop Googlebot, Bingbot and Applebot too.
  • Allowlisting by user agent alone. Anyone can send "GPTBot". Match the published IP ranges.
  • Treating robots.txt as security. It's a request, not access control.
  • Expecting llms.txt to rank you. Google Search ignores it.

#FAQ

#Can ChatGPT crawl my website?

Yes, unless you block it. OpenAI uses OAI-SearchBot to surface sites in ChatGPT search, ChatGPT-User to fetch pages when a user asks, and GPTBot for training. Allow OAI-SearchBot in robots.txt and make sure your firewall isn't blocking OpenAI's published IP ranges.

#Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot only covers training. OpenAI says each setting is independent, so you can disallow GPTBot and still appear in ChatGPT search as long as OAI-SearchBot is allowed. Blocking OAI-SearchBot is what removes you from ChatGPT search answers.

#What's the difference between GPTBot and ChatGPT-User?

GPTBot crawls automatically for content that may be used to train OpenAI's models. ChatGPT-User visits a page only when a ChatGPT user's action calls for it, and OpenAI says robots.txt rules may not apply to it because a person initiated the request.

#Does Google-Extended block AI Overviews?

No. Google-Extended controls whether content is used for Gemini training and grounding, and Google says it doesn't affect inclusion or ranking in Google Search. AI Overviews are part of Search, controlled with Googlebot access and snippet rules like nosnippet and max-snippet.

#Do AI crawlers respect robots.txt?

The major companies say their training and search crawlers do, including OpenAI, Anthropic, Perplexity, Apple, Meta and Common Crawl. User-triggered fetchers are different: OpenAI, Perplexity, Google and Meta say theirs may not follow robots.txt. Anthropic says all three of its bots do.

#Does Cloudflare block AI crawlers by default?

Not across the board. Since September 15, 2026, new Cloudflare domains are offered presets: sites without ads allow everything, while ad-supported sites disallow AI training and block agents on pages with ads. Existing settings migrate automatically, so check your Search, Training and Agent controls.

#Is llms.txt worth adding?

It's cheap, so usually yes, as long as expectations are low. It's a proposal rather than a standard, and Google Search ignores it. It can help agents and developer tools that read it, and Chrome's Lighthouse checks for it in an experimental audit.

#Final recommendation

Allow the search crawlers and user-triggered fetchers, decide on training deliberately, and keep your * rules copied into any bot-specific group. Then check your CDN settings, especially on Cloudflare after the September 2026 changes, and confirm in your logs that real bots get 200s. Add an llms.txt if you like, but spend more time on what crawlers actually read: fast, server-rendered pages that say clearly what you do. To see where you stand, run the free audit for a robots.txt and llms.txt check in one pass, or compare options in our AI visibility tools roundup.

Check which AI crawlers your robots.txt lets in, and whether you have an llms.txt, with a free Verdy audit.

Run a free audit

Prices, plans and platform rules change. Anything current in this guide was checked on September 28, 2026; confirm on the vendor's own site before you buy.