AI & Crawlers4 min readUpdated September 2026RFC 9309 Verified

How to Block AI Scrapers and LLM Crawlers in Robots.txt

A definitive guide to managing AI user agents, preventing content scraping for model training, and configuring selective access directives.

Advertisement
AdSense Placeholder: Top Leaderboard Ad (728x90 / 320x50)

Displayed below main page header or above the tool container. • Zero CLS Container

Quick Answer

To block AI scrapers from indexing or training on your content without impacting your organic search rankings, declare explicit User-agent blocks followed by Disallow: / in your root robots.txt file. Standard search engine bots (like Googlebot and Bingbot) must remain allowed, while dedicated training and retrieval bots (such as GPTBot, ClaudeBot, PerplexityBot, CCBot, and Google-Extended) can be selectively restricted. Directives are case-sensitive and must precede universal wildcard rules.

Generate & Validate Your Robots.txt Client-Side

Quickly toggle AI crawlers, validate syntax rules, and generate companion LLMs.txt files with zero telemetry.

The Complete AI User-Agent Reference Matrix

Use this comprehensive reference matrix to understand the organization, exact user-agent token, primary function, and SEO safety profile for each major AI crawler:

Bot / CrawlerOrganizationUser-Agent TokenPrimary FunctionSafe to Block?
GPTBotOpenAIGPTBotAI Model Training & Data ScrapingYes (No impact on organic Google SEO)
ChatGPT-UserOpenAIChatGPT-UserReal-time browsing inside ChatGPTYes (Blocks live user browsing)
ClaudeBotAnthropicClaudeBotAnthropic Claude model trainingYes
Google-ExtendedGoogleGoogle-ExtendedGemini & Vertex AI training dataYes (Does NOT impact Google Search index)
PerplexityBotPerplexity AIPerplexityBotReal-time web indexation for answersYes
CCBotCommon CrawlCCBotOpen-web repository used by multiple LLMsYes
BytespiderByteDanceBytespiderTikTok & ByteDance AI crawlingYes

Robots.txt AI Directives Code Examples

Select between blocking all foundational training crawlers or implementing a selective policy that allows real-time conversational citations while rejecting bulk dataset scraping:

Explicitly blocks major LLM foundation model training crawlers while keeping standard search engines (Googlebot, Bingbot) and universal user-agents allowed.Content-Type: text/plain
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: *
Allow: /
Want to toggle AI crawlers interactively and test syntax?

Snippet 1: Block All AI Training Scrapers

Zero Scraping

Disallow GPTBot, ClaudeBot, Google-Extended, PerplexityBot, CCBot, and Bytespider while allowing search engines via wildcard:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: *
Allow: /

Snippet 2: Allow Citations & Block Training

Citation Friendly

Disallow bulk foundation scrapers (CCBot, GPTBot) while explicitly granting access to live search bots (ChatGPT-User, PerplexityBot):

User-agent: CCBot
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: ChatGPT-User
Allow: /

User-agent: PerplexityBot
Allow: /
Advertisement
AdSense Placeholder: Native In-Feed Ad (Responsive)

Separates the interactive tool output from the deep technical guide. • Zero CLS Container

4-Step Implementation & Verification Protocol

Follow this verified engineering workflow to safely implement AI bot restrictions across your website without endangering organic search engine indexation:

  1. 1

    Identify Target AI Agents

    Determine whether you want to block raw model training only (GPTBot, ClaudeBot, CCBot, Bytespider, Google-Extended), or also restrict conversational AI search citations (such as ChatGPT-User and PerplexityBot). Distinguishing between model training and real-time citation retrieval is essential to avoid cutting off referral traffic from AI search engines.

    Strategic Choice: If you want your content cited as an answer in ChatGPT Search or Perplexity, keep ChatGPT-User and PerplexityBot set to Allow: /.
  2. 2

    Append Directives to robots.txt

    Place explicit AI bot blocks at the top of your public/robots.txt file or configure them via your CMS / framework settings (such as Next.js App Router app/robots.ts). Ensure that specific crawler tokens precede universal wildcard rules (User-agent: *) to guarantee deterministic parsing across all crawler implementations.

    Case Sensitivity: User-agent tokens are case-sensitive. Always write GPTBot, ClaudeBot, and Google-Extended with exact casing.
  3. 3

    Verify Header Responses

    Ensure your web server returns a valid HTTP 200 OK status code and Content-Type: text/plain header for the /robots.txt endpoint. If your web server returns Content-Type: text/html, redirect chains (301/302), or HTTP 500 server errors, crawlers like Googlebot and GPTBot may treat the entire site as unreachable or completely disallowed.

    curl -IL https://yourdomain.com/robots.txt
    HTTP/1.1 200 OK | Content-Type: text/plain; charset=utf-8
  4. 4

    Test in Robots Validator

    Run your complete file through our Robots.txt Validator to confirm there are no syntax errors, malformed path expressions, or accidental universal disallow directives. The tool parses rules client-side according to RFC 9309 standards and highlights conflicting permissions instantly.

    Zero Leakage: Validating client-side ensures your proprietary robots rules and unpublished staging paths are never transmitted to external analytics servers.

Frequently Asked Questions

Understanding crawler behavior, search engine safety, and machine readability

Q1.Does blocking Google-Extended remove my site from Google Search?

No, Google-Extended only controls Gemini and AI training, whereas Googlebot handles search indexing. Google has explicitly verified that Google-Extended operates as a separate standalone token. Blocking Google-Extended has zero impact on your organic search rankings, indexing status, or search snippet visibility.

Q2.Do all AI companies honor robots.txt directives?

Robots.txt is voluntary; major players like OpenAI, Anthropic, and Google honor it, but smaller scrapers may ignore it. While reputable frontier AI labs strictly adhere to RFC 9309 robots directives, unverified web scrapers or anonymous botnets may disregard robots.txt entirely. For comprehensive defense, combine robots.txt with Cloudflare WAF bot management rules or IP rate limiting.

Q3.What is the difference between robots.txt and llms.txt?

Robots.txt restricts bot access, while llms.txt provides clean markdown context for AI models that are allowed. In other words, robots.txt serves as the security perimeter determining which bots may crawl your endpoints, whereas /llms.txt acts as a token-efficient semantic map and API directory for permitted AI search agents.

Ready to Build Your Custom Robots.txt?

Configure AI scraper blocks, validate path syntax, and export certified files in seconds.

Open Robots.txt Validator
Advertisement
AdSense Placeholder: Top Leaderboard Ad (728x90 / 320x50)

Displayed below main page header or above the tool container. • Zero CLS Container