How to Block AI Scrapers and LLM Crawlers in Robots.txt
A definitive guide to managing AI user agents, preventing content scraping for model training, and configuring selective access directives.
Displayed below main page header or above the tool container. • Zero CLS Container
Quick Answer
To block AI scrapers from indexing or training on your content without impacting your organic search rankings, declare explicit User-agent blocks followed by Disallow: / in your root robots.txt file. Standard search engine bots (like Googlebot and Bingbot) must remain allowed, while dedicated training and retrieval bots (such as GPTBot, ClaudeBot, PerplexityBot, CCBot, and Google-Extended) can be selectively restricted. Directives are case-sensitive and must precede universal wildcard rules.
Generate & Validate Your Robots.txt Client-Side
Quickly toggle AI crawlers, validate syntax rules, and generate companion LLMs.txt files with zero telemetry.
The Complete AI User-Agent Reference Matrix
Use this comprehensive reference matrix to understand the organization, exact user-agent token, primary function, and SEO safety profile for each major AI crawler:
| Bot / Crawler | Organization | User-Agent Token | Primary Function | Safe to Block? |
|---|---|---|---|---|
| GPTBot | OpenAI | GPTBot | AI Model Training & Data Scraping | Yes (No impact on organic Google SEO) |
| ChatGPT-User | OpenAI | ChatGPT-User | Real-time browsing inside ChatGPT | Yes (Blocks live user browsing) |
| ClaudeBot | Anthropic | ClaudeBot | Anthropic Claude model training | Yes |
| Google-Extended | Google-Extended | Gemini & Vertex AI training data | Yes (Does NOT impact Google Search index) | |
| PerplexityBot | Perplexity AI | PerplexityBot | Real-time web indexation for answers | Yes |
| CCBot | Common Crawl | CCBot | Open-web repository used by multiple LLMs | Yes |
| Bytespider | ByteDance | Bytespider | TikTok & ByteDance AI crawling | Yes |
Robots.txt AI Directives Code Examples
Select between blocking all foundational training crawlers or implementing a selective policy that allows real-time conversational citations while rejecting bulk dataset scraping:
User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: PerplexityBot Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: * Allow: /
Snippet 1: Block All AI Training Scrapers
Disallow GPTBot, ClaudeBot, Google-Extended, PerplexityBot, CCBot, and Bytespider while allowing search engines via wildcard:
User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: PerplexityBot Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: * Allow: /
Snippet 2: Allow Citations & Block Training
Disallow bulk foundation scrapers (CCBot, GPTBot) while explicitly granting access to live search bots (ChatGPT-User, PerplexityBot):
User-agent: CCBot Disallow: / User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Allow: / User-agent: PerplexityBot Allow: /
Separates the interactive tool output from the deep technical guide. • Zero CLS Container
4-Step Implementation & Verification Protocol
Follow this verified engineering workflow to safely implement AI bot restrictions across your website without endangering organic search engine indexation:
- 1
Identify Target AI Agents
Determine whether you want to block raw model training only (GPTBot, ClaudeBot, CCBot, Bytespider, Google-Extended), or also restrict conversational AI search citations (such as ChatGPT-User and PerplexityBot). Distinguishing between model training and real-time citation retrieval is essential to avoid cutting off referral traffic from AI search engines.
Strategic Choice: If you want your content cited as an answer in ChatGPT Search or Perplexity, keepChatGPT-UserandPerplexityBotset toAllow: /. - 2
Append Directives to robots.txt
Place explicit AI bot blocks at the top of your
public/robots.txtfile or configure them via your CMS / framework settings (such as Next.js App Routerapp/robots.ts). Ensure that specific crawler tokens precede universal wildcard rules (User-agent: *) to guarantee deterministic parsing across all crawler implementations.Case Sensitivity: User-agent tokens are case-sensitive. Always writeGPTBot,ClaudeBot, andGoogle-Extendedwith exact casing. - 3
Verify Header Responses
Ensure your web server returns a valid
HTTP 200 OKstatus code andContent-Type: text/plainheader for the/robots.txtendpoint. If your web server returnsContent-Type: text/html, redirect chains (301/302), or HTTP 500 server errors, crawlers like Googlebot and GPTBot may treat the entire site as unreachable or completely disallowed.curl -IL https://yourdomain.com/robots.txt
HTTP/1.1 200 OK | Content-Type: text/plain; charset=utf-8 - 4
Test in Robots Validator
Run your complete file through our Robots.txt Validator to confirm there are no syntax errors, malformed path expressions, or accidental universal disallow directives. The tool parses rules client-side according to RFC 9309 standards and highlights conflicting permissions instantly.
Zero Leakage: Validating client-side ensures your proprietary robots rules and unpublished staging paths are never transmitted to external analytics servers.
Frequently Asked Questions
Understanding crawler behavior, search engine safety, and machine readability
Q1.Does blocking Google-Extended remove my site from Google Search?
No, Google-Extended only controls Gemini and AI training, whereas Googlebot handles search indexing. Google has explicitly verified that Google-Extended operates as a separate standalone token. Blocking Google-Extended has zero impact on your organic search rankings, indexing status, or search snippet visibility.
Q2.Do all AI companies honor robots.txt directives?
Robots.txt is voluntary; major players like OpenAI, Anthropic, and Google honor it, but smaller scrapers may ignore it. While reputable frontier AI labs strictly adhere to RFC 9309 robots directives, unverified web scrapers or anonymous botnets may disregard robots.txt entirely. For comprehensive defense, combine robots.txt with Cloudflare WAF bot management rules or IP rate limiting.
Q3.What is the difference between robots.txt and llms.txt?
Robots.txt restricts bot access, while llms.txt provides clean markdown context for AI models that are allowed. In other words, robots.txt serves as the security perimeter determining which bots may crawl your endpoints, whereas /llms.txt acts as a token-efficient semantic map and API directory for permitted AI search agents.
Ready to Build Your Custom Robots.txt?
Configure AI scraper blocks, validate path syntax, and export certified files in seconds.
Displayed below main page header or above the tool container. • Zero CLS Container