GEO

ROBOTS POLICY · OFFICIAL-SOURCE MATRIX

Decide which AI crawlers your robots.txt should address.

Separate search discovery, model training, and user-triggered retrieval, then generate only the groups current vendor documentation says robots.txt can control.

Start from an operating posture

A preset fills all controllable tokens. You can change each directive before copying.

CONTROL MATRIX

One token, one explicit decision

Only rows marked controllable enter the draft. Advisory rows explain user-triggered behavior but never generate misleading rules.

OpenAIOAI-SearchBot
Search discoveryControllable

OpenAI search crawler used to surface and link sites in ChatGPT search.

Independent from GPTBot; allowing it does not allow training crawling.

OAI-SearchBot: Your directive
OpenAIGPTBot
Model trainingControllable

OpenAI crawler for content that may be used to train foundation models.

A disallow is a signal about future content; it is separate from ChatGPT search inclusion.

GPTBot: Your directive
OpenAIChatGPT-User
User retrievalAdvisory only

User-triggered OpenAI fetcher, not an automatic web crawler.

OpenAI says robots.txt rules may not apply; it does not determine ChatGPT search inclusion, so no group is generated.

Advisory only
AnthropicClaudeBot
Model trainingControllable

Anthropic crawler for web content that may contribute to model training.

Anthropic currently says its bots respect robots.txt directives.

ClaudeBot: Your directive
AnthropicClaude-SearchBot
Search discoveryControllable

Anthropic search crawler used to improve search result quality.

Keep this decision separate from ClaudeBot training access.

Claude-SearchBot: Your directive
AnthropicClaude-User
User retrievalControllable

Anthropic user-triggered fetcher for a user request.

Anthropic currently documents robots.txt control; blocking can reduce user-directed web retrieval.

Claude-User: Your directive
GoogleGoogle-ExtendedControl token, not HTTP user agent
Training + groundingControllable

Google control token for Gemini training and grounding with Google Search.

It has no separate HTTP user agent and does not affect Google Search inclusion or ranking.

Google-Extended: Your directive
PerplexityPerplexityBot
Search discoveryControllable

Perplexity search crawler used to surface and link websites.

Perplexity says it is not used for foundation-model training.

PerplexityBot: Your directive
PerplexityPerplexity-User
User retrievalAdvisory only

Perplexity user-triggered fetcher, not a crawling or training bot.

Perplexity says it generally ignores robots.txt, so no group is generated.

Advisory only

Deployment boundary

The generated text is deliberately narrow. Treat it as a review artifact, not a drop-in replacement.

  1. 01Merge the draft into your existing robots.txt; replacing the file can erase unrelated search directives.
  2. 02Verify vendor documentation again before release. Names, purposes, and robots behavior can change.
  3. 03A WAF, CDN, hosting rule, or IP block can still deny a crawler that robots.txt allows.
  4. 04robots.txt is neither an access guarantee nor a promise of indexing, model visibility, mentions, or citations. User-triggered fetchers can behave differently.

Before you publish

Does Allow guarantee an AI engine will access or cite the site?

No. robots.txt is a crawler preference signal. Network controls, vendor systems, relevance, and product behavior remain separate.

Why are ChatGPT-User and Perplexity-User not in the generated text?

Their official documentation describes user-triggered behavior where robots.txt may not apply or is generally ignored. Showing a rule would overstate control.

Can I replace my current robots.txt with this output?

No. Merge only the reviewed groups into the current file, preserve existing search and asset rules, then test the deployed response.