/ THE SHORT ANSWER
OAI-SearchBot is used for search discovery and surfacing public web content in ChatGPT search. GPTBot relates to potential model training. A publisher can make different decisions for each user agent in robots.txt. If a page should not appear in search at all, use a noindex directive while allowing the relevant crawler to access the page long enough to read it; blocking crawling alone is not a reliable substitute for noindex.
- 01OAI-SearchBot and GPTBot serve different purposes.
- 02Robots.txt and noindex solve different control problems.
- 03Document policy, owner, intended outcome, and review date.
/ dotSuper point of view
Crawler policy is a publishing decision, not a binary opinion about AI. Separate search visibility, model training, indexing, and private-content controls.
What the evidence says
OpenAI’s official publisher FAQ distinguishes OAI-SearchBot access for ChatGPT search from GPTBot controls for content that publishers wish to exclude from potential training.
OpenAI notes that a noindex directive may be needed to prevent a disallowed URL from surfacing as a link and title, and that the crawler must be allowed to read the directive.
A practical decision framework
The following framework is dotSuper’s operating synthesis of the cited guidance. It is designed to make the decision inspectable, not to imitate a platform ranking formula, certification checklist, or legal test.
- Public and discoverable: allow search crawling, permit indexing, and publish complete content.
- Public but not discoverable: allow crawler access to read noindex, then verify removal.
- Searchable but excluded from potential training: allow OAI-SearchBot and disallow GPTBot.
- Private: require authentication and avoid relying on robots.txt as access control.
| Step | Decision to record |
|---|---|
| 01 | Public and discoverable: allow search crawling, permit indexing, and publish complete content. |
| 02 | Public but not discoverable: allow crawler access to read noindex, then verify removal. |
| 03 | Searchable but excluded from potential training: allow OAI-SearchBot and disallow GPTBot. |
| 04 | Private: require authentication and avoid relying on robots.txt as access control. |
How to put it into practice
Create a crawler-policy table for site sections and user agents. Review it with publishing, legal, security, and growth owners before deployment.
Test the served robots.txt and page-level directives from the public origin. Recheck after CDN, framework, or domain changes, and avoid contradictory signals across environments.
- Name the accountable owner and the decision this work must enable.
- Record the current evidence, assumptions, exclusions, and next review trigger.
- Measure a useful outcome rather than treating publication or deployment as success.
What this page cannot conclude
- 01Crawler names, documentation, and product behaviour can change and should be reviewed against current official guidance.
- 02Robots directives are not security boundaries and do not make sensitive information private.
- 03Publication, technical eligibility, or good practice cannot guarantee ranking, referral traffic, citation, adoption, or a business outcome.
Sources
Turn useful expertise into an inbound system.
dotSuper connects research, evidence-led pages, technical discovery, AI visibility, analytics, and qualified lead routing as one controlled operating loop.
Explore the Inbound Engine