Your service pages are meant to be discovered. Our position: welcome AI access, including for training. That creates an opportunity for learning, without guaranteeing a model will remember or recommend your business.
Start by checking robots.txt. But first, understand that AI crawlers do different jobs.
What does robots.txt do?
A crawler is a program that automatically visits web pages. A robots.txt file tells crawlers that follow the protocol which paths they may explore.
Read it at https://example.com/robots.txt. Subdomains, such as your store, have their own rules. Google’s file-creation guide
This public file does not protect a client portal or guarantee removal from search results. Private areas need proper access controls. Google’s introduction
AI-related agents do different jobs
There are three useful categories:
| Use | What happens | Examples |
|---|---|---|
| Training | Collect content that may contribute to model development | GPTBot, ClaudeBot |
| Search | Explore the web to support search features | OAI-SearchBot, Claude-SearchBot |
| User request | Fetch a page when someone uses an assistant | ChatGPT-User, Claude-User |
Rules differ: ChatGPT-User may not follow robots.txt; Anthropic says Claude-User does. OpenAI, Anthropic
How to read a rule without being a developer
Here is a deliberately fictional example:
User-agent: ExampleBot
Disallow: /archives/
User-agent names the crawler. Disallow asks it not to explore the specified path, here /archives/ and anything below it. Allow permits a path; / covers the root and paths below it. RFC 9309
Three examples that show the difference
These fragments explain your options. Do not paste them over your file: have its complete rules reviewed before changing anything.
1. Allow OpenAI training and search
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
This allows training collection and ChatGPT search crawling. The controls are independent: replacing the first Allow: / with Disallow: / would refuse GPTBot without changing this search group. OpenAI documentation
2. Understand Anthropic’s training opt-out option
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
For Anthropic, this refuses training collection while allowing search and user-requested visits. It is an option, not our default recommendation for public pages. Anthropic documentation
3. Recognize a much broader restriction
User-agent: *
Disallow: /
In a file containing only these lines, all crawlers following this protocol are told not to explore the site. This is not a training-only restriction. In a longer file, you must also inspect agent-specific groups. RFC 9309
Why Google-Extended needs its own explanation
Google-Extended is not a separate HTTP crawler. It is a control token covering use of already-crawled content for Gemini training and certain Gemini answers grounded with Google Search.
User-agent: Google-Extended
Disallow: /
This rule therefore does not mean “refuse training only.” It does not affect the site’s inclusion or ranking in Google Search, and it should not be confused with blocking Googlebot. Google documentation
What about llms.txt?
llms.txt offers AI agents a Markdown guide: an introduction and useful links. It already has uses, including software documentation. It is not training permission. The llms.txt proposal
Google says it does not improve Search visibility, including in its generative features. Google’s official guide
Our advice: once your pages are clear and accessible, a simple, maintained file can be a reasonable small investment. Introduce your business and link to your services, service area and public contact details. Adoption could broaden, with no guaranteed timeline. Allow for some upkeep; the pages themselves come first.
A checklist for your own website
Start by observing, not editing:
- Open
/robots.txton your site’s exact domain. - Record complete groups and classify their uses: training, search or user-requested visits.
- Open
/llms.txttoo. If it exists, check the business description, links and information. Its absence at that address does not mean your site blocks AI.
If the file fails to load, record “not verified,” not “AI blocked.”
A trap to check before pasting an example
An agent-specific group does not necessarily inherit restrictions from *. Have exclusions checked before adding Allow: /. Google’s group-selection rules
You can do this first reading yourself without changing anything. It does not verify every real access condition: a firewall can block requests that robots.txt permits. If you need to investigate that or change a rule, review hosting and existing restrictions too. Google’s AI features guide
Next article: your site is accessible, but what can AI actually read?
Want an initial look at your website?
Free BREACH examines your website’s public signals, not assistant recommendations.
For buyer-question tests and competitor comparisons, explore BREACH Pro.