How AI assistants find your website
Published:
Short answer: AI assistants have no secret route to your site. They read public pages through crawlers and search indexes. To be mentioned you need (1) to let them read you, (2) to be indexed and (3) to have content that answers questions clearly. There is no guarantee, and anyone promising one is not telling the truth.
Three kinds of bot, not one
The large companies use separate crawlers for separate purposes. This matters in practice: if you block one, you do not automatically block the others.
| Company | Model training | Search | User request |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Google-Extended (usage control) | Googlebot (normal indexing) | not applicable |
Perplexity uses PerplexityBot (search) and Perplexity-User (user request).
- According to OpenAI, if you block OAI-SearchBot you will not appear in ChatGPT search results. Blocking GPTBot only concerns model training.
- OpenAI states that robots.txt rules may not apply to ChatGPT-User, because the action is started by a user.
- Anthropic runs three independent crawlers, each with its own name, and states that they respect robots.txt.
- According to Google, Google-Extended does not affect inclusion or ranking in Google Search. It is a control for the use of content in Gemini products.
What Google says about its AI features
For a page to appear as a supporting link in AI Overviews or AI Mode, it must be indexed and eligible to be shown in Google Search with a snippet. Google writes that there are no additional requirements, no special optimizations and no special schema.org markup. The fundamental SEO practices remain the foundation.
What about llms.txt?
llms.txt is a file proposed to guide language models to your important content. It is not an established standard. Google states that it is not needed to appear in its AI features, and in May 2026 its updated guide named it explicitly as unnecessary. For the other providers its use is not documented.
We have one on our own site too, because it costs almost nothing, but we do not expect results from it. If a consultant sells you llms.txt as a solution, ask them to show you evidence.
What we did on our own site
- Static HTML pages: the content is in the HTML itself and needs no JavaScript to be read.
- One H1 per page, and the answer to the question at the start of the text, before the details.
- A visible FAQ section, and schema.org data that matches what is shown on the page.
- Sitemap, canonical and hreflang for Greek and English.
- A robots.txt that allows the search and user crawlers of the main providers. The choice for training crawlers is our own decision and may differ for your business.
- Lightweight pages with few requests, so crawlers read them quickly.
We cannot promise that AI assistants will mention us, and nobody can. These are the measures that depend on us.
If you use Cloudflare or another firewall
robots.txt is not the only place where a crawler gets blocked. Services such as Cloudflare have AI crawler controls (AI Crawl Control) and firewall rules that can block them even if robots.txt allows them. Check both.
How to check whether they read you
- Read your robots.txt and make sure there is no blanket block, or block of the search crawlers.
- Look at your server or Cloudflare logs for the crawler names (OAI-SearchBot, Claude-SearchBot, PerplexityBot).
- Make sure your important pages are indexed in Google and Bing, using Search Console and Bing Webmaster Tools.
- Ask the assistants the questions a customer of yours would ask and keep notes each month. Answers change, so the trend matters, not a single test.
What matters most
- Content with specific, verifiable facts and clear answers.
- A consistent name and details everywhere (site, LinkedIn, Google Business Profile), so you can be told apart from other entities with the same name.
- Mentions from reliable sources, such as partners, organisations and publications.
If you want to see how we work on SEO and AI visibility, see our service.
Frequently asked questions
Should I block GPTBot?
It is your choice. GPTBot concerns model training. If you want to appear in ChatGPT search, you should not block OAI-SearchBot.
Do I need llms.txt?
Not for Google, which states it is not needed. For the other providers its use is not documented. You can have one because it costs little, but do not expect results.
Do I need special schema for AI?
No. Google says there is no special schema for its AI features. schema.org remains useful for page understanding and rich results, as long as it matches the visible content.
How quickly will they mention me?
There is no guaranteed time and no guarantee that you will be mentioned. It depends on indexing, content and competition. Measure the trend month by month.