Definition
An AI crawler is an automated bot run by an AI company that fetches web pages to train models, to build a search index, or to answer a user's question.
AI crawler explained
AI companies run several kinds of bots, and it pays to tell them apart. Training crawlers, such as OpenAI's GPTBot and Anthropic's ClaudeBot, collect content that may be used to train future models. Search crawlers, such as OAI-SearchBot, Claude-SearchBot and PerplexityBot, build the indexes that AI search features draw on. User-triggered fetchers, such as ChatGPT-User and Perplexity-User, visit a page because a person asked the assistant about it.
Each identifies itself with a user-agent name that you can allow or block in robots.txt, although OpenAI and Perplexity note that their user-triggered fetchers may not follow robots.txt rules. Many sites block AI crawlers without meaning to, through firewall rules, bot protection at the CDN, or an old line in robots.txt.
A common approach is to allow search and user-triggered bots, so your pages can be cited, and to decide separately about training bots based on your content and licensing position.
Example
Your marketing team wonders why Perplexity never cites your guides. A check shows your CDN's bot protection serves a challenge page to PerplexityBot. After allowing it, your guides become eligible to appear as sources again.
Why it matters
If an assistant's search crawler can't reach your pages, it can't cite them. Checking crawler access is usually the first step of a GEO audit.
Related service
Generative Engine Optimization (GEO)
We help your brand get found, understood and cited by ChatGPT, Perplexity, Gemini and Google AI Overviews.
Generative engine optimization servicesSources
- OpenAI: Overview of OpenAI crawlers (checked 30 Sep 2026)
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler? (checked 30 Sep 2026)
- Perplexity: Perplexity crawlers (checked 30 Sep 2026)
Published by Vidern, founded and led by Malhar Shah. Updated .
See how AI assistants describe your brand
The free GEO audit scores your readiness for ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews from 0 to 100, with an action plan, within 24 hours.
Related terms
- GPTBotGPTBot is OpenAI's web crawler for collecting content that may be used to train its generative AI models, and it follows robots.txt.
- OAI-SearchBotOAI-SearchBot is OpenAI's crawler for surfacing websites in ChatGPT search, and blocking it keeps your pages out of ChatGPT search answers.
- PerplexityBotPerplexityBot is Perplexity's crawler for surfacing and linking websites in its answers, and Perplexity says it isn't used to train AI models.
- ClaudeBotClaudeBot is Anthropic's crawler for web content that may be used to train its AI models, separate from its Claude-SearchBot and Claude-User agents.
- robots.txtrobots.txt is a plain-text file at the root of a website that tells crawlers which URLs they may or may not request.