Definition
Crawling is the process by which search engine and AI bots discover web pages by following links and sitemaps, then download their content for processing.
Crawling explained
Crawlers such as Googlebot, Bingbot and OAI-SearchBot start from known URLs, fetch each page, and add any new links they find to a queue. Sitemaps give them extra URLs to try. What they fetch is then processed and, if it qualifies, stored in an index.
Several things can stop a page from being crawled: robots.txt rules that disallow it, no internal links or sitemap entries pointing to it, server errors or timeouts, login walls, and firewall or bot-protection rules that block particular crawlers. Very large sites can also run into crawl budget limits, where search engines don't get to every URL often.
You can see crawling in action in Google Search Console's crawl stats report and in your server logs, which record every request a bot makes. Checking both is the reliable way to know whether important pages are being fetched, how often, and with what response.
Example
Your blog's older posts are only reachable through a “load more” button that uses JavaScript. Crawlers don't click it, so those posts are rarely crawled. Adding plain paginated links lets crawlers reach every post.
Why it matters
A page that isn't crawled can't be indexed, ranked or cited, however good it is.
Related service
Technical SEO
Crawlability, indexing, structured data and rendering fixes so search engines and AI crawlers can read your site.
Technical SEO servicesPublished by Vidern, founded and led by Malhar Shah. Updated .
Find out what's holding your site back
Get a free SEO audit with prioritized fixes, delivered within 48 hours.
Related terms
- IndexingIndexing is when a search engine processes a crawled page and stores it in its index, making it eligible to appear in search results.
- Crawl budgetCrawl budget is the number of URLs a search engine can and wants to crawl on a site in a given time, set by crawl capacity and crawl demand.
- robots.txtrobots.txt is a plain-text file at the root of a website that tells crawlers which URLs they may or may not request.
- XML sitemapAn XML sitemap is a file that lists the URLs you want search engines to crawl, with optional details such as when each page last changed.
- AI crawlerAn AI crawler is an automated bot run by an AI company that fetches web pages to train models, to build a search index, or to answer a user's question.