Glossary · Technical SEO

Crawling

Also called: web crawling, crawler, spider, Googlebot

Definition

Crawling is the process by which search engine and AI bots discover web pages by following links and sitemaps, then download their content for processing.

Crawling explained

Crawlers such as Googlebot, Bingbot and OAI-SearchBot start from known URLs, fetch each page, and add any new links they find to a queue. Sitemaps give them extra URLs to try. What they fetch is then processed and, if it qualifies, stored in an index.

Several things can stop a page from being crawled: robots.txt rules that disallow it, no internal links or sitemap entries pointing to it, server errors or timeouts, login walls, and firewall or bot-protection rules that block particular crawlers. Very large sites can also run into crawl budget limits, where search engines don't get to every URL often.

You can see crawling in action in Google Search Console's crawl stats report and in your server logs, which record every request a bot makes. Checking both is the reliable way to know whether important pages are being fetched, how often, and with what response.

Example

Your blog's older posts are only reachable through a “load more” button that uses JavaScript. Crawlers don't click it, so those posts are rarely crawled. Adding plain paginated links lets crawlers reach every post.

Why it matters

A page that isn't crawled can't be indexed, ranked or cited, however good it is.

Related service

Technical SEO

Crawlability, indexing, structured data and rendering fixes so search engines and AI crawlers can read your site.

Technical SEO services

Published by Vidern, founded and led by Malhar Shah. Updated .

Find out what's holding your site back

Get a free SEO audit with prioritized fixes, delivered within 48 hours.