Definition
An XML sitemap is a file that lists the URLs you want search engines to crawl, with optional details such as when each page last changed.
XML sitemap explained
Google describes a sitemap as a file where you provide information about the pages, videos and other files on your site and the relationships between them. Search engines use it to crawl more efficiently. It doesn't guarantee indexing, but it helps them find pages, especially new pages and pages with few internal links.
A clean sitemap:
- Lists only canonical, indexable URLs that return a 200 status.
- Leaves out redirected, noindexed, duplicate and error URLs.
- Uses accurate last-modified dates, updated only when content really changes.
- Is referenced in robots.txt and submitted in Google Search Console.
- Is split into several files, linked from a sitemap index, on large sites.
Comparing the URLs in your sitemap with what Search Console reports as indexed is a quick health check: large gaps usually point to quality, duplication or crawling problems.
Example
Your sitemap includes thousands of redirected and out-of-stock URLs from an old platform. Regenerating it to list only live, canonical pages, with honest last-modified dates, helps Google focus on the pages you actually want indexed.
Why it matters
An accurate sitemap helps search engines discover and recrawl your important pages quickly, and a messy one sends mixed signals about what matters.
Related service
Technical SEO
Crawlability, indexing, structured data and rendering fixes so search engines and AI crawlers can read your site.
Technical SEO servicesSources
- Google Search Central: What is a sitemap? (updated 10 Dec 2025)
Published by Vidern, founded and led by Malhar Shah. Updated .
Find out what's holding your site back
Get a free SEO audit with prioritized fixes, delivered within 48 hours.
Related terms
- CrawlingCrawling is the process by which search engine and AI bots discover web pages by following links and sitemaps, then download their content for processing.
- IndexingIndexing is when a search engine processes a crawled page and stores it in its index, making it eligible to appear in search results.
- robots.txtrobots.txt is a plain-text file at the root of a website that tells crawlers which URLs they may or may not request.
- Canonical tagA canonical tag is an HTML link element that tells search engines which URL is the preferred version of a page when similar or duplicate versions exist.
- llms.txtllms.txt is a proposed Markdown file at the root of a website that gives AI tools a short summary of the site and links to its most useful pages.