Definition
robots.txt is a plain-text file at the root of a website that tells crawlers which URLs they may or may not request.
robots.txt explained
The file lives at yourdomain.com/robots.txt. It contains groups of rules, each starting with a User-agent line naming a crawler, or an asterisk for all crawlers, followed by Allow and Disallow lines for URL paths. It can also point to your XML sitemap.
Google explains that robots.txt is mainly used to avoid overloading a site with requests, and warns that it is not a mechanism for keeping a page out of Google. A disallowed URL can still be indexed without its content if other pages link to it. To keep a page out of search results, use a noindex rule or password protection instead, and don't block the page in robots.txt, or crawlers won't see the noindex.
robots.txt is also where you manage AI crawlers by name, such as GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and the Google-Extended token. A single overly broad rule can block every one of them, so review the file whenever you change your AI crawler policy.
Example
After launching a new site, your traffic drops sharply. The robots.txt copied from staging still says “Disallow: /” for all user agents. Replacing it with the production version, which allows crawling and lists the sitemap, restores access.
Why it matters
One wrong line in robots.txt can hide an entire site from search engines or AI assistants, which makes it one of the first files to check in any audit.
Related service
Technical SEO
Crawlability, indexing, structured data and rendering fixes so search engines and AI crawlers can read your site.
Technical SEO servicesSources
- Google Search Central: Introduction to robots.txt (updated 10 Dec 2025)
Published by Vidern, founded and led by Malhar Shah. Updated .
Find out what's holding your site back
Get a free SEO audit with prioritized fixes, delivered within 48 hours.
Related terms
- NoindexNoindex is a rule, set in a robots meta tag or an X-Robots-Tag HTTP header, that tells search engines not to include a page in their results.
- AI crawlerAn AI crawler is an automated bot run by an AI company that fetches web pages to train models, to build a search index, or to answer a user's question.
- CrawlingCrawling is the process by which search engine and AI bots discover web pages by following links and sitemaps, then download their content for processing.
- XML sitemapAn XML sitemap is a file that lists the URLs you want search engines to crawl, with optional details such as when each page last changed.
- Google-ExtendedGoogle-Extended is a robots.txt token controlling whether Google-crawled content can train and ground Gemini models; it doesn't affect Google Search.