TL;DR: In September 2026 we tested 1,319 well-known websites (Majestic Million ranks 1,001–3,000) on five basic checks that decide whether AI assistants like ChatGPT, Perplexity and Claude can reach, read and identify a site. 76.3% failed at least one. Only 23.7% passed all five. The most expensive failure is easy to miss: 15.4% of sites turn AI crawlers away at the firewall while their robots.txt says they are welcome.
This study is written for the people who own a website's results: marketing leads, founders and SEO managers. Every finding comes with the way to check your own site, and the whole method is below so you can repeat it.
What is AI search readiness?
AI search readiness is whether AI assistants can do three things with your website: reach it (your robots.txt and firewall let their crawlers in), read it (the content exists in the HTML, not only after JavaScript runs), and identify it (machine-readable signals say who you are and what the pages are).
It is not a bag of new tricks. Google says its AI features rest on the same crawling and ranking systems as Search, and that no special files or markup are needed to appear in them (Google Search Central, 2026). So readiness comes down to technical basics. Our data shows that a large share of well-known websites still get those basics wrong.
How we ran the study
- Sample: the 2,000 domains ranked 1,001–3,000 in the Majestic Million, the free list of the web's most-linked domains. This is the tier below the global giants: established brands, publishers, software companies, universities and public bodies.
- Exclusions: 183 domains were unreachable, 318 returned an error or bot challenge even to a normal browser request, 179 redirected to a different website, and 1 was a duplicate. That leaves 1,319 measurable sites.
- Five checks: each homepage was tested for (1) AI answer bots allowed in robots.txt, (2) the firewall letting AI crawlers in, (3) content visible without JavaScript, (4) Organization schema in JSON-LD and (5) an XML sitemap plus a meta description.
- Controlled firewall test: for every site we sent the same homepage request from the same machine within the same minute, first as a normal Chrome browser and then as GPTBot, ClaudeBot and PerplexityBot. Only the user-agent changed. A site counted as refusing a crawler only when it returned an explicit refusal (403, 402, 429 or a bot challenge page) to the crawler after serving the page to the browser. Timeouts were treated as inconclusive, not as blocks.
- Honest caveats: some firewalls verify real crawlers by IP address, so the genuine GPTBot may get in where our user-agent test did not. What we measured is how servers treat unverified AI traffic, which is also how they treat the growing number of AI agents that are on no allow-list. We tested homepages only. Site type (commercial, .org, government and education) was assigned by domain ending.
The scorecard: where websites fail
| Check | All sites (1,319) | Commercial (977) | .org (190) | Gov & edu (152) |
|---|---|---|---|---|
| robots.txt blocks AI answer bots | 13.6% | 16.8% | 6.3% | 2.0% |
| Firewall refuses AI crawlers | 23.1% | 23.5% | 26.8% | 15.8% |
| Content invisible without JavaScript | 9.7% | 10.5% | 10.0% | 3.9% |
| No Organization schema | 56.4% | 49.6% | 74.7% | 77.0% |
| Missing sitemap or meta description | 32.8% | 29.0% | 46.8% | 39.5% |
| Passed all five | 23.7% | 26.4% | 17.4% | 13.8% |
76.3% of 1,319 well-known websites fail at least one basic AI-search check, and only 23.7% pass all five.
Source: Vidern AI Search Readiness Study, September 2026
Finding 1: One in four sites turns AI crawlers away at the door
23.1% of sites refused at least one AI crawler that their own homepage had just served to a browser, and 11.8% refused all three we tested. By crawler, 18.8% refused GPTBot, 20.2% refused ClaudeBot and 13.7% refused PerplexityBot.
The bigger problem is that most of this blocking is invisible from the outside. 15.4% of sites allow every AI crawler in robots.txt but refuse at least one at the firewall. The published policy says "welcome", while the server says 403. That matches the 17.6% we found among the top 1,000 sites in our July 2026 AI crawler study, so the pattern holds well beyond the very largest sites.
Examples of sites whose robots.txt allows AI crawlers but whose server refused all three: Goldman Sachs, Morningstar, LastPass, Unity and Padlet. Six Hearst magazine sites (Esquire, Elle, Cosmopolitan, Popular Mechanics, Good Housekeeping and Harper's Bazaar) refused GPTBot and ClaudeBot but let PerplexityBot through, and none of their robots.txt files mention it. Even two well-known digital-marketing sites, WordStream and NeilPatel.com, served a bot challenge to all three crawlers.
Some of these are deliberate choices enforced by IP verification. In our client audits, though, the usual cause is simpler: bot protection was switched on for security reasons, and nobody checked what it does to AI crawlers. Cloudflare-served sites in our sample refused AI crawlers far more often (33.8%) than sites on other infrastructure (19.4%). That is a sign that bot-protection defaults, not written policy, decide a lot of outcomes.
Check your site: run the free AI Crawler Access Checker, which runs the same robots.txt and live-response tests on your domain. If you use a CDN or WAF, open its bot settings and look for AI-crawler rules before you assume you are visible.
Finding 2: Many sites block the wrong AI bot
AI companies run several crawlers with different jobs. Blocking one does not block the others. Here is what each major crawler does and how often our sample blocked it in robots.txt:
| Crawler | Company | What it does | Blocked in robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Collects content for model training | 16.6% |
| OAI-SearchBot | OpenAI | Builds the index for ChatGPT search results | 7.1% |
| ChatGPT-User | OpenAI | Fetches pages when a ChatGPT user asks for them | 10.1% |
| ClaudeBot | Anthropic | Collects content for model training | 15.3% |
| PerplexityBot | Perplexity | Builds Perplexity's search index | 11.8% |
| Google-Extended | Controls use for Gemini models; no effect on Search | 13.6% | |
| CCBot | Common Crawl | Builds an open dataset many AI models train on | 17.7% |
| Bingbot | Microsoft | Bing search, which also feeds Copilot | 1.3% |
OpenAI's own documentation is clear on the difference. GPTBot is used to make its foundation models more useful and safe, while OAI-SearchBot is used to surface websites in ChatGPT search, and "sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers" (OpenAI, Overview of OpenAI Crawlers). Anthropic splits its crawlers the same way. Blocking ClaudeBot opts you out of training, but Claude-SearchBot and Claude-User need their own rules (Anthropic Help Center). Google states that Google-Extended does not affect inclusion or ranking in Search (Google, common crawlers).
Of the sites that block GPTBot, 58.0% still allow OAI-SearchBot. That is a precise choice: no training, but stay visible in ChatGPT search. The other 42.0%, 7.0% of all sites, block both, and by OpenAI's own rules they will not appear in ChatGPT search answers. For a news publisher negotiating content licences, that can be a reasonable choice. For a business that wants buyers to find it, it is usually a mistake copied from someone else's robots.txt file.
58% of websites that block OpenAI's training crawler still allow its search crawler, choosing "no training, but stay findable in ChatGPT". The other 42% have opted out of ChatGPT search entirely.
Source: Vidern AI Search Readiness Study, September 2026
Finding 3: Publishers have started charging AI crawlers
36 sites (2.7%) answered at least one AI crawler with HTTP 402 Payment Required. That is the status code behind Cloudflare's pay per crawl, which lets site owners charge AI crawlers per request and return a 402 with a price to crawlers that have not agreed to pay (Cloudflare).
Almost all of them are publishers: Allrecipes, Deadline, Billboard, the San Francisco Chronicle, NZZ, Food & Wine, Serious Eats and Martha Stewart among them. Stack Exchange sites Super User and Ask Ubuntu returned 402 to ClaudeBot, which fits Stack Overflow's public move to pay per crawl (Stack Overflow).
For publishers, content is the product, so charging for it makes sense. For a brand that wants to be recommended, the AI assistant is effectively a salesperson, and putting your product pages behind a paywall only hides them from it.
Finding 4: About one in ten homepages is blank without JavaScript
9.7% of homepages delivered fewer than 100 words of text in their HTML, and 5.4% delivered none at all. Their content exists only after JavaScript runs in a browser. Delta's homepage arrives as an empty <idp-root> element. Kaggle's arrives as an empty <div id="root">. Duolingo, Qualcomm, Ryanair and Roblox also served zero words of text before scripts ran.
This matters because the major AI crawlers do not run JavaScript. Vercel's analysis of crawler traffic found that "none of the major AI crawlers currently render JavaScript", including OpenAI's three crawlers, ClaudeBot and PerplexityBot (Vercel). A page that needs JavaScript to show its text looks empty to them. The median homepage in our sample carried 930 words in its raw HTML, so this failure is avoidable.
Check your site: in Chrome, open DevTools, press Ctrl+Shift+P (Cmd+Shift+P on Mac), type "Disable JavaScript" and reload the page. What you still see is roughly what an AI crawler sees. If key pages go blank, you need server-side rendering or pre-rendering, which is a job for your developers or a technical SEO partner.
Finding 5: Most sites never tell machines who they are
56.4% of homepages had no Organization schema, and 49.3% had no JSON-LD structured data of any kind. Government and education sites were weakest, with 77.0% missing Organization schema.
To be clear, Google says there is no special schema for its AI features (Google Search Central). Organization markup is not an AI trick. It is standard structured data that states your official name, logo, URL, contact points and social profiles in a form machines can read without guessing (Google, Organization structured data). When AI systems describe a company, that clear identity matters. Adding it to your homepage takes about 30 minutes.
The remaining basics were patchy too: 19.0% of homepages had no meta description, 22.6% had no XML sitemap we could find, and 31.2% had no H1 heading.
What about llms.txt?
13.6% of sites publish an llms.txt file, the proposed markdown summary that tells AI tools where a site's key content lives. We report it separately and did not score it, because Google states that Search ignores llms.txt, and that creating one neither helps nor harms visibility in Google Search, including its AI features (Google Search Central).
One result stands out: 16.2% of the sites that publish llms.txt also refuse AI crawlers at the firewall. They wrote a guide for AI crawlers and then blocked the crawlers from reading it. If you are weighing where to spend an hour, spend it on the five checks above first.
How to check your own site in 15 minutes
- robots.txt (2 min): open
yoursite.com/robots.txt. Look forUser-agent: OAI-SearchBot,ChatGPT-User,PerplexityBotorClaude-SearchBotfollowed byDisallow: /, and for aUser-agent: *group that disallows everything. Decide separately whether you want to block training bots (GPTBot, ClaudeBot, CCBot, Google-Extended). - Firewall (3 min): run your domain through the AI Crawler Access Checker. Any 403, 402, 429 or challenge result means your firewall is making the decision, whatever robots.txt says. Fix it in your CDN or WAF bot settings.
- JavaScript (3 min): disable JavaScript in Chrome DevTools and reload your homepage and one key product or service page. If the main text disappears, raise server-side rendering with your developers.
- Organization schema (4 min): paste your homepage URL into the Schema Markup Validator. You should see an Organization (or more specific type) with your name, logo, URL and sameAs links to your official profiles.
- Sitemap and meta (3 min): open
yoursite.com/sitemap.xmlor check the Sitemap line in robots.txt, then view your homepage source and search forname="description".
What to fix first
| Priority | Fix | Typical effort | Why first |
|---|---|---|---|
| 1 | Let AI crawlers through your firewall/CDN | Under 1 hour | Nothing else matters while requests get a 403 |
| 2 | Allow answer bots in robots.txt | 15 minutes | Opt out of training if you want, but keep search bots |
| 3 | Add Organization schema | 30 minutes | Cheapest way to state who you are, clearly |
| 4 | Sitemap + meta descriptions | 1–2 hours | Basic discovery signals every crawler uses |
| 5 | Server-render key content | Days to weeks | Highest effort, but a blank page cannot be cited |
These five checks are the entry ticket, not the whole game. Once AI crawlers can reach and read your site, whether they cite you depends on content quality, brand mentions across the web and how directly your pages answer buyers' questions. Our free GEO audit scores those as well, across six categories.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT?
Not from ChatGPT search. OpenAI uses GPTBot for model training and OAI-SearchBot for ChatGPT search results. Only blocking OAI-SearchBot keeps you out of ChatGPT search answers, according to OpenAI's crawler documentation. In our study, 58% of sites that block GPTBot still allow OAI-SearchBot.
How do I know if my firewall blocks AI crawlers?
Your robots.txt cannot tell you, because the firewall decides before robots.txt is ever read. Test with a tool that requests your pages using real AI-crawler user-agents, such as our free AI Crawler Access Checker, and review the bot-management settings in your CDN or WAF. In our study, 15.4% of sites allowed AI crawlers in robots.txt but refused them at the firewall.
Can AI crawlers read JavaScript-rendered websites?
Generally no. Vercel's crawler analysis found that none of the major AI crawlers, including OpenAI's crawlers, ClaudeBot and PerplexityBot, render JavaScript. If your content only appears after scripts run, AI crawlers see an empty page. 9.7% of homepages in our study had fewer than 100 words before JavaScript.
Do I need an llms.txt file?
It is optional. Google says Search ignores llms.txt and that it neither helps nor harms visibility in Google's AI features. It is cheap to add, but it should come after crawler access, rendering and structured data, which decide whether AI systems can use your site at all.
Should I block AI crawlers to protect my content?
It depends on how you make money. Publishers who sell content may reasonably block training bots or charge for access. 2.7% of sites in our study already return HTTP 402 Payment Required to some AI crawlers. Most businesses want to be recommended, so a common middle ground is to block training crawlers and allow search and user-request crawlers.
Cite this study
Figures come from Vidern's scan on 28 September 2026 of Majestic Million ranks 1,001–3,000 (1,319 measurable websites), using the method described above. Journalists, researchers and bloggers are welcome to cite this study with attribution to Vidern and a link to this page. The per-site results are available to download: AI Search Readiness Study dataset (CSV). For category breakdowns or questions about the method, contact us.
Related reading
- We Tested the World's Top 1,000 Websites: 41% Are Unreadable to ChatGPT
- Is Your Website Blocked to ChatGPT? How to Check & Fix It
- AI Search Visibility: How to Check If ChatGPT, Gemini & Perplexity Recommend Your Brand
- What Is Generative Engine Optimization (GEO)? The 2026 Guide
Want to know where your site stands? Our free GEO audit runs these checks and more on your domain and sends you a 0–100 score with a prioritised action plan within 24 hours. No credit card, no sales call required.
Free SEO Audit · No Commitment
Is Your Site Invisible to AI Search?
Get a free audit of your technical health, AI visibility, and growth roadmap — delivered within 48 hours.
Trusted by 150+ companies globally
Related services
Vidern helps brands grow through SEO audits, technical SEO, content marketing, and SaaS SEO. Check your AI visibility with our free AI Visibility Checker and AI Crawler Access Checker, or get a free GEO audit.
The Vidern SEO Team consists of certified SEO specialists with 10+ years of combined experience helping SaaS companies, e-commerce brands, and local businesses across India grow their organic search presence.