Can ChatGPT, Claude, Gemini and Perplexity read your website? I tested 1,000 top sites
On 5 October 2026 I loaded the homepages of the world’s 1,000 most popular domains twice, once with JavaScript off and once with it on, and counted the words. JavaScript off is how the crawlers behind ChatGPT, Claude and Perplexity have been found to read a page. I did it all again from a second country, and read each site’s robots.txt to see which AI crawlers it lets in.
Original data, collected on 5 October 2026 from the US and Pakistan, and checked on 6 and 7 October. Every number from my test comes from it, and you can check it yourself.
Data reviewed by Muhammad Afnan, AI Engineer at Softlixx, on 7 October 2026. What was checked
- top domains loaded
- 1,000
- homepages measured
- 494
- robots.txt files read
- 550
- countries
- 2

The short answer
Not always. It depends on how your website is built. The crawlers behind ChatGPT, Claude and Perplexity have been found to skip JavaScript, so any words your page adds with JavaScript stay hidden from them. On 5 October 2026 I loaded the 1,000 most popular domains with JavaScript off and on. Of the 494 homepages I could measure, 116 (23.5%) lost more than half their words with it off, and 61 (12%) were close to empty. Gemini is likely the exception, because it uses Google’s crawler, which runs JavaScript.
Key takeaways
- On 5 October 2026 I loaded the 1,000 most popular domains the way the crawlers behind ChatGPT, Claude and Perplexity have been found to read pages, and nearly 1 in 4 of the homepages I could measure hid most of their words.
- Gemini is likely the exception: it relies on Google’s crawler, which runs JavaScript, so it should see most of what you see.
- A page can look fine and still be almost empty to a crawler: autodesk.com looked complete, but only 51 of its 925 words were there.
- Very few sites seem to make up for it: under the name of ChatGPT’s search crawler, only 10 of 96 hidden-text homepages sent a fuller page.
- How a site is built is linked to it: 37% of React homepages hid most of their words, against 12% of sites with no framework I could detect.
- Many big sites keep AI out of training but let it into search: 18.7% block GPTBot, while only 8.7% block OAI-SearchBot, which ChatGPT search uses.
What exactly did I test?
I loaded the homepages of the world’s 1,000 most popular domains twice. Once with JavaScript off, the way AI crawlers read a page, and once with it on, the way you see it. Then I compared the words.
The sites come from the Tranco top 1,000, a ranking of the most popular domains built for research. I used the list generated on 4 October 2026 (list ID 647LX).
A browser loaded each homepage two times, each in a fresh session. First with JavaScript switched off. That’s roughly what a crawler that doesn’t run JavaScript gets: the page’s HTML and nothing more. Then with JavaScript on, the way a person sees it. Each time, I counted the words you can see on the page. The share visible is the first count divided by the second.
I also read each site’s robots.txt, the file that tells crawlers where they may go, and checked whether 14 crawlers were allowed on the homepage.
I ran the whole thing twice on 5 October 2026, from two countries. Once from a server in the US, which gives the headline numbers, and once from a computer in Pakistan as a cross-check. I decided which run would be the headline before I saw either result, and wrote it down with the rest of my rules before loading a single site.
A note on how. I didn’t open 1,000 websites by hand. Nobody could, and it wouldn’t be a fair test, because I’d look at some more carefully than others. So I used a small script that opens every homepage in a normal browser, the same way every time, and counts the words. Then I went through the results, starting with the strangest ones, had the drafts checked a second time, and fixed what I’d got wrong. The method page lists those mistakes too.
This isn’t a sample of the top 1,000. It’s all of them. But not all 1,000 domains are websites. Many are ad servers, content networks or app back ends with no homepage. Here’s what happened to each one in the US run.
What happened to the 1,000 domains
United States run, 5 October 2026
| Outcome | Domains |
|---|---|
| Measured: a real homepage with at least 50 words | 494 |
| No homepage: ad servers, content networks, errors | 316 |
| A copy of a site already counted, like youtu.be, which opens youtube.com | 65 |
| Refused our browser outright, with JavaScript on too | 57 |
| Too little text to judge, even with JavaScript on | 44 |
| A bot check or error page, but only with JavaScript off | 24 |
Source: My AI crawler study, United States run, 5 October 2026.
Want to check my work? Here are both datasets, and the page that explains every step.
Why does JavaScript matter for ChatGPT, Claude, Gemini and Perplexity?
Because the crawlers behind ChatGPT, Claude and Perplexity have been found not to run it, while Gemini relies on Google’s crawler, which does. If your words only appear after JavaScript runs, three of the four may never see them.
Many websites send a nearly empty page first, then fill it in with JavaScript inside your browser. People never notice. A crawler that doesn’t run JavaScript only gets the first version.
The evidence comes from Vercel and MERJ, who watched how AI crawlers behaved on real websites for several months and published what they found in December 2024. None of the major AI crawlers ran JavaScript. That included OpenAI’s three crawlers, and those of Anthropic (ClaudeBot), Meta, ByteDance and Perplexity (PerplexityBot). It didn’t test newer crawlers like Claude-SearchBot, Claude-User and Perplexity-User.
“ChatGPT and Claude don’t execute JavaScript, so any important content should be server-rendered.”
Gemini works differently. The same study says Gemini uses Googlebot’s setup, and Google says a headless Chromium browser renders the pages it crawls and runs their JavaScript, once its resources allow. So Gemini should get the full page, much as you see it, though I didn’t test Gemini itself. Bing says its crawler renders pages with the latest version of Microsoft Edge.
That matters for ChatGPT, because ChatGPT search doesn’t rely only on its own crawler. OpenAI says it “sometimes partners with other search providers”, and its help page links to Microsoft’s privacy statement as one of them. So a page ChatGPT’s crawler can’t read may still reach ChatGPT through a partner. Even Google, which does run JavaScript, calls sending pages ready-made “still a great idea”, partly because “not all bots can run JavaScript”.
A page that looks fine in your browser tells you nothing about what an AI crawler gets. You have to look at it with JavaScript off. The rest of this study does that for the 1,000 biggest domains on the web.
How much can a crawler without JavaScript read?
Most of it, on most sites. Half the homepages showed 90% or more of their words without JavaScript. But nearly 1 in 4 showed less than half, and 1 in 8 were close to empty.
How much of each homepage is visible without JavaScript
494 homepages, United States run, 5 October 2026
- 90% or more visible · 258 homepages
- 50% to 89% · 120 homepages
- 10% to 49% · 55 homepages
- Under 10%, close to empty · 61 homepages
Source: My AI crawler study, United States run, 5 October 2026.
The typical homepage (the median) showed 92% of its words without JavaScript. For comparison, the HTTP Archive’s 2024 Web Almanac, which looks at millions of websites, found a 17.5% difference in visible words on the typical desktop home page before and after JavaScript ran. Its method isn’t the same as mine, so it’s only a rough guide. But the top 1,000 as a whole don’t look worse than the wider web.
The trouble is at the bottom. 116 homepages (23.5%) showed less than half their words, and 61 of them showed under a tenth. Some are household names. twitch.tv showed 0 of its 376 words, just a logo and a spinning circle. airbnb.com showed 26 of 1,062, with grey boxes where the listings should be.


Looking fine isn’t the same as being readable. autodesk.com looked almost complete without JavaScript: the headline, the intro and the button were all there. But the menu and the cards below were empty, and the page had 51 of its 925 words.


microsoft.com sat in between, with 155 of its 552 words (28%). Others were fine. apple.com showed all 909 of its words, and theguardian.com showed 2,873 of 2,885.
Even the AI companies differ. anthropic.com showed 98% of its words without JavaScript, and openai.com 58%. The homepages of claude.ai and perplexity.ai showed none.
What about text that’s in the page but hidden from people? A crawler reading the raw HTML can pick up words a browser hides, like the links in a closed menu. So I also counted every word in the saved HTML, hidden or not. Even then, 80 of the 493 homepages I could count this way (16.2%) had less than half their words. So the real share for these crawlers probably sits somewhere between 16% and 23.5%. That count leaves out code, so any text a site keeps inside its scripts isn’t in it. For a crawler that reads that text too, the share could be a little lower.
Are the very biggest sites different?
I didn’t plan this comparison. I noticed it while going through the results, so treat it as a pattern to watch, not a rule. The 100 most popular domains hid more. 18 of their 51 homepages (35%) showed less than half their words without JavaScript, against 23.5% across the top 1,000. Many of them are apps you sign in to: youtube.com, instagram.com, tiktok.com and spotify.com all showed none of their words. From Pakistan it was 31%, and with only 51 homepages, one site moves the share by 2 points.
The top 100 against the top 1,000
United States run, 5 October 2026
| Top 100 | Top 1,000 | |
|---|---|---|
| Homepages measured | 51 | 494 |
| Less than half visible without JavaScript | 35.3% | 23.5% |
| robots.txt files read | 51 | 550 |
| Block GPTBot | 23.5% | 18.7% |
| Block OAI-SearchBot | 15.7% | 8.7% |
Source: My AI crawler study, United States run, 5 October 2026.
They also seemed to block AI crawlers more often: 23.5% of their robots.txt files block GPTBot, against 18.7% across the top 1,000. With only 51 files, though, that gap is two or three sites. If you compare your own site with the giants, compare it with the whole top 1,000, not just the first 100.
These are the biggest sites on the web, with big teams behind them, and nearly a quarter still hide most of their words from a crawler that doesn’t run JavaScript. The problem isn’t JavaScript itself. It’s whether the words are already in the HTML before JavaScript runs.
Do websites send AI crawlers a fuller page?
A few do. When I loaded the homepages again using the name of ChatGPT’s search crawler, 10 of the 96 that hid most of their words sent a fuller page. 77 sent about the same page, and 9 turned the visit away.
A site can treat crawlers differently from people. So on 6 October, from a US server, I loaded every measured homepage whose robots.txt lets OAI-SearchBot in twice more, both times with JavaScript off and minutes apart. Once as the same browser as before, and once using OAI-SearchBot’s name, exactly as OpenAI publishes it. That’s 439 homepages. I never used OAI-SearchBot’s name on a site whose robots.txt keeps it out.
I counted a page as fuller when it had at least one and a half times the words, and at least 100 more. 10 of the 96 hidden-text homepages sent a fuller page under OAI-SearchBot’s name. bestbuy.com sent my browser 117 words and OAI-SearchBot 2,384. yahoo.com went from 126 to 542, and target.com from 407 to 815. airbnb.com sent more too, 190 words instead of 41, but that’s still under a fifth of its 1,062.
Pages also change by themselves, so I checked that too. Between my browser’s own visits on 5 and 6 October, 5 of these 96 homepages grew by that much anyway. Two of the 10 are among those 5, so they may just be pages that change from visit to visit. The other 8, bestbuy.com, yahoo.com and target.com among them, gave my browser about the same page on both days.
77 sent OAI-SearchBot about the same page as my browser, 70 of them with exactly the same number of words. twitch.tv sent nothing to either, and autodesk.com sent both the same 51 words. One more hidden-text homepage, trendyol.com, answered the OAI-SearchBot visit with an error and no page, so it isn’t in the 96.
32 sites turned the OAI-SearchBot visit away, with a 403 or similar, while giving my browser a normal page. openai.com and perplexity.ai were among them. That’s probably a firewall catching an impostor: my visits used OAI-SearchBot’s name but didn’t come from OpenAI’s published addresses. Some sites may block that name outright, though. Either way, it shows how closely sites now check who’s visiting.
“Dynamic rendering was a workaround and not a long-term solution…”
Sending crawlers their own ready-made page is called dynamic rendering. A few big sites seem to do it, and Best Buy clearly does. Google now calls it a workaround and recommends server-side rendering instead. For most businesses, putting the words in the HTML for everyone is simpler, and it works for every crawler at once.
Do the basics survive without JavaScript?
Mostly. Nearly every homepage had its title without JavaScript, and 88% had a meta description. The main heading and structured data did worse.
The basics, with JavaScript off and on
494 homepages, United States run, 5 October 2026
| What I checked | JavaScript off | JavaScript on |
|---|---|---|
| Page title | 483 (97.8%) | 493 (99.8%) |
| Meta description | 437 (88.5%) | 447 (90.5%) |
| Main heading (H1) | 303 (61.3%) | 347 (70.2%) |
| Canonical link | 378 (76.5%) | 391 (79.1%) |
| Structured data (JSON-LD) | 250 (50.6%) | 266 (53.8%) |
Source: My AI crawler study, United States run, 5 October 2026.
48 homepages only had a main heading once JavaScript ran. Without it, a crawler sees no H1 at all. And 18 only had structured data, the code that tells search engines what a page is about, after JavaScript ran.
The title and description sit in the top part of the page’s code, and most sites send that ready-made. The body is where JavaScript does its work. So when you check your own site, look at the main heading and the text first.
Does it depend on how the site is built?
It’s linked. In both runs, homepages built with JavaScript frameworks like React and Vue hid most of their words far more often than homepages with no framework I could detect.
Homepages showing less than half their words without JavaScript
By what the site is built with, in both runs. A site can match more than one, like Next.js and React. Detection is by clues in the code, so it’s approximate.
| Built with | Sites (US run) | US run | Pakistan run |
|---|---|---|---|
| React | 187 | 37% | 37% |
| Vue | 33 | 30% | 34% |
| Next.js | 86 | 27% | 26% |
| WordPress | 62 | 15% | 7% |
| None detected | 186 | 12% | 12% |
I spotted what each site is built with from clues in its code, like the “wp-content” folder that WordPress sites have. Angular stood out most: 4 of the 5 Angular homepages showed less than half their text, in both runs. Five sites are too few to say much, so Angular isn’t in the table.
Read WordPress with care. It’s 9 hidden-text sites in the US run and 4 from Pakistan, so the two runs differ by almost 8 points. React, Next.js and sites with no framework came out nearly the same in both.
A framework doesn’t decide it on its own. 63% of the React homepages and 73% of the Next.js homepages showed half their words or more.
What counts is whether the site sends its words in the HTML (often called server-side rendering) or builds them in the browser. Both can be done with React. If your site uses a framework, ask your developer which way yours works.
Which AI crawlers do websites block?
Training crawlers more than search crawlers. 18.7% of the robots.txt files I read blocked GPTBot. Only 8.7% blocked OAI-SearchBot, the crawler OpenAI uses for ChatGPT search.
AI companies now run separate crawlers for separate jobs. OpenAI’s GPTBot collects pages that may be used to train its models. OAI-SearchBot finds pages to show in ChatGPT search. Anthropic splits its crawlers the same way. Perplexity’s main crawler, PerplexityBot, is for search, and Perplexity says it’s “not used to crawl content for AI foundation models”. Your robots.txt can let one crawler in and keep another out.
I could read a real robots.txt file, or confirm there wasn’t one, for 550 of the 619 sites with a homepage. The rest refused the request, didn’t answer, or sent a web page instead of a robots.txt file, so they’re left out.
How many sites block each crawler from their homepage
550 robots.txt files, United States run, 5 October 2026. Google-Extended and Applebot-Extended are names for robots.txt rules, not separate crawlers.
| Crawler | Company | What the company says it’s for | Blocked by |
|---|---|---|---|
| GPTBot | OpenAI | Training AI models | 18.7% |
| ClaudeBot | Anthropic | Training AI models | 18.2% |
| Google-Extended | Gemini models and apps, not Google Search | 15.6% | |
| Applebot-Extended | Apple | Training AI models | 14.4% |
| PerplexityBot | Perplexity | Perplexity search results | 14.0% |
| Claude-User | Anthropic | Fetching pages a Claude user asks about | 10.9% |
| ChatGPT-User | OpenAI | Fetching pages a ChatGPT user asks about | 10.7% |
| Claude-SearchBot | Anthropic | Claude search results | 10.2% |
| OAI-SearchBot | OpenAI | ChatGPT search results | 8.7% |
| Bingbot | Microsoft | Bing search | 2.0% |
| Googlebot | Google Search | 1.6% |
Source: My AI crawler study, United States run, 5 October 2026.
The dataset has three more crawlers, and two were blocked even more often than GPTBot: ByteDance’s Bytespider and Common Crawl’s CCBot, each by 20.7%. Meta’s Meta-ExternalAgent was blocked by 16.0%.
143 sites (26.0%) blocked at least one training crawler. 86 (15.6%) blocked at least one AI search or user crawler, and 84 of those 86 blocked a training crawler too.
56 sites blocked GPTBot but let OAI-SearchBot in, like linkedin.com, spotify.com and reuters.com. That’s the split OpenAI describes: stay out of training, stay in ChatGPT search. 47 blocked both. OpenAI says sites that block OAI-SearchBot “will not be shown in ChatGPT search answers”, though they can still appear as plain links.
14 of the 48 sites blocking OAI-SearchBot never named it. facebook.com, twitter.com and pinterest.com, for example, list the crawlers they welcome, Googlebot among them, and end with a rule that blocks everyone else. OAI-SearchBot isn’t on their lists. And 42 sites blocked all three AI search crawlers, OAI-SearchBot, Claude-SearchBot and PerplexityBot. 33 of those still let Googlebot in.
This is what “keep me out of training, but show me in ChatGPT search” looks like in a robots.txt file:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /OpenAI says a change to robots.txt can take about 24 hours to reach its search.
Blocking training and allowing search is a choice many big sites make on purpose. The blanket rule is the one to watch. If your robots.txt blocks every crawler it doesn’t name, it blocks ChatGPT search too. For Facebook that’s a policy. On a small business site it’s usually a setting nobody remembers choosing.
Do bot checks get in the way?
Sometimes. 24 sites showed their full homepage with JavaScript on, but only a bot check or an error page with it off. A crawler that doesn’t run JavaScript would get that instead.
10 of the 24 showed the very same short page without JavaScript. It said “we need to verify that you’re not a robot. This requires JavaScript.” Six of them were Amazon stores, amazon.com included. The other four were imdb.com, booking.com, ieee.org and binance.com. Others, like trustpilot.com, fandom.com and iso.org, showed a “Verifying Connection” or “Just a moment...” page, slideshare.net a “Client Challenge”, and avito.ru a security check in Russian.
Where you load from changes this. From the US server, 24 sites showed a bot check or an error page without JavaScript. From Pakistan, 14 did. The US run came from a data centre, and AI crawlers come from servers too. It also ran Edge on Linux while saying it was Windows, which some firewalls may notice. OpenAI and Perplexity both publish the IP addresses of their crawlers, so sites can let them through.
57 more sites refused our browser outright in the US run, with JavaScript on as well, against 40 from Pakistan. robots.txt can’t show you any of this.
You never see your own bot check, because your browser passes it. If you use a firewall or a service like Cloudflare, check its bot settings and make sure the crawlers you want can get through.
Do the results hold from another country?
Yes. I ran the whole test again from Pakistan, with the same script and rules. Every share in the table below came out within about 2 points of the US run.
The US run and the Pakistan run side by side
Same script, same rules, both on 5 October 2026
| Number | United States | Pakistan |
|---|---|---|
| Homepages measured | 494 | 502 |
| Less than half visible without JavaScript | 23.5% | 22.3% |
| Close to empty (under 10%) | 12.3% | 11.6% |
| Median share visible | 92.0% | 92.3% |
| Main heading (H1) without JavaScript | 61.3% | 63.3% |
| robots.txt blocks GPTBot | 18.7% | 18.8% |
| robots.txt blocks OAI-SearchBot | 8.7% | 8.9% |
| Bot check or error page without JavaScript only | 24 | 14 |
Site by site, 455 homepages were measured in both runs. 416 of them (91.4%) landed in the same band. For robots.txt, all 511 files read in both runs gave the same answer for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot.
Most of the differences were about reaching a site at all. Some sites refused one location but not the other. Before I saw any result, I set a rule: the US run gives the headline, and any headline number that differs by more than 5 points gets both versions shown. None did. Small groups moved more, because one site shifts them a lot: Vue by 4 points and the top 100 by almost 5. WordPress moved by almost 8, so section 06 shows it both ways.
10 · Your checklist
How can you check if ChatGPT can read your website?
It takes about 10 minutes, and you don’t need any tools. Do these checks on your homepage, and on the page that brings you the most customers.
Step 1: Can AI tools read your words?
- Search your page’s code for a sentenceOpen your page and press Ctrl+U (Cmd+Option+U on a Mac). A page full of code opens. That’s what a crawler downloads. Press Ctrl+F and type a sentence from your page. If it’s there, AI tools can read it. If it isn’t, ChatGPT, Claude and Perplexity probably can’t.
- Look at your page with JavaScript offIn Chrome or Edge, press F12, then Ctrl+Shift+P. Type “Disable JavaScript”, press Enter and reload the page. Scroll all the way down. What’s left is roughly what those AI tools get. Close the panel to switch JavaScript back on.
- Make sure the important words are thereWith JavaScript still off, look for your main heading, what you do, where you work and how to reach you. 48 of the top homepages in my test only had a main heading once JavaScript ran.
Step 2: Do you let AI tools in?
- Open your robots.txtGo to yoursite.com/robots.txt. Look for these names: GPTBot and OAI-SearchBot (ChatGPT), ClaudeBot and Claude-SearchBot (Claude), PerplexityBot (Perplexity) and Google-Extended (Gemini). “Disallow: /” under a name keeps that crawler out of your whole site.
- Know which name does whatGPTBot collects pages that may be used to train OpenAI’s models. OAI-SearchBot finds pages for ChatGPT search. Blocking OAI-SearchBot keeps you out of ChatGPT’s search answers.
- Watch for the catch-all rule“User-agent: *” followed by “Disallow: /” blocks every crawler the file doesn’t name, ChatGPT’s included. 14 top sites in my test blocked OAI-SearchBot this way.
- Check your firewall tooIf you use Cloudflare or a similar service, open its bot settings and make sure the crawlers you want can get through. 24 top sites showed a bot check or an error page to a browser without JavaScript. robots.txt can’t show you that.
Step 3: Found a problem? Here’s what to do
- Your words only show with JavaScriptAsk your developer one question: “Is our main text in the HTML the server sends, or added by JavaScript in the browser?” If it’s the second, ask about server-side rendering or pre-rendering. Both put your words in the HTML, for every crawler at once. Even Google calls it “still a great idea”.
- A crawler you want is blockedDelete the rule that blocks it, or give it its own group with “Allow: /”. OpenAI says a change to robots.txt can take about 24 hours to reach ChatGPT search.
- You’d rather stay out of AI trainingBlock GPTBot and ClaudeBot, but leave OAI-SearchBot, Claude-SearchBot and PerplexityBot open. That asks OpenAI and Anthropic not to train on your pages, while ChatGPT, Claude and Perplexity can still find you in search. Section 07 shows what that looks like in a robots.txt file.
A good result looks like this: your sentence is in the code, your page still makes sense with JavaScript off, your robots.txt doesn’t block OAI-SearchBot, Claude-SearchBot or PerplexityBot, and your firewall lets them through. Then nothing on your side stops AI tools from reading your page.
I ran the same test on this website before publishing it, on 6 October 2026. Every main page showed all or nearly all of its words with JavaScript off, each one had its main heading, description and structured data in the HTML, and a request for the homepage under each of the names GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot got the same full page.
Questions people ask
Does ChatGPT run JavaScript when it reads a website?
It didn’t when it was tested. Vercel and MERJ watched AI crawlers on real sites for several months, and in December 2024 published that none of the major ones, OpenAI’s three included, ran JavaScript. In my test on 5 October 2026, 23.5% of the top homepages showed less than half their words without JavaScript. ChatGPT search also uses partner search providers, which may see more.
Should I block GPTBot?
That’s your call. It depends on whether you want your pages used to train OpenAI’s models. OpenAI says blocking GPTBot means your content shouldn’t be used in training. ChatGPT search uses a different crawler, OAI-SearchBot. In my October 2026 test, 56 of the 550 big sites whose robots.txt I read blocked GPTBot but let OAI-SearchBot in.
How can I see my website the way an AI crawler does?
Press Ctrl+U to see your page’s source, then Ctrl+F for a sentence from your main text. If it’s missing, a crawler that doesn’t run JavaScript misses it too. For the full picture, turn JavaScript off in Chrome’s developer tools (F12, Ctrl+Shift+P, “Disable JavaScript”) and reload. In my October 2026 test, nearly 1 in 4 top homepages lost more than half their words that way.
Can Gemini read my website?
Probably, yes, though I didn’t test Gemini itself. Gemini relies on Google’s crawler, which renders pages and runs their JavaScript, so it should see most of what you see. robots.txt still matters: Google says the Google-Extended token controls whether your pages are used to train Gemini and to ground its answers in Gemini Apps, without affecting Google Search. In my test on 5 October 2026, 15.6% of the robots.txt files I read blocked it.
Are WordPress sites safe?
Mostly, but check yours. In my test on 5 October 2026, WordPress homepages in the top 1,000 hid most of their words less often than React ones: 15% in the US run and 7% from Pakistan, against 37% for React in both. The WordPress groups were small and themes differ, so check your own site with Ctrl+U or with JavaScript off.
Sources
- AI crawler view dataset, United States run (CSV, 1,000 rows)Tayyab Hussain · Collected 5 October 2026
- AI crawler view dataset, Pakistan run (CSV, 1,000 rows)Tayyab Hussain · Collected 5 October 2026
- How I checked what AI crawlers see (the full method)Tayyab Hussain · 6 October 2026
- Information on the Tranco list with ID 647LXTranco · Generated 4 October 2026
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against ManipulationLe Pochat et al., NDSS Symposium · 2019
- The rise of the AI crawlerVercel and MERJ · Published 17 December 2024
- SEO chapter, Web Almanac 2024HTTP Archive · Published 2 December 2024
- Overview of OpenAI CrawlersOpenAI · Checked 6 October 2026
- Searching the web with ChatGPTOpenAI Help Center · Checked 6 October 2026
- Does Anthropic crawl data from the web, and how can site owners block the crawler?Claude Help Center · Checked 6 October 2026
- Perplexity CrawlersPerplexity · Checked 7 October 2026
- Google’s common crawlersGoogle for Developers · Checked 7 October 2026
- Understand JavaScript SEO BasicsGoogle Search Central · Checked 6 October 2026
- Dynamic Rendering as a workaroundGoogle Search Central · Checked 6 October 2026
- Which Crawlers Does Bing Use?Bing Webmaster Tools · Checked 6 October 2026
- About ApplebotApple Support · Checked 6 October 2026
- RFC 9309: Robots Exclusion ProtocolIETF · September 2022
How I checked this. I opened every source linked here and checked each figure and each quotation against the source page, most recently on 7 October 2026. Anything in a box marked “My reading” is my own interpretation, not a finding of the sources. Studies change and new ones appear, so I will update this piece and change its dates when the evidence does.