1. Home
  2. Research
  3. How I checked it
MethodAI search

How I checked what AI crawlers see

Everything you need to check my numbers, or to run the test yourself. Which sites, how each one was loaded, how I counted the words, how I tested the method, what I got wrong at first, and what this can’t tell you.

Original research

The method behind the AI crawler study, collected on 5 October 2026 and checked on 6 and 7 October.

Data reviewed by Muhammad Afnan, AI Engineer at Softlixx, on 7 October 2026. What was checked

top domains
1,000
page loads per site
2
crawlers checked in robots.txt
14
countries
2

The short answer

I loaded each homepage twice in a normal browser, once with JavaScript off and once with it on. The sites were the 1,000 most popular domains on the Tranco research list. Each time I counted the words you can see, then compared the two counts. I also read every site’s robots.txt, the file that tells crawlers where they may go, by the official rules. I ran it all on 5 October 2026 from a US server and from a computer in Pakistan, then tested the method on 6 October.

Key takeaways

  • I took all of the top 1,000 domains from Tranco, a ranking built for research, using list 647LX from 4 October 2026.
  • A script loaded each homepage twice in a normal browser, with JavaScript off and then on, in a fresh session each time.
  • I wrote the rules down before the first site was loaded, and this page lists every fix I made after.
  • I ran it from two countries on 5 October 2026: the US for the headline numbers, and Pakistan as a cross-check.
  • I also tested the method, loading 439 homepages again under the name of ChatGPT’s search crawler and counting every word in the saved HTML.
  • The limits are real: it’s one day, homepages only, on a desktop, and the real AI crawlers weren’t run.

Which sites did I check?

The 1,000 most popular domains in the Tranco list, a ranking built for research.

Tranco combines five popularity rankings over 30 days: Chrome UX Report, Farsight, Majestic, Cloudflare Radar and Cisco Umbrella. It’s designed to be hard to manipulate, and every version has a fixed ID, so anyone can download the exact same list. I used the list with ID 647LX, generated on 4 October 2026 from the 30 days before, and took its top 1,000 on 5 October.

Its authors ask to be cited, and they should be: Le Pochat and colleagues, “Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation”, NDSS Symposium 2019.

Many of the 1,000 aren’t websites. They’re content networks, ad servers and app back ends, like gstatic.com or akamai.net. I kept them all in the dataset, with what happened to each, rather than quietly swapping in other sites.

How was each site loaded?

Twice, in Microsoft Edge, each time in a fresh session. First with JavaScript off, then with it on.

  • One homepage per domainhttps://domain/ first. If that couldn’t be reached (no connection, a certificate error or no answer within 30 seconds), https://www.domain/ instead. Redirects were followed, and the final address recorded.
  • A normal desktop browserMicrosoft Edge, run automatically, in a 1,366 × 900 window, language English. It used Edge’s normal user agent, without the “Headless” marker automated browsers add.
  • View A, JavaScript offScripts switched off before the page loaded. Measured when the page finished loading, plus 2 seconds.
  • View B, JavaScript onA new, separate session. Measured once the page had loaded and no new network requests had started for 2 seconds, or after 15 seconds at most.
  • No clickingNo scrolling, clicking or accepting cookie banners in either view.
  • Light on the sitesIn each run, a site got one page load with JavaScript off (two if the first address failed), one with it on, and one robots.txt request.

The two runs

Same script and rules. Both on 5 October 2026, times in UTC.

United States (headline)Pakistan (cross-check)
ComputerA GitHub Actions server (Ubuntu 24.04)An office PC (Windows 10)
Location, from Cloudflare’s trace pageUS, data centre IAD (Washington, DC area)Pakistan, data centre ISB (Islamabad)
BrowserMicrosoft Edge 154.0.4258.37Microsoft Edge 154.0.4258.53
Time12:22 to 13:2912:02 to 12:48
Sites at a time45 for the first 10, then 6

A script did the loading, not me. Opening 1,000 sites by hand would take weeks, and I’d never treat every site exactly the same way. The script does. I went through its results afterwards, the strangest ones first.

Both runs sent the same user agent, the one Edge 154 sends on Windows. The exact text is at the top of each run’s collection log. The location of each run comes from Cloudflare’s own trace page, saved during the run.

Why two countries? Sites tailor their pages to the visitor’s country, and some refuse visitors from certain places, so one location could give a lopsided picture. I decided before seeing either result that the US run gives the headline numbers, and that any headline number differing by more than 5 points would be shown both ways.

Every extra visit, counted. The checks loaded some sites again. From Pakistan, from 5 to 7 October: 555 pages once more with JavaScript off (for a code correction I later withdrew), 43 pages twice (the web components check), 36 sites twice for screenshots (4 of them on 7 October), 3 sites twice to test the crawler check, and 38 pages once more for a second check of the drafts, which found the wrong correction. From the US, on 6 October: 439 homepages twice for the crawler check. Everything else was worked out from the pages saved on 5 October, without loading anything.

How did I count the words?

From the text you can see on the page, counted with the browser’s own word splitter. Then the JavaScript-off count divided by the JavaScript-on count.

The text is the page’s visible text, as the browser reports it (document.body.innerText). The browser’s language-aware word splitter (Intl.Segmenter) counts the words, so Chinese and Japanese pages, which don’t put spaces between words, are counted fairly.

Share visible without JavaScript is the words with JavaScript off divided by the words with it on. It’s capped at 100%, because 62 homepages showed more words without JavaScript than with it. The bands use the exact share, before rounding. That only matters for one page, calendly.com in the Pakistan run: 906 of 1,007 words is 89.97%, so it’s partly visible, though the dataset shows 90.

The four bands

Fixed before any site was loaded

Share visible without JavaScriptBand
90% or moreVisible without JavaScript
50% to 89%Partly visible
10% to 49%Mostly hidden
Under 10%Empty without JavaScript

Code is taken out. When a page hides its whole body until JavaScript runs, the browser hands back all of the body’s text, code included. For those pages I count the text word by word and leave the code out. That changed two measured homepages in each run, doctolib.fr and verisign.com. The dataset shows how many words came out of each page, so anyone can rebuild the original count.

The basics were checked in both views: a page title with text, a meta description, at least one H1 with text, a canonical link, and structured data (JSON-LD).

What a site is built with comes from clues in its code with JavaScript on. Next.js (“__NEXT_DATA__” or “/_next/”), Nuxt (“__NUXT__” or “/_nuxt/”), Angular (“ng-version”), Vue (its “data-v-” marks), Gatsby (“___gatsby”), SvelteKit (“__sveltekit”), React (its own marks on the page’s elements), WordPress (“/wp-content/” or “/wp-includes/”), Shopify (“cdn.shopify.com”), Wix (“static.wixstatic.com”), Squarespace (“static1.squarespace.com”), Webflow (“data-wf-page”) and Drupal (“/sites/default/files”). A site can match more than one, and many match none. Matching clues like these is approximate: a page that embeds something from a WordPress site, for example, can pick up the WordPress clue.

Did I test the method itself?

Yes, in four ways. Each one answers a fair doubt about the main numbers, and none of them changed the overall picture.

Four checks of the method

Reported apart from the main numbers. The study shows each result.

The doubtThe checkWhat it found
“A site might send AI crawlers a different page.”On 6 October, from a US server, 439 homepages loaded again with JavaScript off, minutes apart, as the same browser and under OAI-SearchBot’s name, as OpenAI publishes it. Only sites whose robots.txt allows OAI-SearchBot.Of the 96 hidden-text homepages, 10 sent a fuller page under OAI-SearchBot’s name (at least one and a half times the words, and 100 more), 77 sent about the same page (70 with exactly the same word count), and 9 refused it. trendyol.com answered OAI-SearchBot with an error and no page, so it isn’t counted. 32 sites in all refused the OAI-SearchBot visit while serving the browser, probably because it didn’t come from OpenAI’s addresses.
“Crawlers read hidden text too.”Every word in the saved HTML of 5 October counted again, hidden or not, without loading anything.80 of 493 homepages (16.2%) still had less than half their words, against 23.5% counting visible words. It leaves out code, so text kept inside scripts isn’t counted.
“One country could be a fluke.”The whole run repeated from Pakistan.The main numbers were all within about 2 points. The biggest gap was the H1 rate: 61.3% in the US, 63.3% from Pakistan. All 511 robots.txt files read in both runs gave the same answers.
“Some text lives in web components.”The 43 US homepages that had too little text on 6 October, counted again from Pakistan, including text inside web components.9 had 50 words or more, and 7 of those showed less than half without JavaScript. Leaving them out makes the hidden share look slightly smaller, not bigger.

Why loading as OAI-SearchBot isn’t the real thing. My visits used OAI-SearchBot’s name, but they came from a GitHub server, not from OpenAI’s published addresses. A firewall that checks both will turn that visit away, which is probably what happened on the 32 sites that refused it, openai.com included. So the check shows which sites treat a crawler’s name differently. It can’t show what the real crawler gets on sites that check addresses too.

One detail of the crawler check. It counts both loads the same way, so the comparison is fair. But it counts the two hidden-body pages, doctolib.fr and verisign.com, the earlier way, before words were counted one by one (see “Words run together” below). doctolib.fr refused OAI-SearchBot, and verisign.com got the same 643 words both times, so no result changes.

Who was left out, and why?

Every domain without a real, readable homepage. Each one is counted, named in the dataset, and given a reason.

What happened to the 1,000 domains, in both runs

5 October 2026

OutcomeUnited StatesPakistan
Measured494502
No homepage316338
A copy of a site already counted6559
Refused our browser5740
Too little text to judge4447
Bot check or error page without JavaScript only2414

Source: My AI crawler study, both runs, 5 October 2026.

  • No homepageNo answer on either address, Edge’s own error page, not an HTML page even with JavaScript on, or an error status of 400 or more with no real page behind it.
  • Refused our browserWith JavaScript on, the site answered 401, 403, 429 or 503, or showed a short block page (under 200 words) with words like “Just a moment”, “Access denied”, “not a robot” or “Client Challenge”.
  • Too little textUnder 50 words even with JavaScript on, like login walls and splash pages.
  • Bot check or error page without JavaScript onlyA block page or an error status with JavaScript off, but the real page with it on. That’s what a crawler without JavaScript would get, so it’s counted on its own, not mixed into the bands.
  • A copyThe site ended up on the same website (host name) as a higher-ranked domain, like youtu.be opening youtube.com. The higher-ranked one is kept.

Two choices worth knowing. A page with under 50 words even with JavaScript on counts as too little text, even if it also showed a bot check without JavaScript. And a 503 answer counts as a refusal, even when the page says the site is briefly down.

How did I read robots.txt?

By the official rules, RFC 9309, for 14 crawlers. Only real robots.txt files that I could actually read are counted.

For each site with a homepage, I fetched robots.txt from the final website’s address and asked one question per crawler: is the homepage allowed?

RFC 9309 sets the rules. A crawler follows the group with its own name, or the “*” group if there isn’t one. The longest matching rule wins, and Allow wins a tie. A file that doesn’t exist (404 or 410) means everything is allowed.

Only clear answers count. A robots.txt request refused by a firewall, or one that failed on my side, says nothing reliable about the site’s rules. Neither does an answer that isn’t a robots.txt file at all: 21 sites in the US run sent an ordinary web page instead, like instagram.com, 2 sent an empty bot challenge, and 1 led to a script instead of a robots.txt file (unpkg.com). So only real files that were read, or clearly absent, are in the percentages. That’s 550 of 619 sites in the US run and 526 of 603 in Pakistan. The rest are named in the dataset.

The 14 crawlers

Grouped by what each company says the crawler is for. Google-Extended and Applebot-Extended are names for robots.txt rules, not separate crawlers.

GroupCrawlers
AI trainingGPTBot, ClaudeBot, Google-Extended, CCBot, Applebot-Extended, Meta-ExternalAgent, Bytespider
AI search and answersOAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot
Search engines, for comparisonGooglebot, Bingbot

Two crawlers come with a warning from their own companies. OpenAI says robots.txt rules “may not apply” to ChatGPT-User, because a person asked for the page. Perplexity says Perplexity-User “generally ignores robots.txt rules”, so I didn’t count it at all.

Other lines don’t split a group. RFC 9309 says lines like Crawl-delay or Sitemap mustn’t interfere with the rules. So when a file names a crawler, then only a Crawl-delay line, then another crawler and a block, the block covers both. That’s how I read it. It blocks CCBot on eventbrite.com and Bingbot on goodreads.com, though those sites probably meant only to slow them down.

robots.txt is a request, not a lock. It shows what a site asks crawlers to do, not what they do, and not what a firewall lets through.

What did I get wrong at first?

A few things. I found them by going through the strangest results, in a second check of the drafts, and in a full re-check on 7 October. I fixed each one for every site, not just the one that showed it.

  • A page with 170,884 “words”doctolib.fr hides its whole body until JavaScript runs, and the browser then counted everything inside it, code included. Its real text without JavaScript was 237 words. Code now comes out of every count.
  • A code correction that was wrongI also took out code that a page’s stylesheet seemed to show without JavaScript, on 40 US homepages. A second check of the drafts reloaded 38 of them and found that code was never drawn on the page, so it had never been counted. The correction had cut real words: arubanetworks.com went from 1,343 words to 34. I withdrew it, and the share of homepages with less than half their words fell from 24.2% to 23.6%, and to 23.5% after the slideshare.net fix below.
  • Words run togetherMy first fix for hidden-body pages read the text in one piece, which joined neighbouring words (“Home” and “Products” became one word). Counting them one by one took doctolib.fr from 151 to 237 words. Neither of the two pages affected changed band.
  • Security companies that looked like bot checkshcaptcha.com and checkpoint.com use words like “captcha” on their real homepages. A bot check now has to be a short page, or a refusal status.
  • Refusals counted as missing homepagesIn the Pakistan run, sites like nih.gov answered 403 because they refused the browser, not because they had no homepage. They now have their own group.
  • Web pages read as robots.txt files21 sites in the US run answered the robots.txt request with an ordinary web page, and 2 with an empty challenge. I first read those as “everything allowed”. They now count as not read. No site’s answer changed, but each blocking share rose by up to a point, because fewer files are in the total.
  • Bot-check wordings I missedamazon.com’s page without JavaScript said “not a robot”, and slideshare.net’s US page was called “Client Challenge”. My list of bot-check words had neither, so slideshare.net first counted as an empty page.
  • Data servers counted as refusalsDomains like appsflyersdk.com and btloader.com answer with data or plain text, not a web page. Because they answered 403, they counted as “refused our browser”. My rules said a page that isn’t HTML has no homepage, so that’s where they are now: 7 in each run.
  • Error pages and a noticeavito.ru answered with an unusual error code, 439, without JavaScript and its real page with it. It counted as having no homepage, and now counts as a bot check without JavaScript. pixiv.net counted as refused only because its page says it’s “protected by reCAPTCHA”. It now counts as too little text.
  • A Crawl-delay line in robots.txtMy reader ended a list of crawler names at a Crawl-delay line, which RFC 9309 says it shouldn’t. That changed one answer in each of two files.

The last three fixes, on 7 October, left the headline as it was. They moved the robots.txt shares by 0.2 points at most in the US run, and by 0.4 in the Pakistan run. I wrote the rules down before the first site was loaded, and the bands and the thresholds never changed. None of these fixes was made because of which way a result went, and two of them made the results look better for the sites, not worse.

What can’t this study tell you?

Quite a bit, and it’s better to say so up front. This is a careful snapshot, not the final word.

It’s one day. Sites change their homepages all the time. The numbers are what the pages showed on 5 October 2026.

It’s homepages only, on a desktop. Inner pages, like articles or product pages, can be built differently. Phone versions weren’t loaded.

It’s two places. A US data centre and an office in Pakistan. AI crawlers fetch from their own servers, which may be treated differently again.

It didn’t run the real AI crawlers. It copies how they read a page, based on the Vercel and MERJ finding, published in December 2024, that the major AI crawlers didn’t run JavaScript. If a crawler starts running JavaScript, the text results stop applying to it.

“none of the major AI crawlers currently render JavaScript”
Vercel and MERJ, The rise of the AI crawler

Sites can treat crawlers differently. The crawler check tested this with OAI-SearchBot’s name, but a firewall can also check where a visit comes from, and I couldn’t send visits from OpenAI’s addresses.

Visible text isn’t everything a crawler reads. A crawler reading raw HTML may also see text that’s hidden from people. The headline numbers count what a person would see. Counting hidden text too gives the lower end of the range shown in the study.

Some text isn’t drawn until you scroll. A few sites tell the browser to skip drawing sections until they come into view, to load faster. The visible-text count leaves those out, with JavaScript on and off alike. The count of all the words in the HTML includes them.

These are the biggest domains on the web, not typical small business websites. Small sites may do better or worse.

Who checked this?

Muhammad Afnan, an AI engineer at Softlixx, checked the data and the method before I published this. He’s a friend and colleague of mine.

He looked over the two datasets and how I collected them, then approved the study for publishing. You can find him on LinkedIn.

A second pair of eyes doesn’t make a study perfect. If you spot something we both missed, tell me and I’ll fix it.

How can you check my numbers?

Download the data, or open any site yourself with JavaScript off. Every number in the study can be rebuilt from the two datasets.

Each dataset has one row for every one of the 1,000 domains. It shows what happened to the site, the words with JavaScript off and on, the share and band, the code words taken out, the basics in both views, what the site is built with, whether each of the 14 crawlers is allowed and where the rule was found, all the words in the HTML, and in the US file, the crawler check. It’s a CSV file, so it opens in Excel or Google Sheets.

To check a site yourself: press Ctrl+U to see the page source and search it for a sentence you can see on the page. Or open it in Chrome or Edge, press F12, then Ctrl+Shift+P, type “Disable JavaScript”, press Enter and reload. Pages change every day, so small differences are normal.

I also saved screenshots of 25 sites picked at random, with a fixed seed so the pick can be repeated, with JavaScript off and on. Every site section 03 of the study names as an example of hidden or visible text has screenshots too. For the bot checks, each site’s saved record keeps the text of the page it got.

I’ve kept the scripts, the saved record of every site in both runs, and the raw HTML pages behind them. If you spot a mistake, tell me. I’ll fix it and note the change on the page.

Questions people ask

How can I check my own site the way an AI crawler sees it?

Press Ctrl+U to see your page’s source, then Ctrl+F for a sentence from your main text. If it’s missing, a crawler that doesn’t run JavaScript misses it too. Then open yoursite.com/robots.txt and look for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot, and for a “User-agent: *” group with “Disallow: /”. In my October 2026 test, 14 top sites blocked OAI-SearchBot that way.

Why use the Tranco list?

Because it’s built for research. It combines five popularity rankings over 30 days, it’s designed to be hard to manipulate, and every version has a fixed ID anyone can download. I used list 647LX, generated on 4 October 2026, and its top 1,000 domains, all of them, not a sample.

Why run it from two countries?

Because sites tailor their pages to where you are, and some block visitors from certain places. I ran the same script from a US server and from Pakistan on 5 October 2026. The US run gives the headline numbers, and the main numbers were all within about 2 points of each other.

Did sites send ChatGPT’s crawler a different page?

A few did. On 6 October 2026 I loaded 439 homepages as a browser and under OAI-SearchBot’s name, both with JavaScript off. Of the 96 that hid most of their words, 10 sent a fuller page under OAI-SearchBot’s name, like bestbuy.com, and 77 sent about the same page. My visits didn’t come from OpenAI’s addresses, so sites that check addresses may treat the real crawler differently.

Sources

How I checked this. I opened every source linked here and checked each figure and each quotation against the source page, most recently on 7 October 2026. Anything in a box marked “My reading” is my own interpretation, not a finding of the sources. Studies change and new ones appear, so I will update this piece and change its dates when the evidence does.

Send me your URL.

I’ll tell you what Google and AI actually see when they look at Your Brand, and what I’d fix first.

No obligation. If we’re a good fit, we’ll plan the next step together. If not, the findings are still yours to keep.