Two Webs

Top 100 / who.int

92

Mostly one web

11 of 14 AI agents get the full page at https://who.int/. Checked 2026-09-25 09:19 UTC.

who.int scores 92 out of 100 for machine visibility, number 37 of the 100 most visited websites checked on 2026-09-25. 11 of 14 AI agents receive the full page: 5 of 8 training crawlers, 3 of 3 AI search indexes and 3 of 3 assistant fetchers. The one refusal is written in robots.txt, which is the transparent way to do it. The pattern is "cite me, don't train on me": 1 of 8 training crawlers are refused while 6 of 6 search and assistant fetchers get through. There is no llms.txt.

youfull pagedegradedblockedunknown

Human

Status
200 → https://www.who.int/
Title
World Health Organization (WHO)
H1
WHO expands safe options for contraception
Words
683
Scripts
40 tags · 10 KB
Meta robots
none
Canonical
https://www.who.int
Server
cloudflare

Machine

robots.txt
527 groups · 1 sitemaps
llms.txt
missing
llms-full.txt
missing
X-Robots-Tag
none
Blocked
CCBot
Degraded
none
Challenged
none

Gaps

  • medium
    1 of 14 AI agents are blockedCCBot: robots.txt Disallow: / (group: CCBot)
  • info
    Training blocked, citation allowedTraining crawlers are refused while search and assistant fetches get through. This is the common "cite me, don't train on me" posture.
  • low
    No llms.txtOptional, but a short curated index at /llms.txt is the cheapest way to tell assistants what matters on this site.
  • info
    Canonical points elsewhereCanonical is https://www.who.int/. Indexes will credit that URL, not this one.

Every agent

AgentTyperobots.txtFetchWordsResult
Human (Chrome)YouHumanno file2001582 ms683visible
HTTP 200, 683 words
GooglebotGoogle · Classic search index. Also feeds AI Overviews.Search engineno file2001159 ms683visible
HTTP 200, 683 words
Google-ExtendedGoogle · Robots-only token controlling Gemini training use. Never fetches on its own.AI trainingno filen/aunknown
Robots-only token; allowed by robots.txt
GPTBotOpenAI · Training crawler for OpenAI models.AI trainingno file2001595 ms683visible
HTTP 200, 683 words
OAI-SearchBotOpenAI · Indexes for ChatGPT search results and citations.AI search indexno file2001897 ms683visible
HTTP 200, 683 words
ChatGPT-UserOpenAI · Live fetch when a user asks ChatGPT about a page.AI assistant (live fetch)no file2001476 ms683visible
HTTP 200, 683 words
ClaudeBotAnthropic · Training crawler for Claude models.AI trainingno file2001711 ms683visible
HTTP 200, 683 words
Claude-SearchBotAnthropic · Indexes for Claude web search.AI search indexno file2001344 ms683visible
HTTP 200, 683 words
Claude-UserAnthropic · Live fetch when a user asks Claude about a page.AI assistant (live fetch)no file2001193 ms683visible
HTTP 200, 683 words
PerplexityBotPerplexityAI search indexno file2001565 ms683visible
HTTP 200, 683 words
Perplexity-UserPerplexityAI assistant (live fetch)no file2001884 ms683visible
HTTP 200, 683 words
BingbotMicrosoft · Bing index. Also feeds Copilot and ChatGPT search.Search engineno file2001190 ms683visible
HTTP 200, 683 words
Applebot-ExtendedApple · Robots-only token for Apple Intelligence training.AI trainingno filen/aunknown
Robots-only token; allowed by robots.txt
Meta-ExternalAgentMetaAI trainingno file2001695 ms683visible
HTTP 200, 683 words
AmazonbotAmazon · Alexa and Amazon AI.AI trainingno file2001550 ms683visible
HTTP 200, 683 words
BytespiderByteDance · Known for ignoring robots.txt.AI trainingno file2001704 ms683visible
HTTP 200, 683 words
CCBotCommon Crawl · Open dataset used to train most LLMs.AI trainingdisallowedDisallow: /2001577 ms683blocked
robots.txt Disallow: / (group: CCBot)

Snapshot

Measured 2026-09-25 09:19 UTC from a residential connection. Sites change their rules often, and sites that verify crawler IPs answer differently to different networks. Run a live comparison to see what who.int does right now, from Cloudflare's network.

← triplinkintl.com (92)All 100 top 100example.com (90) →