Two Webs

Publishers / npr.org

20

Invisible to machines

0 of 14 AI agents get the full page at https://npr.org/. Checked 2026-09-25 09:33 UTC.

npr.org scores 20 out of 100 for machine visibility, number 80 of the 100 content publishers checked on 2026-09-25. Even a plain request carrying a Chrome user agent was refused (HTTP 402), so npr.org serves its pages only to clients that pass a full browser check. Every crawler below got the same treatment, which is why the score is so low. All 11 refusals are written in robots.txt, which is the transparent way to do it. Googlebot itself is unknown: Fetch failed: Timed out. There is no llms.txt.

youfull pagedegradedblockedunknown

Human

Status
402 → https://www.npr.org/
Title
none
H1
none
Words
0
Scripts
0 tags · 0 KB
Meta robots
none
Canonical
none
Server
hidden

Machine

robots.txt
47 groups · 4 sitemaps
llms.txt
missing
llms-full.txt
missing
X-Robots-Tag
noindex
Blocked
Google-Extended, GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Applebot-Extended, Bytespider, CCBot
Degraded
none
Challenged
none

Gaps

  • high
    Humans cannot load this page eitherHTTP 402. Every other check is measured against a broken baseline.
  • high
    11 of 14 AI agents are blockedAssistants asked about this page will guess, or cite someone else. Deliberate for a paywalled publisher. Otherwise this is the biggest gap on the list.
  • low
    No llms.txtOptional, but a short curated index at /llms.txt is the cheapest way to tell assistants what matters on this site.

Every agent

AgentTyperobots.txtFetchWordsResult
Human (Chrome)YouHumanallowed4021254 ms0blocked
HTTP 402
GooglebotGoogle · Classic search index. Also feeds AI Overviews.Search engineallowedAllow: /Timed outunknown
Fetch failed: Timed out
Google-ExtendedGoogle · Robots-only token controlling Gemini training use. Never fetches on its own.AI trainingdisallowedDisallow: /n/ablocked
robots.txt Disallow: / (group: Google-Extended)
GPTBotOpenAI · Training crawler for OpenAI models.AI trainingdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: GPTBot)
OAI-SearchBotOpenAI · Indexes for ChatGPT search results and citations.AI search indexdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: OAI-SearchBot)
ChatGPT-UserOpenAI · Live fetch when a user asks ChatGPT about a page.AI assistant (live fetch)disallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: ChatGPT-User)
ClaudeBotAnthropic · Training crawler for Claude models.AI trainingdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: ClaudeBot)
Claude-SearchBotAnthropic · Indexes for Claude web search.AI search indexdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: Claude-SearchBot)
Claude-UserAnthropic · Live fetch when a user asks Claude about a page.AI assistant (live fetch)disallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: Claude-User)
PerplexityBotPerplexityAI search indexdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: PerplexityBot)
Perplexity-UserPerplexityAI assistant (live fetch)allowedTimed outunknown
Fetch failed: Timed out
BingbotMicrosoft · Bing index. Also feeds Copilot and ChatGPT search.Search engineallowedTimed outunknown
Fetch failed: Timed out
Applebot-ExtendedApple · Robots-only token for Apple Intelligence training.AI trainingdisallowedDisallow: /n/ablocked
robots.txt Disallow: / (group: Applebot-Extended)
Meta-ExternalAgentMetaAI trainingallowedTimed outunknown
Fetch failed: Timed out
AmazonbotAmazon · Alexa and Amazon AI.AI trainingallowedTimed outunknown
Fetch failed: Timed out
BytespiderByteDance · Known for ignoring robots.txt.AI trainingdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: Bytespider)
CCBotCommon Crawl · Open dataset used to train most LLMs.AI trainingdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: CCBot)

Snapshot

Measured 2026-09-25 09:33 UTC from a residential connection. Sites change their rules often, and sites that verify crawler IPs answer differently to different networks. Run a live comparison to see what npr.org does right now, from Cloudflare's network.

← dw.com (29)All 100 publishersjapantimes.co.jp (15) →