Two Webs

Publishers / theintercept.com

67

Two different webs

7 of 14 AI agents get the full page at https://theintercept.com/. Checked 2026-09-25 09:33 UTC.

theintercept.com scores 67 out of 100 for machine visibility, number 25 of the 100 content publishers checked on 2026-09-25. 7 of 14 AI agents receive the full page: 3 of 8 training crawlers, 2 of 3 AI search indexes and 2 of 3 assistant fetchers. 6 are refused in robots.txt and 1 is refused at the HTTP level. The pattern is "cite me, don't train on me": 5 of 8 training crawlers are refused while 4 of 6 search and assistant fetchers get through. Bytespider received a bot-management challenge instead of the page. There is no llms.txt.

youfull pagedegradedblockedunknown

Human

Status
200
Title
The Intercept
H1
Flock Wants the Most Detailed Map of Its Surveillance Cameras Taken Offline
Words
846
Scripts
18 tags · 5 KB
Meta robots
index, follow, max-image-preview:large, max-snippet:-1, max-video-preview:-1
Canonical
https://theintercept.com/
Server
nginx

Machine

robots.txt
9 groups · 1 sitemaps
llms.txt
missing
llms-full.txt
missing
X-Robots-Tag
none
Blocked
Google-Extended, GPTBot, ChatGPT-User, PerplexityBot, Amazonbot, Bytespider, CCBot
Degraded
none
Challenged
Bytespider

Gaps

  • medium
    7 of 14 AI agents are blockedGoogle-Extended: robots.txt Disallow: / (group: Google-Extended) · GPTBot: robots.txt Disallow: / (group: GPTBot) · ChatGPT-User: robots.txt Disallow: / (group: ChatGPT-User) · PerplexityBot: robots.txt Disallow: / (group: PerplexityBot) · Amazonbot: robots.txt Disallow: / (group: Amazonbot) · Bytespider: Bot challenge served (HTTP 429) · CCBot: robots.txt Disallow: / (group: CCBot)
  • medium
    Bot-management challenge served to crawlersBytespider received an interstitial instead of the page. A robots.txt rule is an honest block. A challenge hides the page without saying so.
  • low
    No llms.txtOptional, but a short curated index at /llms.txt is the cheapest way to tell assistants what matters on this site.

Every agent

AgentTyperobots.txtFetchWordsResult
Human (Chrome)YouHumanallowedAllow: (empty)200140 ms846visible
HTTP 200, 846 words
GooglebotGoogle · Classic search index. Also feeds AI Overviews.Search engineallowedAllow: (empty)200138 ms846visible
HTTP 200, 846 words
Google-ExtendedGoogle · Robots-only token controlling Gemini training use. Never fetches on its own.AI trainingdisallowedDisallow: /n/ablocked
robots.txt Disallow: / (group: Google-Extended)
GPTBotOpenAI · Training crawler for OpenAI models.AI trainingdisallowedDisallow: /200144 ms846blocked
robots.txt Disallow: / (group: GPTBot)
OAI-SearchBotOpenAI · Indexes for ChatGPT search results and citations.AI search indexallowedAllow: (empty)200128 ms846visible
HTTP 200, 846 words
ChatGPT-UserOpenAI · Live fetch when a user asks ChatGPT about a page.AI assistant (live fetch)disallowedDisallow: /200155 ms846blocked
robots.txt Disallow: / (group: ChatGPT-User)
ClaudeBotAnthropic · Training crawler for Claude models.AI trainingallowedAllow: (empty)200153 ms846visible
HTTP 200, 846 words
Claude-SearchBotAnthropic · Indexes for Claude web search.AI search indexallowedAllow: (empty)200157 ms846visible
HTTP 200, 846 words
Claude-UserAnthropic · Live fetch when a user asks Claude about a page.AI assistant (live fetch)allowedAllow: (empty)200159 ms846visible
HTTP 200, 846 words
PerplexityBotPerplexityAI search indexdisallowedDisallow: /200123 ms846blocked
robots.txt Disallow: / (group: PerplexityBot)
Perplexity-UserPerplexityAI assistant (live fetch)allowedAllow: (empty)200130 ms846visible
HTTP 200, 846 words
BingbotMicrosoft · Bing index. Also feeds Copilot and ChatGPT search.Search engineallowedAllow: (empty)200163 ms846visible
HTTP 200, 846 words
Applebot-ExtendedApple · Robots-only token for Apple Intelligence training.AI trainingallowedAllow: (empty)n/avisible
Robots-only token; allowed by robots.txt
Meta-ExternalAgentMetaAI trainingallowedAllow: (empty)200167 ms846visible
HTTP 200, 846 words
AmazonbotAmazon · Alexa and Amazon AI.AI trainingdisallowedDisallow: /200140 ms846blocked
robots.txt Disallow: / (group: Amazonbot)
BytespiderByteDance · Known for ignoring robots.txt.AI trainingallowedAllow: (empty)429 challenge119 ms9blocked
Bot challenge served (HTTP 429)
CCBotCommon Crawl · Open dataset used to train most LLMs.AI trainingdisallowedDisallow: /200159 ms846blocked
robots.txt Disallow: / (group: CCBot)

Snapshot

Measured 2026-09-25 09:33 UTC from a residential connection. Sites change their rules often, and sites that verify crawler IPs answer differently to different networks. Run a live comparison to see what theintercept.com does right now, from Cloudflare's network.

← techcrunch.com (67)All 100 publisherscorriere.it (66) →