Two Webs

Publishers / telegraph.co.uk

8

Invisible to machines

1 of 14 AI agents get the full page at https://telegraph.co.uk/. Checked 2026-09-25 09:33 UTC.

telegraph.co.uk scores 8 out of 100 for machine visibility, number 82 of the 100 content publishers checked on 2026-09-25. Even a plain request carrying a Chrome user agent was refused (HTTP 402), so telegraph.co.uk serves its pages only to clients that pass a full browser check. Every crawler below got the same treatment, which is why the score is so low. All 13 refusals are written in robots.txt, which is the transparent way to do it. Bytespider received a bot-management challenge instead of the page. Googlebot itself is unknown: Fetch failed: Timed out. There is no llms.txt.

youfull pagedegradedblockedunknown

Human

Status
402 → https://www.telegraph.co.uk/
Title
none
H1
Access Issue Help
Words
118
Scripts
0 tags · 0 KB
Meta robots
none
Canonical
none
Server
hidden

Machine

robots.txt
59 groups · 6 sitemaps
llms.txt
missing
llms-full.txt
missing
X-Robots-Tag
none
Blocked
GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Applebot-Extended, Meta-ExternalAgent, Amazonbot, Bytespider, CCBot
Degraded
none
Challenged
Bytespider

Gaps

  • high
    Humans cannot load this page eitherHTTP 402. Every other check is measured against a broken baseline.
  • high
    13 of 14 AI agents are blockedAssistants asked about this page will guess, or cite someone else. Deliberate for a paywalled publisher. Otherwise this is the biggest gap on the list.
  • medium
    Bot-management challenge served to crawlersBytespider received an interstitial instead of the page. A robots.txt rule is an honest block. A challenge hides the page without saying so.
  • low
    No llms.txtOptional, but a short curated index at /llms.txt is the cheapest way to tell assistants what matters on this site.
  • info
    Google-Extended not listedYou block other training crawlers but Gemini training via Google-Extended is still allowed. Add it to the same group if the policy is meant to be even-handed.

Every agent

AgentTyperobots.txtFetchWordsResult
Human (Chrome)YouHumanallowed4021034 ms118blocked
HTTP 402
GooglebotGoogle · Classic search index. Also feeds AI Overviews.Search engineallowedTimed outunknown
Fetch failed: Timed out
Google-ExtendedGoogle · Robots-only token controlling Gemini training use. Never fetches on its own.AI trainingallowedAllow: /n/avisible
Robots-only token; allowed by robots.txt
GPTBotOpenAI · Training crawler for OpenAI models.AI trainingdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: GPTBot)
OAI-SearchBotOpenAI · Indexes for ChatGPT search results and citations.AI search indexdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: OAI-SearchBot)
ChatGPT-UserOpenAI · Live fetch when a user asks ChatGPT about a page.AI assistant (live fetch)disallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: ChatGPT-User)
ClaudeBotAnthropic · Training crawler for Claude models.AI trainingdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: ClaudeBot)
Claude-SearchBotAnthropic · Indexes for Claude web search.AI search indexdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: Claude-SearchBot)
Claude-UserAnthropic · Live fetch when a user asks Claude about a page.AI assistant (live fetch)disallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: Claude-User)
PerplexityBotPerplexityAI search indexdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: PerplexityBot)
Perplexity-UserPerplexityAI assistant (live fetch)disallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: Perplexity-User)
BingbotMicrosoft · Bing index. Also feeds Copilot and ChatGPT search.Search engineallowedTimed outunknown
Fetch failed: Timed out
Applebot-ExtendedApple · Robots-only token for Apple Intelligence training.AI trainingdisallowedDisallow: /n/ablocked
robots.txt Disallow: / (group: Applebot-Extended)
Meta-ExternalAgentMetaAI trainingdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: Meta-ExternalAgent)
AmazonbotAmazon · Alexa and Amazon AI.AI trainingdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: Amazonbot)
BytespiderByteDance · Known for ignoring robots.txt.AI trainingdisallowedDisallow: /403 challenge610 ms16blocked
robots.txt Disallow: / (group: Bytespider)
CCBotCommon Crawl · Open dataset used to train most LLMs.AI trainingdisallowedDisallow: /Timed outblocked
robots.txt Disallow: / (group: CCBot)

Snapshot

Measured 2026-09-25 09:33 UTC from a residential connection. Sites change their rules often, and sites that verify crawler IPs answer differently to different networks. Run a live comparison to see what telegraph.co.uk does right now, from Cloudflare's network.

← japantimes.co.jp (15)All 100 publisherspolitico.com (5) →