Two Webs

Publishers / japantimes.co.jp

15

Invisible to machines

1 of 14 AI agents get the full page at https://japantimes.co.jp/. Checked 2026-09-25 09:33 UTC.

japantimes.co.jp scores 15 out of 100 for machine visibility, number 81 of the 100 content publishers checked on 2026-09-25. Even a plain request carrying a Chrome user agent was refused (HTTP 403), so japantimes.co.jp serves its pages only to clients that pass a full browser check. Every crawler below got the same treatment, which is why the score is so low. All 11 refusals happen at the HTTP level with no matching robots.txt rule, so a crawler cannot tell the block is deliberate. GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Bingbot, Amazonbot, Bytespider, CCBot received a bot-management challenge instead of the page. The raw HTML is an app shell with 3 words; content appears only after JavaScript runs. There is no llms.txt.

youfull pagedegradedblockedunknown

Human

Status
403
Title
Just a moment...
H1
none
Words
3
Scripts
1 tags · 3 KB · app shell
Meta robots
noindex,nofollow
Canonical
none
Server
cloudflare

Machine

robots.txt
none (HTTP 403)
llms.txt
missing
llms-full.txt
missing
X-Robots-Tag
none
Blocked
GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Amazonbot, Bytespider, CCBot
Degraded
none
Challenged
GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Bingbot, Amazonbot, Bytespider, CCBot

Gaps

  • high
    Humans cannot load this page eitherHTTP 403. Every other check is measured against a broken baseline.
  • high
    11 of 14 AI agents are blockedAssistants asked about this page will guess, or cite someone else. Deliberate for a paywalled publisher. Otherwise this is the biggest gap on the list.
  • medium
    Bot-management challenge served to crawlersGPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Bingbot, Amazonbot, Bytespider, CCBot received an interstitial instead of the page. A robots.txt rule is an honest block. A challenge hides the page without saying so.
  • low
    No robots.txtEvery crawler is allowed by default. Add one if you want any say in the matter.
  • low
    No llms.txtOptional, but a short curated index at /llms.txt is the cheapest way to tell assistants what matters on this site.

Every agent

AgentTyperobots.txtFetchWordsResult
Human (Chrome)YouHumanno file403 challenge242 ms3blocked
HTTP 403
GooglebotGoogle · Classic search index. Also feeds AI Overviews.Search engineno file2001867 ms1,694visible
HTTP 200, 1694 words
Google-ExtendedGoogle · Robots-only token controlling Gemini training use. Never fetches on its own.AI trainingno filen/aunknown
Robots-only token; allowed by robots.txt
GPTBotOpenAI · Training crawler for OpenAI models.AI trainingno file403 challenge242 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
OAI-SearchBotOpenAI · Indexes for ChatGPT search results and citations.AI search indexno file403 challenge242 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
ChatGPT-UserOpenAI · Live fetch when a user asks ChatGPT about a page.AI assistant (live fetch)no file403 challenge241 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
ClaudeBotAnthropic · Training crawler for Claude models.AI trainingno file403 challenge241 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
Claude-SearchBotAnthropic · Indexes for Claude web search.AI search indexno file403 challenge242 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
Claude-UserAnthropic · Live fetch when a user asks Claude about a page.AI assistant (live fetch)no file403 challenge241 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
PerplexityBotPerplexityAI search indexno file403 challenge240 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
Perplexity-UserPerplexityAI assistant (live fetch)no file403 challenge240 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
BingbotMicrosoft · Bing index. Also feeds Copilot and ChatGPT search.Search engineno file403 challenge241 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
Applebot-ExtendedApple · Robots-only token for Apple Intelligence training.AI trainingno filen/aunknown
Robots-only token; allowed by robots.txt
Meta-ExternalAgentMetaAI trainingno file2001827 ms1,695visible
HTTP 200, 1695 words
AmazonbotAmazon · Alexa and Amazon AI.AI trainingno file403 challenge240 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
BytespiderByteDance · Known for ignoring robots.txt.AI trainingno file403 challenge241 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)
CCBotCommon Crawl · Open dataset used to train most LLMs.AI trainingno file403 challenge240 ms3blocked
Bot challenge served (HTTP 403, cf-mitigated=challenge)

Snapshot

Measured 2026-09-25 09:33 UTC from a residential connection. Sites change their rules often, and sites that verify crawler IPs answer differently to different networks. Run a live comparison to see what japantimes.co.jp does right now, from Cloudflare's network.

← npr.org (20)All 100 publisherstelegraph.co.uk (8) →