Two Webs

Publishers / huffpost.com

50

Two different webs

0 of 14 AI agents get the full page at https://huffpost.com/. Checked 2026-09-25 09:33 UTC.

huffpost.com scores 50 out of 100 for machine visibility, number 54 of the 100 content publishers checked on 2026-09-25. None of the 14 AI agents receive the page. All 14 refusals are written in robots.txt, which is the transparent way to do it. There is no llms.txt.

youfull pagedegradedblockedunknown

Human

Status
200 → https://www.huffpost.com/
Title
HuffPost - Breaking News, Politics, Entertainment & Opinion
H1
HUFFPOST
Words
2,307
Scripts
27 tags · 26 KB
Meta robots
max-snippet:-1,max-image-preview:large,max-video-preview:-1
Canonical
https://www.huffpost.com/
Server
nginx

Machine

robots.txt
33 groups · 5 sitemaps
llms.txt
missing
llms-full.txt
missing
X-Robots-Tag
none
Blocked
Google-Extended, GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Applebot-Extended, Meta-ExternalAgent, Amazonbot, Bytespider, CCBot
Degraded
none
Challenged
none

Gaps

  • high
    14 of 14 AI agents are blockedAssistants asked about this page will guess, or cite someone else. Deliberate for a paywalled publisher. Otherwise this is the biggest gap on the list.
  • low
    No llms.txtOptional, but a short curated index at /llms.txt is the cheapest way to tell assistants what matters on this site.
  • info
    Canonical points elsewhereCanonical is https://www.huffpost.com/. Indexes will credit that URL, not this one.

Every agent

AgentTyperobots.txtFetchWordsResult
Human (Chrome)YouHumanallowed200870 ms2,307visible
HTTP 200, 2307 words
GooglebotGoogle · Classic search index. Also feeds AI Overviews.Search engineallowedAllow: /2001154 ms2,307visible
HTTP 200, 2307 words
Google-ExtendedGoogle · Robots-only token controlling Gemini training use. Never fetches on its own.AI trainingdisallowedDisallow: /n/ablocked
robots.txt Disallow: / (group: Google-Extended)
GPTBotOpenAI · Training crawler for OpenAI models.AI trainingdisallowedDisallow: /2001227 ms2,307blocked
robots.txt Disallow: / (group: GPTBot)
OAI-SearchBotOpenAI · Indexes for ChatGPT search results and citations.AI search indexdisallowedDisallow: /2001227 ms2,307blocked
robots.txt Disallow: / (group: OAI-SearchBot)
ChatGPT-UserOpenAI · Live fetch when a user asks ChatGPT about a page.AI assistant (live fetch)disallowedDisallow: /2001140 ms2,307blocked
robots.txt Disallow: / (group: ChatGPT-User)
ClaudeBotAnthropic · Training crawler for Claude models.AI trainingdisallowedDisallow: /2001140 ms2,307blocked
robots.txt Disallow: / (group: ClaudeBot)
Claude-SearchBotAnthropic · Indexes for Claude web search.AI search indexdisallowedDisallow: /2001272 ms2,307blocked
robots.txt Disallow: / (group: Claude-SearchBot)
Claude-UserAnthropic · Live fetch when a user asks Claude about a page.AI assistant (live fetch)disallowedDisallow: /2001171 ms2,307blocked
robots.txt Disallow: / (group: Claude-User)
PerplexityBotPerplexityAI search indexdisallowedDisallow: /200915 ms2,307blocked
robots.txt Disallow: / (group: PerplexityBot)
Perplexity-UserPerplexityAI assistant (live fetch)disallowedDisallow: /2001170 ms2,307blocked
robots.txt Disallow: / (group: Perplexity-User)
BingbotMicrosoft · Bing index. Also feeds Copilot and ChatGPT search.Search engineallowed200867 ms2,307visible
HTTP 200, 2307 words
Applebot-ExtendedApple · Robots-only token for Apple Intelligence training.AI trainingdisallowedDisallow: /n/ablocked
robots.txt Disallow: / (group: Applebot-Extended)
Meta-ExternalAgentMetaAI trainingdisallowedDisallow: /2001177 ms2,307blocked
robots.txt Disallow: / (group: Meta-ExternalAgent)
AmazonbotAmazon · Alexa and Amazon AI.AI trainingdisallowedDisallow: /2001115 ms2,307blocked
robots.txt Disallow: / (group: Amazonbot)
BytespiderByteDance · Known for ignoring robots.txt.AI trainingdisallowedDisallow: /2001096 ms4,353blocked
robots.txt Disallow: / (group: Bytespider)
CCBotCommon Crawl · Open dataset used to train most LLMs.AI trainingdisallowedDisallow: /2001035 ms2,307blocked
robots.txt Disallow: / (group: CCBot)

Snapshot

Measured 2026-09-25 09:33 UTC from a residential connection. Sites change their rules often, and sites that verify crawler IPs answer differently to different networks. Run a live comparison to see what huffpost.com does right now, from Cloudflare's network.

← theage.com.au (50)All 100 publishersslate.com (50) →