Two Webs

Top 100 / cloud.microsoft

70

Two different webs

1 of 14 AI agents get the full page at https://cloud.microsoft/. Checked 2026-09-25 09:17 UTC.

cloud.microsoft scores 70 out of 100 for machine visibility, number 63 of the 100 most visited websites checked on 2026-09-25. 1 of 14 AI agents receive the full page: 0 of 8 training crawlers, 1 of 3 AI search indexes and 0 of 3 assistant fetchers. Googlebot itself is blocked: HTTP 403 for this user agent. There is no llms.txt.

youfull pagedegradedblockedunknown

Human

Status
200 → https://m365.cloud.microsoft/
Title
Microsoft 365 - Sign into Copilot
H1
Work smarter across Microsoft 365 with Copilot
Words
1,133
Scripts
24 tags · 123 KB
Meta robots
none
Canonical
https://m365.cloud.microsoft
Server
hidden

Machine

robots.txt
none (HTTP 200)
llms.txt
missing
llms-full.txt
missing
X-Robots-Tag
none
Blocked
none
Degraded
GPTBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Meta-ExternalAgent, Amazonbot, Bytespider, CCBot
Challenged
none

Gaps

  • medium
    11 agents get a degraded pageGPTBot: Different content than humans receive · ChatGPT-User: Different content than humans receive · ClaudeBot: Different content than humans receive · Claude-SearchBot: Different content than humans receive · Claude-User: Different content than humans receive · PerplexityBot: Different content than humans receive · Perplexity-User: Different content than humans receive · Meta-ExternalAgent: Different content than humans receive · Amazonbot: Different content than humans receive · Bytespider: Different content than humans receive · CCBot: Different content than humans receive
  • low
    No robots.txtEvery crawler is allowed by default. Add one if you want any say in the matter.
  • low
    No llms.txtOptional, but a short curated index at /llms.txt is the cheapest way to tell assistants what matters on this site.
  • high
    Googlebot is blockedHTTP 403 for this user agent
  • info
    Canonical points elsewhereCanonical is https://m365.cloud.microsoft/. Indexes will credit that URL, not this one.

Every agent

AgentTyperobots.txtFetchWordsResult
Human (Chrome)YouHumanno file200886 ms1,133visible
HTTP 200, 1133 words
GooglebotGoogle · Classic search index. Also feeds AI Overviews.Search engineno file403822 ms35blocked
HTTP 403 for this user agent
Google-ExtendedGoogle · Robots-only token controlling Gemini training use. Never fetches on its own.AI trainingno filen/aunknown
Robots-only token; allowed by robots.txt
GPTBotOpenAI · Training crawler for OpenAI models.AI trainingno file200855 ms1,134partial
Different content than humans receive
OAI-SearchBotOpenAI · Indexes for ChatGPT search results and citations.AI search indexno file200855 ms1,133visible
HTTP 200, 1133 words
ChatGPT-UserOpenAI · Live fetch when a user asks ChatGPT about a page.AI assistant (live fetch)no file200890 ms1,134partial
Different content than humans receive
ClaudeBotAnthropic · Training crawler for Claude models.AI trainingno file200863 ms1,134partial
Different content than humans receive
Claude-SearchBotAnthropic · Indexes for Claude web search.AI search indexno file200874 ms1,134partial
Different content than humans receive
Claude-UserAnthropic · Live fetch when a user asks Claude about a page.AI assistant (live fetch)no file200878 ms1,134partial
Different content than humans receive
PerplexityBotPerplexityAI search indexno file200853 ms1,134partial
Different content than humans receive
Perplexity-UserPerplexityAI assistant (live fetch)no file200884 ms1,134partial
Different content than humans receive
BingbotMicrosoft · Bing index. Also feeds Copilot and ChatGPT search.Search engineno file403836 ms35blocked
HTTP 403 for this user agent
Applebot-ExtendedApple · Robots-only token for Apple Intelligence training.AI trainingno filen/aunknown
Robots-only token; allowed by robots.txt
Meta-ExternalAgentMetaAI trainingno file200865 ms1,134partial
Different content than humans receive
AmazonbotAmazon · Alexa and Amazon AI.AI trainingno file200861 ms1,134partial
Different content than humans receive
BytespiderByteDance · Known for ignoring robots.txt.AI trainingno file200899 ms1,095partial
Different content than humans receive
CCBotCommon Crawl · Open dataset used to train most LLMs.AI trainingno file200871 ms1,134partial
Different content than humans receive

Snapshot

Measured 2026-09-25 09:17 UTC from a residential connection. Sites change their rules often, and sites that verify crawler IPs answer differently to different networks. Run a live comparison to see what cloud.microsoft does right now, from Cloudflare's network.

← t-online.de (72)All 100 top 100samsung.com (70) →