Two Webs

Top 100 / github.io

95

Mostly one web

12 of 14 AI agents get the full page at https://github.io/. Checked 2026-09-25 09:18 UTC.

github.io scores 95 out of 100 for machine visibility, number 18 of the 100 most visited websites checked on 2026-09-25. 12 of 14 AI agents receive the full page: 6 of 8 training crawlers, 3 of 3 AI search indexes and 3 of 3 assistant fetchers. There is no llms.txt.

youfull pagedegradedblockedunknown

Human

Status
200 → https://pages.github.com/
Title
GitHub Pages | Websites for you and your projects, hosted directly from your GitHub repository. Just edit, push, and your changes are live.
H1
Websites for you and your projects.
Words
700
Scripts
4 tags · 1 KB
Meta robots
none
Canonical
https://pages.github.com/
Server
GitHub.com

Machine

robots.txt
none (HTTP 200)
llms.txt
missing
llms-full.txt
missing
X-Robots-Tag
none
Blocked
none
Degraded
none
Challenged
none

Gaps

  • medium
    This URL is a meta-refresh stubThe HTML only redirects to https://docs.github.com/pages. Browsers follow it; most crawlers do not. Compare that URL instead, and consider a real HTTP 301.
  • low
    No robots.txtEvery crawler is allowed by default. Add one if you want any say in the matter.
  • low
    No llms.txtOptional, but a short curated index at /llms.txt is the cheapest way to tell assistants what matters on this site.
  • info
    Canonical points elsewhereCanonical is https://pages.github.com/. Indexes will credit that URL, not this one.

Every agent

AgentTyperobots.txtFetchWordsResult
Human (Chrome)YouHumanno file200187 ms700visible
HTTP 200, 700 words
GooglebotGoogle · Classic search index. Also feeds AI Overviews.Search engineno file200187 ms700visible
HTTP 200, 700 words
Google-ExtendedGoogle · Robots-only token controlling Gemini training use. Never fetches on its own.AI trainingno filen/aunknown
Robots-only token; allowed by robots.txt
GPTBotOpenAI · Training crawler for OpenAI models.AI trainingno file200187 ms700visible
HTTP 200, 700 words
OAI-SearchBotOpenAI · Indexes for ChatGPT search results and citations.AI search indexno file200186 ms700visible
HTTP 200, 700 words
ChatGPT-UserOpenAI · Live fetch when a user asks ChatGPT about a page.AI assistant (live fetch)no file200186 ms700visible
HTTP 200, 700 words
ClaudeBotAnthropic · Training crawler for Claude models.AI trainingno file200186 ms700visible
HTTP 200, 700 words
Claude-SearchBotAnthropic · Indexes for Claude web search.AI search indexno file200186 ms700visible
HTTP 200, 700 words
Claude-UserAnthropic · Live fetch when a user asks Claude about a page.AI assistant (live fetch)no file200186 ms700visible
HTTP 200, 700 words
PerplexityBotPerplexityAI search indexno file200187 ms700visible
HTTP 200, 700 words
Perplexity-UserPerplexityAI assistant (live fetch)no file200185 ms700visible
HTTP 200, 700 words
BingbotMicrosoft · Bing index. Also feeds Copilot and ChatGPT search.Search engineno file200185 ms700visible
HTTP 200, 700 words
Applebot-ExtendedApple · Robots-only token for Apple Intelligence training.AI trainingno filen/aunknown
Robots-only token; allowed by robots.txt
Meta-ExternalAgentMetaAI trainingno file200186 ms700visible
HTTP 200, 700 words
AmazonbotAmazon · Alexa and Amazon AI.AI trainingno file200185 ms700visible
HTTP 200, 700 words
BytespiderByteDance · Known for ignoring robots.txt.AI trainingno file200186 ms700visible
HTTP 200, 700 words
CCBotCommon Crawl · Open dataset used to train most LLMs.AI trainingno file200185 ms700visible
HTTP 200, 700 words

Snapshot

Measured 2026-09-25 09:18 UTC from a residential connection. Sites change their rules often, and sites that verify crawler IPs answer differently to different networks. Run a live comparison to see what github.io does right now, from Cloudflare's network.

← bit.ly (95)All 100 top 100apache.org (95) →