28081999.me

A daily census · reading 2026-08-25

How the web treats machines.

Every day, 2,110 fixed origins are asked two questions: do you publish anything written for agents, and do you let agents read you at all. Published adoption estimates for these files disagree by two orders of magnitude. This one states its method, lists its hits, and repeats itself forever.

26.85% of the top 1,000 block CCBot

Who invites agents in — 2026-08-25 14,770 probes

Origins publishing files written for machines rather than people. Measured against a fixed cohort — a reading that cannot be reconstructed later.

SignalWeb head
Tranco top 1,000
n=1,000
Web tail
sampled to rank 100k
n=1,000
AI-native
model & dev platforms
n=110
Any agent-facing signal 10.8% (108)5.8% (58)74.55% (82)
llms.txt (apex domain) 8.8% (88)4.9% (49)58.18% (64)
llms.txt (docs subdomain) 3.4% (34)1% (10)43.64% (48)
.well-known/mcp.json 0.5% (5)0.3% (3)7.27% (8)
.well-known/agents.txt 0.2% (2)0.2% (2)0% (0)
ai.txt 0.2% (2)0.1% (1)0% (0)
ai-plugin.json (deprecated) 0.6% (6)0.3% (3)0.91% (1)

Who shuts agents out robots.txt

Share of origins whose robots.txt disallows each crawler at the site root. Training crawlers and answer-engine fetchers are separated, because refusing to be trained on is a different decision from refusing to be cited.

CrawlerOperatorPurposeWeb head
Tranco top 1,000
Web tail
sampled to rank 100k
AI-native
model & dev platforms
GPTBot OpenAI training 23.85% (119)22.92% (116)2.11% (2)
ChatGPT-User OpenAI user-fetch 15.63% (78)8.7% (44)1.05% (1)
OAI-SearchBot OpenAI search 13.23% (66)7.51% (38)1.05% (1)
ClaudeBot Anthropic training 24.25% (121)21.15% (107)2.11% (2)
Claude-User Anthropic user-fetch 15.43% (77)6.32% (32)1.05% (1)
CCBot Common Crawl training 26.85% (134)21.94% (111)3.16% (3)
Google-Extended Google training 21.04% (105)20.75% (105)3.16% (3)
PerplexityBot Perplexity search 19.04% (95)8.7% (44)1.05% (1)
Bytespider ByteDance training 25.65% (128)22.73% (115)4.21% (4)
Applebot-Extended Apple training 20.24% (101)18.58% (94)1.05% (1)
meta-externalagent Meta training 22.04% (110)19.76% (100)2.11% (2)
Amazonbot Amazon search 18.64% (93)20.36% (103)2.11% (2)

Denominator is origins that served a parseable robots.txt: head 499/1,000 · tail 506/1,000 · ai_native 95/110. A site counts as blocking when the group that applies to that crawler disallows the root, with an exact user-agent match taking precedence over the wildcard group.

Roughly half the ranked strata return no robots.txt, and that is expected rather than a gap: Tranco ranks by DNS resolution volume, so its head contains a great many CDN and infrastructure hostnames — akamaized.net, googleapis.com, apple-dns.net, gtld-servers.net — which are not browsable websites and correctly serve nothing at the apex. Rates are computed over origins that answered, never over the whole stratum, so infrastructure hostnames cannot dilute the result.

Why published numbers disagree

  1. 01cause

    Who you sample decides the answer.

    The head of the web, its long tail, and the AI-native platforms are three different worlds. A survey of one, reported as a fact about the web, is how you get figures an order of magnitude apart. This census keeps the strata separate and never blends them into one headline.

  2. 02cause

    These files usually live on the docs subdomain.

    Probing apex domains alone undercounts adoption. On this reading anthropic.com, perplexity.ai, fireworks.ai, deepinfra.com, anyscale.com and 13 others publish an llms.txt that an apex-only survey scores as absent. This census probes both, and counts a 200 only when the body validates as the file it claims to be — an HTML page served with status 200 is not an llms.txt.

The series building

This is the first reading at this scale. The value of the layer is entirely in its repetition — one measurement is a fact, a year of them is a trend nobody else holds. A new reading is taken every morning and frozen.

Method

Panel v2 · 14,770 probes · 2,450 errors · 3,344 hosts that do not exist.

Fixed stratified cohort. Each domain is probed for six agent-facing files, on both the apex and the docs subdomain where applicable, plus robots.txt. A 200 counts only when the body validates as the declared type — an HTML page returned with status 200 is not an llms.txt. A subdomain that does not resolve is recorded as absent, not as an error. robots.txt is parsed with exact user-agent precedence over the wildcard group; a group counts as blocking when it disallows the root.

curl https://28081999.me/api/census/latest.json
curl https://28081999.me/api/census/series.json
curl https://28081999.me/api/census/2026-08-25.json