We asked 67 companies for an llms.txt file, and not one publisher or retailer had written one.
The people most worried about AI eating their traffic have not written the file that is supposed to help.
llms.txt is a proposed standard: a markdown index at your root telling a language model what your site contains and where the good pages are. Two years in, nobody has measured who actually publishes one. So we did, on 6 September 2026, and re-ran it from scratch rather than reprint an earlier count. Here are the 7 lessons.
Check who has actually published one
A few domains refused our request outright — those are unknowns, not absences. Of the ones that answered, 53% returned a real file.
Split by what the company sells, the pattern is not subtle:
| What they sell | Publish one |
|---|---|
| SaaS | 88% |
| Developer tools | 81% |
| AI labs | 60% |
| SEO tools | 25% |
| Publishers | none |
| Retailers | none |
The New York Times, the BBC, the Guardian, Wired, CNN, Healthline, wikiHow, The Verge. Amazon, eBay, Etsy, Walmart, Booking.com, Airbnb, Zillow. Some of the most-read commercial sites on the web, and not one llms.txt between them.
If this file were a defensive measure against AI, that is exactly backwards.
So before you plan yours, look at who in your own market has bothered. It takes an afternoon to check a list of competitors, and if the answer is nobody, you have learned something more useful than the file would have told you.
Put it on your docs, not marketing
Here is the finding that explains the whole thing.
anthropic.com/llms.txt returns a 404. docs.anthropic.com/llms.txt returns a substantial file.
The same split holds across the board — Stripe, Perplexity, GitHub, Cloudflare, Mistral and Cohere all publish on the docs subdomain and not on the marketing one. Nearly every documentation subdomain we checked has a file; the marketing domain in front of it usually does not.
An AI company whose own product would consume this file does not put one on its homepage. It puts one on its docs.
Placement is the clearest statement of purpose available. This is a machine-readable table of contents for technical documentation, and the companies publishing it are the ones whose docs a model genuinely needs to navigate.
So put yours where your reference material is. If you serve docs from a subdomain, that is the root that needs the file — publishing it on your marketing homepage indexes the one part of your site a model was never going to get lost in.
Do not generate it from your sitemap
The median file we found is about a page of links. Perfectly reasonable.
Then there is Twilio's, at 2.3 MB and over 9,000 links, with Zendesk, Mailchimp and Dropbox not far behind.
A table of contents with 9,000 entries is not a table of contents. It is the site, flattened, in a file whose entire premise is that it should be short enough to fit in a context window ahead of the pages themselves.
Ours is 16 links. That is not modesty, it is the format working as intended — and Squarespace, Supabase, Netlify and Vercel all sit in the same range.
If you are about to generate yours from your sitemap, that is the failure mode you are walking into. Indexing everything is what the XML sitemap is already for, and it carries a hard size limit precisely because machines choke on unbounded lists.
That is free and it measures the AI side. Keep watching the search side too, because that is where the traffic still is — Semrush and SE Ranking both track it properly, and neither has anything to say about llms.txt either.
Write yours by hand instead. If you cannot list your own important pages without a script, that is worth knowing on its own.
Discount advice the adviser has not taken
This one is uncomfortable, and we are in it.
Of the SEO tool vendors we checked, only Semrush and ourselves publish one. Ahrefs, Moz, SE Ranking, Surfer, Sitebulb and Screaming Frog all return a 404.
An industry that sells advice about being visible to machines has a 25% adoption rate on the file it keeps recommending. Read that either way you like — as hypocrisy, or as the people closest to the data quietly concluding it does not matter yet. We think it is mostly the second.
Apply the same test to whatever you are told to do next. When somebody who would know has not done the thing themselves, that silence is data, and it is usually cheaper to trust than the article.
An llms.txt file cannot be shown to do anything yet. Plenty of AI visibility work can. If you want to be found by both Google and the assistants, we will tell you honestly which jobs are worth paying for.
- We measure your AI visibility properly
- We only recommend work we can measure
- SEO and GEO delivered by one team
- Free call, no obligation
Never report this file as a result
Now the part every article about this file skips.
llms.txt is a proposal. No search engine and no major model provider publishes documentation saying it reads the file, ranks with it, or treats it as an input. There is no conformance test, no reporting, and no way from the outside to tell whether a single answer you have ever seen was shaped by one.
We cannot measure that it does anything. Neither can anyone else, which is why nobody has shown you a before-and-after. What you can measure is whether AI assistants currently name and link your brand at all — that is what our AI Visibility Checker reports, and it is the number that would have to move for this file to have earned its place.
Publish one if you like. Just do not put it in your quarterly review as work that produced a result, because you will not be able to defend it when somebody asks what changed.
Take the reading before you publish, and take it again afterwards. That is the whole difference between doing this and knowing whether it mattered.
Write one only if you have docs
The instruction, and it follows directly from where the files actually are.
If you run documentation — an API reference, a knowledge base, a help centre with hundreds of pages and a structure a stranger cannot guess — an llms.txt is a genuinely useful artefact. You are handing a model a map of a maze you built.
If you run a marketing site with 40 pages, a blog and a pricing page, you are describing a room the model can already see. Nothing in our data suggests it helps, and the companies best placed to know have not done it.
Spend the afternoon on the thing that is measurable instead. Run your top pages through the Content Optimizer, or find what your site is not ranking for with Keyword Research. Both change numbers you can see.
Keep it under 20 KB
If you write one anyway, here is the shape the working examples share.
An H1 with the product name. A blockquote of one sentence saying what it is. Then linked sections — Docs, API, Guides, Changelog — with a handful of entries each and a short description per link. That is it, and it lands between 2 KB and 20 KB.
Every file we found over 100 KB was a dump. Every file under 20 KB was written by a person. The format has no enforcement, so the only thing keeping it useful is the restraint of whoever generated it.
And check the file actually serves as text. Several sites in our sample answer /llms.txt with a styled HTML 404 page and a 200 status — which is how you end up believing you have published one when you have not.
No one can show that llms.txt changes an AI answer. You can, however, measure the answers themselves — which assistants name your brand, and whether any of them link you.
- 6 buyer questions, asked without your brand name in them
- Mentions and citations counted separately
- 3 checks a day, no signup, no card
What we could not measure
And here is what this survey does not show. A couple of the tools named on this site are partners of ours — if you buy through those links we earn a commission, and you do not pay a cent more.
We cannot tell you whether llms.txt works. Nothing here is a controlled test. We measured who publishes the file, where they put it and how big it is — not whether a single AI answer changed because of one. If somebody shows you a causal claim about this file, ask what the control group was.
5 of the 67 domains returned a 403 to our fetcher — npmjs.com, gitlab.com, canva.com, openai.com and perplexity.ai. Those are unknowns, so we excluded them from the rates rather than counting them as absences. One more answered with a 402. That is why the headline counts 62 domains rather than 67.
A 200 is not a file. Many sites answer /llms.txt with a styled HTML 404 page, so we sniff the first 400 bytes for markup and reject those. Counting status codes alone would have inflated this survey substantially.
The sample is 67 domains we chose across 6 categories, not a random sample of the web, and it is weighted towards companies an SEO audience would recognise. A different 67 would give a different number. The method is published in scripts/survey-ai-crawler-files.mjs in our repo so you can run it against your own list.
And one honest correction: an earlier version of this article reported figures from a run on 22 August 2026 whose raw data we did not keep. We could not reproduce it, so the whole survey was re-run from scratch for this version and every number above is from 6 September 2026.
Frequently asked questions
What is llms.txt?
A proposed standard: a markdown file at your site's root that gives a language model an index of what the site contains and where the important pages are. It is a proposal, not a specification any search engine or model provider has committed to.
How many sites actually have an llms.txt?
33 of the 62 domains that answered us — 53% — measured 6 September 2026 across 67 requests. SaaS leads at 88% and developer tools at 81%. Publishers are 0 of 9 and retailers 0 of 7.
Do publishers use llms.txt?
None of the 9 we checked: the New York Times, BBC, Guardian, Wired, CNN, The Verge, Healthline, Investopedia and wikiHow. Nor do any of the 7 retailers. The companies most exposed to AI answers have not published the file.
Does Google use llms.txt?
No search engine or model provider publishes documentation saying it reads the file as a ranking or retrieval input, and there is no way from outside to verify that any answer was shaped by one. We measured adoption, not effect, and nobody has published a controlled test of effect.
How big should an llms.txt be?
The median file we found is 13K bytes and the well-formed ones sit between 2 KB and 20 KB. The largest is Twilio's at 2.3M bytes with 9,100 links — a site dump rather than an index, and the format's main failure mode.
Where should the file live?
On your documentation host, if you have one. anthropic.com returns a 404 while docs.anthropic.com serves 74K bytes, and 7 of the 8 documentation subdomains we checked publish one. The file is a map of technical docs, not a marketing asset.
Do SEO tools publish an llms.txt?
2 of 8: Semrush at 15K bytes and seotools.pro at 2,200. Ahrefs, Moz, SE Ranking, Surfer, Sitebulb and Screaming Frog all return a 404 — a 25% adoption rate in the industry that writes most of the articles recommending it.
Should I write one?
Only if you publish documentation a stranger could not navigate unaided. For a 40-page marketing site there is no evidence it does anything, and the companies with the most to gain from it being real have mostly not bothered.
The one number to take away
0 of 16.
That is how many publishers and retailers in this survey have written an llms.txt — the 16 companies whose business is most directly threatened by an AI answering instead of linking.
They are not ignorant of the problem. As we found in a companion survey the same week, 9 of those same 9 publishers block AI crawlers in robots.txt. They have read the situation carefully and reached the opposite conclusion to the one the industry keeps recommending: not "help the model index me better", but "do not index me at all".
You do not have to agree with them. But before you spend an afternoon writing a file because a blog post told you to, notice that the people with the most at stake, who have thought about this the longest, went the other way.
Measured 6 September 2026. The survey requested /llms.txt from 67 domains across 6 categories, counted a file as present only when a non-HTML body came back with a 200, and excluded the 5 domains that returned a 403. The method is scripts/survey-ai-crawler-files.mjs in our repo, and the raw results are in evidence/surveys/, so you can re-run it against your own list and check every figure here.