llms.txt in the Baltics: a standard adopted by plugins, not people
About one Baltic website in ten publishes an llms.txt. Almost none of their owners chose to, and no AI engine documents reading one.
Somewhere between 19,000 and 20,000 websites in Latvia, Estonia and Lithuania publish an llms.txt file, the proposed âinstructions for AIâ document that is supposed to help assistants understand a site. We know the number because the Baltic Web Audit, our near-census of 293,806 Baltic websites in July and August 2026, asked every one of them for the file and checked what came back.
The interesting part is not the number. It is that almost nobody who has one decided to have one.
In short
Real llms.txt adoption across the Baltics is 12% in Latvia, 10% in Estonia and 11% in Lithuania, and the files are overwhelmingly generated by plugins and site builders rather than written by owners. Google states in writing that Google Search ignores them. The published crawler documentation of OpenAI, Anthropic and Perplexity, checked on 16 September 2026, describes robots.txt, user agents and IP ranges, and says nothing about reading another siteâs llms.txt. The file does have a real job, guiding an agent that is already reading your documentation, and that job is nothing like the one it is being sold for.
First, most published adoption numbers are wrong
Before the finding, the measurement, because this is the metric in the whole study most likely to be quoted incorrectly.
Ask a server for a file that does not exist and a well-behaved server says ânot foundâ. A surprising number do not. They answer âfoundâ to any address whatsoever and serve a page. The audit measured that failure mode directly by asking for addresses that cannot exist: 13.6% of probed Latvian servers (5,257), 9.7% of Estonian (6,230) and 12.1% of Lithuanian (9,024) claim to have a page that does not.
Count llms.txt naively, by trusting a 200 response, and you get 19 to 24% adoption per country. Roughly half of that is those servers. Filter them out, then check that what came back is genuinely a text file rather than a web page wearing a .txt name, and real adoption is 12% in Latvia (4,484 of 38,520 sites), 10% in Estonia (6,249 of 63,982) and 11% in Lithuania (8,374 of 74,501). In a 600-site spot check of what the servers actually sent, 507 came back as plain text and 93 as Markdown, with not one page in disguise.
So: 19,107 real files across three countries. If you see an llms.txt adoption figure roughly double that rate, it was measured by trusting the status code.
For completeness, the rival proposal, ai.txt: 14 sites in Latvia, 30 in Estonia, 90 in Lithuania. It effectively does not exist here.
Who actually put those files there
Reading the bodies answers it. They are overwhelmingly auto-generated, and the generators identify themselves.
The llms.txt proposalâs own site, llmstxt.org, lists the platforms that produce the file automatically: Mintlify and GitBook for documentation sites, Yoast SEO and AIOSEO as WordPress plugins, and Wix, which it describes as generating an llms.txt file for every Wix site. That last one alone explains a slice of the Baltic count, because it is not a decision any of those site owners made. Wixâs own support article says the file helps AI agents âquickly access tools like MCPâ, which is to say it advertises an agent interface that most owners could not name, a detail we ran into from the other direction when we wrote about what Baltic Wix sites do next.
This is the pattern the audit kept finding in every chapter. The AI blocking in the region is mostly a security productâs default. The consent banners are installed and then ignored. Whoever writes the defaults writes the national web, and llms.txt is the cleanest example of it: what looks like a grassroots convention is a handful of vendorsâ release notes.
What the plugins promise
Worth quoting, because it is checkable.
AIOSEOâs feature page, read on 16 September 2026, sells the generator as a way to âshow AI search engines the very best of your siteâ and to âStop letting AI search engines decide what to show from your siteâ, with markdown formatting for âfaster and more efficient indexing by AI-powered enginesâ. Yoastâs page for the same feature is more specific still: the file âhelps AI tools like ChatGPT, Claude, and Gemini understand your site betterâ, and âisnât for search engines. Itâs for AI assistants trying to give accurate answers based on your content.â
Three assistants are named there. All three publish documentation about how their crawlers work. So the claim can simply be checked.
What the engines actually document
Google. Its AI features guide, last updated 2026-07-10 and read on 16 September 2026, is unambiguous: you do not need to create machine-readable files, AI text files, markup or Markdown to appear in Google Search including its generative features, âas Google Search itself doesnât use them.â On the file by name: âItâs completely fine if you decide to create and maintain LLMS.txt files (or other similar files) for other services or systems that use these files. Doing so will neither harm nor help your siteâs visibility or rankings in Google Search, as Google Search ignores them.â
OpenAI. Its bot documentation describes GPTBot, OAI-SearchBot and ChatGPT-User, what each is for, the full user-agent strings, and the published IP ranges at openai.com/gptbot.json. Control is robots.txt. The page mentions llms.txt exactly once, in its own site header, linking to OpenAIâs documentation index.
Anthropic. Its page on web crawling covers ClaudeBot, Claude-User and how to allow or block them in robots.txt. It does not mention llms.txt at all.
Perplexity. Its bot documentation describes PerplexityBot and Perplexity-User, publishes IP ranges at perplexity.com/perplexitybot.json, and walks through allowing them in a firewall. llms.txt appears on the page in a banner offering Perplexityâs own /llms.txt so that an agent can discover Perplexityâs documentation.
That last detail is the whole confusion in miniature, and it happens twice. Two of the three vendors publish an llms.txt for their own documentation. Publishing one and consuming other peopleâs are completely different things, and the first is routinely cited as evidence of the second.
None of this proves no engine reads the file. It proves that on 16 September 2026, four vendors between them documented their crawlers in detail and not one of them said so, while one said the opposite in writing.
Where llms.txt does work
This is not a case for deleting the file, and the proposal is not a scam. It is a case for reading what it was designed to do.
llmstxt.org, now on version two of the proposal, says it plainly: the files âare used most heavily for software documentation, where coding agents follow them to find API references and tutorials.â That is a real and growing use. An agent that has already been pointed at your documentation, by a developer who typed the URL, benefits from a clean Markdown index instead of navigation, ads and JavaScript. Mintlify and GitBook generate the file for that reason and it earns its place there.
What it does not do is make an assistant find you in the first place. Discovery is still crawling, and crawling is still robots.txt, HTML and links. A file that describes your site is useful to something already reading your site.
Our own site publishes an llms.txt because it costs nothing and because a written description of the company is a good thing to have in a fixed place. It is not why anyone cites us.
What to do about it
- If a plugin already generates one, leave it. It will not hurt your rankings, which is the one thing Google has committed to in writing. Switching it off is as much a non-event as switching it on.
- Do not buy it as a service. Anyone charging for llms.txt as the fix for not appearing in AI answers is charging for something no engine has documented using. The same goes for the wider category: Googleâs page says outright that many suggested âhacksâ sold under the AEO and GEO labels are not effective or supported by how Search actually works.
- Use robots.txt for the decisions that bind. That is the file every one of these vendors documents. Whether GPTBot, ClaudeBot and PerplexityBot may read you is a real choice with real consequences, and it is made there.
- If you want to be quoted, the constraint is structural, not technical. We wrote about why assistants cite everyone except you and the short version is that an assistant looks for a source it can trust on a question, and the company selling the thing is the most conflicted source available. No text file changes that.
- Measure before you change anything. Ask ten buyer questions in three assistants, record who got cited instead of you, repeat in a month. That is an afternoon and it is the only way to tell whether anything you did worked.
The Baltic llms.txt number is a small finding with a wide lesson. When 19,107 sites adopt a standard and almost none of their owners know it, the adoption figure is measuring a software release, not a decision. That is worth remembering the next time a percentage is used to argue that everyone is doing something.
If you want the systematic version of the measurement above, run against your competitors as well as you, that is what AI search visibility is. If you would rather run it yourself, the five steps in the piece linked above are complete as written.