← All posts
17 September 2026

Why the media blocks AI, and what publishers know that you don't

One Latvian media site in five blocks AI crawlers in robots.txt. Almost no other business does. The publishers are right. For most companies it is a mistake.

Most Baltic websites have never said a word about AI crawlers. In the Baltic Web Audit, our near-census of 293,806 websites in Latvia, Estonia and Lithuania in July and August 2026, the robots.txt files of nearly every sector are silent on the question. One sector is not, and it is the one whose business model an AI answer threatens most directly.

What the publishers decided, and why, is worth understanding before you copy them. Their reasons are sound. For most companies they point the other way.

In short

Publishers block AI crawlers because their text is the product: a summary in an answer can replace the visit that pays for the article. In Latvia 21% of media and publishing sites write an AI block into robots.txt (278 of 1,325), against 9% in Estonia and 11% in Lithuania, and no other sector in any of the three countries passes 12%. Some treat the block as a bargaining position rather than a wall: the Guardian leaves out the AI company it has a deal with and refuses the rest. If your text exists to win customers rather than to be sold, the crawlers that put pages into AI answers work for you, and blocking them costs you the answer. Training crawlers are a separate, genuinely optional decision.

The one sector that wrote its answer down

The audit read every robots.txt it could reach and sorted sites by what they are for. Cross the two and one sector stands alone in all three countries:

CountryMedia and publishing sites declaring an AI blockNext highest sector
Latvia21% (278 of 1,325)“other” at 12%, finance at 10%
Lithuania11% (256 of 2,374)finance and “other” at 7%
Estonia9% (117 of 1,337)e-commerce at 4%

Latvian publishers declare it at about twice the Lithuanian rate and more than twice the Estonian. Professional services, the largest sector in every country, sits at 4% in Latvia. Public bodies, hotels and professional firms in Estonia sit at 1%.

This is declared policy, the lines a site writes into its own file. It is a different layer from the one in our piece on sites blocking AI assistants without knowing it, which measured what servers actually do and found that most blocking there is a security product’s default rather than anyone’s choice. Here the owners chose.

What publishers block, read from their own files

The Baltic figures are published as aggregates and name no site. The large international publishers make the same decision in public, so their files show what a deliberate policy looks like. These are the robots.txt files of three of them, read on 17 September 2026:

CrawlerWhat it doesThe New York TimesBBCThe Guardian
GPTBotOpenAI, trainingBlockedBlockedNot named
OAI-SearchBotChatGPT search answersBlockedBlockedNot named
ChatGPT-UserChatGPT fetching a page for a userBlockedBlockedNot named
ClaudeBotAnthropic, trainingBlockedBlockedBlocked
Claude-SearchBotClaude search answersBlockedNot namedBlocked
PerplexityBotPerplexity search resultsBlockedBlockedBlocked
Google-ExtendedGemini training and groundingBlockedBlockedNot named

“Not named” means the crawler falls under the file’s general rules, which do not close off articles. So the Times and the BBC refuse almost every AI crawler by name, search crawlers included. The Guardian refuses Anthropic, Perplexity, Amazon and Meta in a single group of 37 crawlers, and leaves OpenAI and Google-Extended out of it.

The file is a price list, not a wall

The Guardian’s exception is not an oversight. On 14 February 2025 Guardian Media Group announced a partnership with OpenAI under which “Guardian reporting and archive journalism will be available as a news source within ChatGPT, alongside the publication of attributed short summaries and article extracts.” The top of its robots.txt says that use of its content for large language models is not permitted under its terms, and gives an email address for licensing. The last line points to a licence file.

Read together, that is what publishers know. A block in robots.txt is not a way to disappear. It is a way to make access something that is negotiated, one AI company at a time, by an organisation whose text has a price. At the Guardian, the AI company whose crawlers are not named is the one that agreed terms.

It is also worth being exact about what the file can do. The standard that defines robots.txt, RFC 9309, says of its own rules: “These rules are not a form of access authorization.” A robots.txt line is a public request that the documented crawlers say they honour. The publishers’ leverage comes from stating the request in public, not from the file stopping anything.

What even a publisher cannot block this way

One company is missing from the “blocked” story, and the reason matters to every site owner. Google’s documentation for Google-Extended, read on 17 September 2026, describes it as a control over whether content may be used “for training future generations of Gemini models” and for grounding in Gemini Apps and Vertex AI, and states that it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”

So the Times and the BBC blocking Google-Extended does not take them out of AI Overviews. Google’s page on AI features and your website says why: “AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search.” To limit what appears, the page points to nosnippet, data-nosnippet, max-snippet and noindex, which limit ordinary search results in the same stroke.

In Google, then, there is no separate AI door to close. A business that blocks Google-Extended hoping to stay out of AI Overviews has changed nothing it can see in Search.

Why the answer is different for almost everyone else

A newspaper’s article is the thing it sells. When an answer summarises the article well, the reader may never visit, and the visit is where the subscription or the advertising revenue was. Blocking protects the product and strengthens the negotiation.

A service company’s page is a different kind of text. It exists so that someone who needs the service finds the company and gets in touch. When an AI answer quotes that page and names the company, the page has done its job, and there is nothing to protect by refusing. The vendors spell out the cost of refusing in their own documentation, both read on 17 September 2026:

  • OpenAI’s bot documentation: sites opted out of OAI-SearchBot “will not be shown in ChatGPT search answers, though can still appear as navigational links.”
  • Anthropic’s page on crawling: disabling Claude-SearchBot “prevents our system from indexing your content for search optimization, which may reduce your site’s visibility and accuracy in user search results.”

Training is the part a non-publisher can reasonably decide either way, because the vendors keep it separate. OpenAI’s page gives the combination as its own example: “a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training”. Whether your pages help train somebody’s model is a real question about your content. It is not the question of whether you get quoted.

So the part of the publishers’ approach worth copying is not the block. It is that they decided crawler by crawler, in writing, knowing what each one does.

Should you block AI crawlers? Decide by what your text is

Your text isTraining crawlers (GPTBot, ClaudeBot, Google-Extended)Search and user crawlers (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot)
The product: journalism, paid research, course materialBlock, unless you license itBlock, or allow the companies you have terms with
The pitch: services, B2B, a shop, a local businessYour choice, and it does not decide whether you are quotedAllow
Help for existing customers: documentation, support, manualsYour choiceAllow, because this is where customers now ask

If you are in the first row, you are a publisher and your robots.txt is part of your commercial strategy. Everyone else is almost always in the second or third row.

What you can do yourself, and what needs a developer

You can set this alone. Open https://yourdomain.com/robots.txt, see which of the crawlers above it names, and write one group per decision. A site that wants to stay in AI answers but keep its text out of training would add:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

and leave the search crawlers unnamed, so the general rules let them in. It is reversible in a minute.

You need a developer in two cases. The first is when the file says one thing and the server does another, because a CDN, firewall or security plugin is refusing crawlers the file allows. That is the failure our piece on accidental AI blocking shows how to test for, and it is common. The second is when you want part of a page kept out of Google’s snippets and AI features but not the whole page, which means data-nosnippet in the templates rather than a setting.

Why we can say any of this

The sector figures come from the Baltic Web Audit: 293,806 of 306,872 Baltic domains measured across July and August 2026, sector labels reported only where the classifier was at least 0.7 confident, and every figure published as an aggregate that names no site. The method and the full sector chapter are in the report.

If you want to know what your own file and your own server are telling each AI crawler, and which of them actually quotes you, that is what AI search visibility measures.

Tell us what’s broken.
We’ll tell you the truth.

Book a free call →
Reply within one business day · EN / LV
↑↓ navigate · ↵ open · esc close