← All posts
28 July 2026

Why AI assistants cite everyone except you

The usual answer is a checklist of website faults, most of which Google says are not required. The real reason is structural: it decides what can be quoted.

You asked ChatGPT who the best supplier in your sector is. It named four companies in a confident paragraph, cited three sources, and none of them were you. You know two of those companies. You are better than both.

So you went looking for why, and you found a checklist. Fix your NAP consistency. Add schema markup. Publish an llms.txt file. Get more third-party mentions. Make your content answer-first. Ensure your business description is not too generic.

We ran that question three ways across the search-grounded models this week, and the checklist is remarkably consistent: roughly forty distinct sources cited across nine answers, and they broadly agree with each other. They also share two problems. The first is that almost every item on the list assumes the cause is a defect in your website. The second is that not one of those forty sources gives you a way to check whether any of it worked.

The actual reason is structural, it is published, and it changes which pages are worth the effort.

The short answer

AI assistants cite other websites because a company is the weakest available source about itself. Google’s published guidance to its own search raters says that when a website’s account of itself disagrees with reputable independent sources, the raters should trust the independent sources, and it instructs them to search for reputation information with the company’s own domain excluded. An assistant answering “who should I hire” is doing the same job, so the citations land on comparison pages, roundups, directories and independent articles. The practical consequence is that your commercial pages are the least citable things you own: they are conflicted, and they contain an offer rather than an answer, so there is nothing in them to quote. The pages that earn citations are the ones that answer a question, and the ceiling on your own domain is why the strongest move is often to be a source that is not you.

Most of the checklist is not a requirement, and Google says so

Start by deleting the parts you do not need, because they are the parts being sold hardest.

Google’s documentation on AI features and your website is unambiguous. It states: “There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” And, on the file formats that have become a small industry: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features.” The same page says you can apply the same foundational SEO best practices for AI features as you do for Google Search overall.

That is Google’s word on Google’s own AI surfaces, which is the only surface here where anyone publishes guidance at all. It does not settle what ChatGPT or Perplexity do, and neither of those publishes a comparable document.

We maintain structured data and an llms.txt file on this site. Both are cheap, both are good hygiene, and neither is why anything gets cited. If a proposal you are reading leads with schema and llms.txt as the fix for AI invisibility, the person writing it is selling you work that the one published source on the subject says is not required.

The NAP consistency advice is worth a separate note. Keeping your name, address and phone number identical across directories is real local SEO, and if you are a dentist or a plumber it matters. If you are a B2B supplier being asked “who does this well”, it is close to irrelevant, and it is on the list because it was on the previous list. A good deal of AI visibility advice is local SEO advice with the dates changed.

You are the weakest source about yourself

Here is the part almost nobody states plainly.

Google publishes the guidelines it gives to the human raters who grade its search results. The current General Guidelines are dated 11 September 2025. Section 2.5 tells raters how to research who is behind a website, and it says this:

“You must also look for reputation information about the website and/or content creators. What do outside, independent sources say about them? When there is disagreement between what the website or content creators say about themselves and what reputable independent sources say, trust the independent sources.”

Section 3.3.1, on the reputation of a website, is blunter:

“Many websites are eager to tell users how great they are. Your job is to independently evaluate the Page Quality of the website, not just accept information that appears on one or two pages of the website without further verification. Be skeptical of claims that websites make about themselves, particularly when there is a clear conflict of interest.”

And then, in the worked examples, Google literally tells raters to run the search with the company’s own domain removed: [ibm -site:ibm.com], [ibm reviews -site:ibm.com]. The accompanying note says to find sources that were not written or created by the website or the company itself, and gives IBM’s own carefully maintained social accounts as an example of what does not count.

Be precise about what this is. The rater guidelines are not the ranking algorithm and they are not the source-selection system behind AI Overviews. Raters produce evaluation data; they do not move any individual result. What the document is, is the clearest published statement of how the largest company in this space thinks about the question, written for humans, in plain language. Read it as a statement of intent rather than a description of mechanism, and it still says the thing you need to hear: on any question about how good you are, your own site is the source that gets discounted.

An assistant answering a buying question is running that same errand. It needs a source it can stand behind for a claim about a company, and the company selling the thing is the most conflicted source in the set.

There is a second reason, and it is more mundane. Assistants quote. A page that sells contains a proposition, a benefits list and a contact form. Even when it is retrieved, there is no sentence in it that answers the question asked, so there is nothing to lift.

Which of your pages can actually win

We can put numbers on the second half of that, because we record what the engines say about a page before it exists and then go back.

The most recent round covered 25 published pieces on a client property against their own pre-publication baselines: 103 buyer questions across three engines, 309 separate checks. The split was not close.

Eleven articles held 30 of the 31 citations. Fourteen commercial pages, the ones that sell the service, produced one citation between them across 168 checks. Every broad commercial query of the “best supplier in city X” shape returned nothing for those pages, at baseline and again weeks later. Three commercial pages that had held a citation at baseline had lost it. Not one article had.

Two caveats travel with that data and we would rather state them than have you find them. It is one sample per question on one day, and these answers vary between runs, so anything you decide on this basis should be re-run three times on separate days first. And it is one property in one sector, which makes it a direction rather than a law. What makes it worth acting on is that it points the same way as the published guidance, for a reason you can articulate: the commercial page is both the conflicted source and the one with no answer in it. It loses twice.

The consequence is uncomfortable if you have been buying AI visibility work. The page you most want quoted is the page least able to be quoted, and effort spent making your pricing page answer-first is mostly effort spent losing slowly. The work that pays is earning the citation with a page that genuinely answers something, and then routing the reader from there to the page that sells. That is a content architecture question, not a markup question.

What actually moves it

Four things, in the order the evidence supports them.

Answer the question in the first two sentences, on a page whose job is to answer it. Not a build-up, not context, not who we are. The claim, then the qualification. If you have a number and a date, they go in the first sentence. This is the one piece of standard advice that survives contact with the evidence, because it is describing the mechanism directly: a model quoting your page needs a self-contained sentence it can lift without inheriting your marketing.

Fix what the engines already get wrong about you. This is the most neglected item in the category and usually the most urgent. Ask the assistants factual questions about your own company and read the answers properly. Wrong services, a former office, a discontinued product, a director who left in 2023, a merger that never happened. Misinformation about your business is being served to buyers with a confident tone and a citation attached, and unlike the ranking question it has no budget category, no benchmark and no competitor to be compared against. It is also the one thing here you can often correct in a week, by publishing the correct fact somewhere checkable and getting the incorrect source updated.

Publish things that answer, and link them to the things that sell. See the 30-of-31 split. An article that answers a real buyer question is the unit that earns citations; the service page is the destination you route to afterwards. This is also why sorting AI work by department rather than by job type produces such unhelpful advice, and it is the same failure here: the category talks about “content” when the thing that matters is what kind of page it is.

Be a source that is not you. This is the ceiling, and it follows directly from the rater guidelines. There is a hard limit to how far your own domain can move an answer about your own quality, because the discount applied to a self-interested source does not go away when your copy improves. The lever is to be somewhere else: independent research others cite, genuine data nobody else publishes, presence in the comparison pages and forum threads where the citations actually land, and where it fits, an independent property that is useful in its own right. We build and run properties like that ourselves, including snowverdict.com, which answers a dated question with a number and a published method on every page. That last option is not for everyone. It is a second brand to be responsible for and a compliance conversation in some sectors, and plenty of companies are right to want nothing outside their own domain.

The step everyone skips: check

Of the roughly forty sources those assistants cited when we asked why businesses go unmentioned, not one told the reader how to find out whether anything had changed. The whole category is prescriptions with no measurement, which is convenient for the people selling the prescriptions.

Here is the baseline, and it costs you an afternoon rather than a retainer.

  1. Write down ten questions a real buyer asks before choosing a supplier like you. Phrase them the way a person types them, not as keywords. “Who does X well in Y” and “is A or B better for C”, not “X services Y”.
  2. Ask each one in ChatGPT, Perplexity and Gemini. Three engines minimum, because they source differently and a win in one says nothing about the others.
  3. Record three things per answer: were you named, which sources were cited, and is anything said about you wrong.
  4. Run each question three times on separate days. These answers are not deterministic. A single sample is an anecdote.
  5. Put the file somewhere with today’s date on it, and do it again in a month.

The fourth column of that spreadsheet is the one that matters most and the one you will be tempted to skip: who got cited instead of you. It tells you what kind of page wins these questions in your sector, and it is usually not a supplier’s website.

What that exercise will not give you is a click figure. A citation frequently produces no visit at all, which is the honest limitation of this entire field and the reason we treat it as authority work rather than a lead channel with attributable revenue. Anyone promising you attributable AI revenue is describing something they cannot measure either.

We do this on our own site as well as for clients, including the daily agent that manages this site’s search visibility and re-checks its own changes at 14 and 28 days. Some of what it published stalled. That is the point of recording the before-state.

If you want the systematic version of the afternoon described above, run against your competitors as well as you, that is what AI search visibility is. If you would rather just do it yourself, the five steps are complete as written and we would rather you ran them than hired anyone on a checklist nobody measured.

Tell us what’s broken.
We’ll tell you the truth.

Book a free call →
Reply within one business day ¡ EN / LV