Measuring AI Citations Without Fooling Yourself
SEO

Measuring AI Citations Without Fooling Yourself

Cited source sets overlap 34 to 42 percent between days. Any figure from one prompt on one day is noise rather than a result.

Run the same prompt through the same AI assistant on two consecutive days and compare which sources it cited. Research published in April 2026 did exactly that, over an observation window of about 45 days, and found the overlap between cited-source sets ran roughly 34% to 42%. The overlap between sets of mentioned brands was somewhat better at around 45% to 59%.

Sit with those numbers for a moment, because they invalidate most of what currently passes for AI visibility reporting.

If a majority of the cited sources change between one day and the next, then a screenshot is not evidence. A single prompt run once is not a measurement. A month-on-month movement drawn from one observation per month is indistinguishable from noise. And an agency presenting a citation win from a single query has either not read the volatility research or is relying on you not having read it.

This article sets out what a defensible measurement looks like instead. It is deliberately specific about definitions, because definitions are where this category hides its problems.

Mention and citation are not the same thing

Start here, because almost every confused conversation about AI visibility traces back to this distinction being skipped.

A mention is your brand name appearing in generated prose. A citation is your domain appearing as an attributed source, usually as a link. These have different causes, different fixes and different commercial value, and they do not move together. One analysis drawing on a very large prompt base reported that the overlap between brands mentioned and domains cited can fall to around 30% on one major assistant. So a brand can be named constantly and cited rarely, or cited from a third-party page while never being named at all.

The fix for a mention problem is usually entity and reputation work: consistent naming, third-party corroboration, earned coverage. The fix for a citation problem is usually your own page being extractable and authoritative on a specific claim. Reporting them as one number means you cannot tell which problem you have, so you cannot tell which fix to buy.

Six metrics, defined before the first report

Vendors define these terms differently. Some count an unlinked brand mention as a citation. Some count a linked URL only. Some average across platforms without disclosing the weighting. That is why cross-vendor comparison is meaningless unless the definitions travel with the numbers, and why we put the definitions into the engagement document before any measurement begins.

Metric Definition What it tells you
Presence rate Valid runs where the brand is named, divided by all valid runs, per engine and per language Whether the system knows you exist for this question at all
Owned-domain citation rate Runs citing your own domain, divided by all valid runs Whether your own pages are doing the work, or someone else's page is
Independent-source citation rate Runs citing a third party that describes you, divided by all valid runs Whether earned coverage and register entries are carrying your entity
Source share Citations to your domain divided by all citations across the prompt set Your weight relative to everything else the system pulls from
Competitive answer share Your presence rate against each named competitor on a frozen competitor list Whether you are gaining or losing ground, rather than whether the category is growing
Descriptive accuracy Material statements classified as accurate, outdated, unsupported, misleading or false Often the highest-value metric in regulated sectors, and the one most often omitted

That last row deserves emphasis for anyone in financial services. A system can cite you correctly and describe your fund structure, your regulatory status or your service scope wrongly in the same answer. Visibility with an inaccurate description is not a win. In a supervised sector it may be a problem to escalate rather than a metric to celebrate.

Three layers, never one score

Each Layer Answers a Different Question

A single blended visibility number hides which of the three actually moved, which is the one thing the reader needs to know.

Layer one

Where am I in a ranked list?

Organic position, impressions, clicks, click-through rate. Mature, well-understood, and measured against a list of links that the user chooses from.

Layer two

Does a system include me, and where did it get me from?

Presence rate, owned and independent citation rates, source share, descriptive accuracy. Newer, noisier, and requiring repeated sampling to mean anything.

Layer three

Did any of it produce business?

AI referral sessions, self-reported discovery, influenced pipeline in euros. Attribution here is imperfect and should be presented as imperfect rather than modelled and rounded.

Movement in one layer is never reported as evidence of movement in another.

Why ranking well is not the same as being cited

There is now direct evidence on this, and it is the strongest single argument for measuring the two separately.

A study released in May 2026 compared the domains cited in AI Overviews against the organic results for the same queries. It found 29.8% of cited domains appeared nowhere on the corresponding first page of organic results. Overlap with the top five organic results was 25.0%, with the top ten 41.4%, and with the whole first page 70.2%.

Read that from both directions. Nearly a third of citations go to pages that were not on page one. And even the whole first page accounts for only about seven in ten citations. So the source pools overlap substantially without being the same pool, which is exactly why a rank improvement cannot be reported as a citation gain, and why a page can hold position one for years and never appear as a cited source.

That study is a preprint rather than a completed peer-reviewed publication, which is worth saying out loud. The direction it establishes is consistent with everything else in this article, but the specific percentages should be quoted with the source and the date attached.

The number everyone quotes, and what it actually measured

The academic paper that named this field reported visibility improvements of up to 40% from its optimisation methods, across a benchmark of 10,000 queries, published at a major data-mining conference in 2024 and properly peer-reviewed.

That 40% now appears in agency decks as though it were a promised citation-rate gain. It was not. It is a visibility metric measured inside a specific experimental setup, on a specific benchmark, using specific optimisation techniques. It is a legitimate and useful finding. It is not a forecast for your brand, and anyone presenting it as one has skipped the method section.

Why you cannot compare two vendors' numbers

Here is a concrete example of the definitional problem, using two respected measurement providers on the same question.

Panel-based measurement put one assistant's share of the AI chatbot category at roughly 78% worldwide and around 75% across Europe in July 2026. A different provider, using clickstream-based methodology, put the same assistant at roughly 53% to 54% globally in May 2026. That is a gap of more than twenty percentage points on what sounds like the same question.

Neither is wrong. They measure different populations with different instruments: one tracks activity across a network of participating sites, the other reconstructs behaviour from clickstream panels, and they differ on what counts as a session and which surfaces are in scope. The mistake is not choosing one. The mistake is averaging them, or quoting whichever is more convenient, or presenting either as a Luxembourg figure when both are global or continental.

On that last point we hold a line that costs us a talking point. Country-level platform share for Luxembourg alone is thin enough that we do not present European or global share as a Luxembourg number, and where we could not open a country-level source ourselves during research, we left the figure out entirely rather than passing along a number a reader cannot verify.

Concentration, and why one platform is not a strategy

Two further large-sample analyses shape how the prompt set should be built.

One examined roughly 200 million prompts over five months and reported that across every model except one, the combined share of the four most-cited domains rarely exceeded 5% of all citations. Citation is diffuse rather than dominated. The same analysis found only around 11% of domains were cited by two major assistants in common, meaning the source pools of different platforms substantially do not overlap.

The operational consequence is direct. Optimising for one assistant and assuming the gains transfer is not supported by that evidence. Neither is chasing a small number of high-authority placements, since no small set of domains accounts for much of the citation volume. Both of those figures come from a single vendor with a commercial interest, so treat the magnitudes as indicative and the direction as informative.

What has to be on the page

A Report a Finance Reader Can Actually Audit

If any of these is missing, the number in front of you cannot be checked, which means it cannot be relied on either.

Prompt universe version and its change log, so nobody quietly swaps the questions between periods

Engines and product surfaces tested, since a chat interface and a search AI mode are different systems

Languages tested, and the location and personalisation settings used for each run

Number of valid repeated runs, with run-level variance shown wherever the sample allows it

Presence and citation broken out by engine and by language, never blended into one figure

Competitor presence against a competitor list frozen for the reporting year

Every inaccurate, outdated or unsupported claim found, with the correction route proposed

Source-language and source-type distribution across all citations captured

AI referral traffic reported strictly separately from zero-click visibility, because they are different phenomena

Archived raw evidence: prompts, timestamps, full answers and cited URLs, handed over so the client can re-run it

The last item is the test of the other nine. A measurement the client cannot reproduce independently is a claim, not a measurement.

Benchmarks that are not about your market

A great deal of what circulates as AI search data describes the United States or Germany. It gets quoted in European plans without the label. Four examples, each useful and each carrying a geography that must travel with it.

Zero-click behaviour. Clickstream analysis put the share of Google searches ending without a click at 68.01% in the United States over the first four months of 2026, up from 60.45% in 2024, with a European figure of 59.7% for 2024. Note also that zero-click is not the same thing as AI Overview coverage: a search can end without a click for many reasons that have nothing to do with generated answers.

AI Overview coverage. One monitoring provider reported generated answers appearing on 48% of tracked queries in February 2026, up from 31% a year earlier. That is a tracked-keyword sample, not a market census, and the keyword mix determines the result.

Click-through impact. Analysis of more than 100 million German keywords found AI Overviews appearing on roughly 20% of them, with organic click-through rate at position one falling from about 27% to about 11%. That is a substantial and credible finding. It is also Germany, and it is being cited in Luxembourg plans as though it were local. No equivalent measurement of AI Overview prevalence exists for Luxembourg, and we state that as unavailable rather than borrowing the German figure.

AI referral traffic. One provider put AI referrals at 0.2689% of total European web traffic in 2026, up from 0.2023% in 2025, while also showing a decline within the year from 0.2858% in February to 0.2506% in April 2026. Both movements come from the same source. Quoting only the year-on-year rise, or only the intra-year fall, produces two opposite headlines from one dataset. That is a good reason to publish the window alongside the number.

What we will not put in a report

Setting out the exclusions is more informative than another metric list, because each one is a claim we could make and choose not to.

No guarantee that an assistant will include you, since the outputs belong to third-party systems that change between sessions. No claim that structured data causes citation, because Google states no additional technical requirement exists for its AI features and no study establishes the causal link. No claim that a machine-readable policy file guarantees visibility. No European or global platform share presented as a Luxembourg figure. No Luxembourg AI Overview prevalence number, because none has been published. No sales-cycle length unsegmented by company size. No performance fee tied to a one-off citation, which would reward precisely the day-to-day volatility the research documents. And no cost-saving percentage attributed to our own location, because no methodologically comparable figure exists for it, and inventing one would be the easiest unverifiable claim in this industry to make.

Everything in that paragraph is available to us. That is rather the point.


Frequently Asked Questions


Why is a single screenshot not evidence of AI visibility?

Because the sources change. Research published in April 2026, observing repeated runs over a window of about 45 days, found overlap between cited-source sets on consecutive days running roughly 34% to 42%, with mentioned-brand sets overlapping around 45% to 59%. If a majority of cited sources differ between one day and the next, a single observation cannot distinguish a genuine change from ordinary movement. The practical minimum is every prompt run more than once on separate days, with run-level variance shown rather than smoothed away.


What is the difference between a mention and a citation?

A mention is your brand name appearing in generated prose. A citation is your domain appearing as an attributed source. They have different causes and different fixes, and they do not move together: one large-sample analysis reported the overlap between brands mentioned and domains cited falling to around 30% on one major assistant. Mention problems are usually solved with entity consistency and earned third-party coverage. Citation problems are usually solved by your own page being extractable and authoritative on a specific claim. Reported as one number, you cannot tell which you have.


If we rank first, will we be cited?

Not reliably. A study released in May 2026 comparing AI Overview citations against organic results for the same queries found 29.8% of cited domains appeared nowhere on the corresponding first page, with overlap of 25.0% for the top five organic results, 41.4% for the top ten and 70.2% for the whole first page. So the two source pools overlap substantially without being identical. A rank improvement is therefore not evidence of a citation gain, and the two must be reported as separate series. That study is a preprint, so quote the percentages with the source and date attached.


Why do two vendors give completely different platform share numbers?

Because they measure different populations with different instruments. Panel-based measurement put one assistant at roughly 78% of the category worldwide and around 75% across Europe in July 2026, while clickstream-based methodology put the same assistant at roughly 53% to 54% globally in May 2026. Neither is wrong. They differ on what counts as a session and which surfaces are in scope. The error is averaging them, quoting whichever suits the argument, or presenting a global or European figure as a country-level one.


Should we just optimise for the largest assistant?

The evidence argues against it. One analysis of roughly 200 million prompts over five months found that only around 11% of domains were cited by two major assistants in common, meaning the source pools substantially do not overlap. The same analysis found that across nearly every model, the combined share of the four most-cited domains rarely exceeded 5% of all citations, so citation is diffuse rather than concentrated in a few authoritative places. Both figures are single-vendor and commercially interested, so read the direction rather than the exact magnitude, but the direction is consistent: track several platforms and do not bet the programme on one.


Can you tell us the AI Overview prevalence for Luxembourg?

No, because nobody has published one. What exists is a German study of more than 100 million keywords finding AI Overviews on roughly 20% of them with position-one click-through falling from about 27% to about 11%, and a tracked-keyword sample putting coverage at 48% of queries in February 2026 against 31% a year earlier. Those are Germany and a general tracked sample respectively. Borrowing either as a Luxembourg figure would be presenting an unmeasured number as measured, so we state the Luxembourg position as unavailable and keep the foreign figures clearly labelled where they are useful as context.


How should AI referral traffic be reported?

Strictly separately from zero-click visibility, and always with the observation window attached. One provider put AI referrals at 0.2689% of total European web traffic in 2026, up from 0.2023% in 2025, while the same source showed a fall within the year from 0.2858% in February to 0.2506% in April 2026. Quote only the annual rise and you have a growth story. Quote only the intra-year fall and you have a decline story. Both come from one dataset, which is why the window belongs next to the number, and why referral traffic and generated-answer visibility should never be blended into a single figure.

Sources & References:

  • Answer volatility: Schulte, Measuring Visibility in AI Search (GEO), April 2026. Overlap between cited-source sets on consecutive observation days approximately 34% to 42%, mentioned-brand sets approximately 45% to 59%, observation window approximately 45 to 46 days. REPORTED, preprint.
  • Ranking and citation overlap: Measuring Google AI Overviews: Activation, Source Quality and Search Overlap, May 2026. 29.8% of cited domains absent from the corresponding first page of organic results. Overlap 25.0% with top five, 41.4% with top ten, 70.2% with the whole first page. REPORTED, arXiv preprint.
  • Founding GEO framework and the up-to-40% visibility figure: Aggarwal and colleagues, GEO: Generative Engine Optimization, ACM SIGKDD 2024, benchmark of 10,000 queries. Peer-reviewed. The figure is a visibility metric within that experimental setup and is not a citation-rate guarantee.
  • Mention against citation overlap falling to approximately 30% on one major assistant, drawn from a base of approximately 126 million prompts: Semrush AI Visibility Index, cited secondhand. REPORTED.
  • Citation concentration and cross-platform overlap: Evertune analysis of approximately 200 million prompts over five months. Combined share of the four most-cited domains rarely exceeding 5% of citations across models other than one, and approximately 11% of domains cited by two major assistants in common. REPORTED, single vendor with commercial interest.
  • Platform share methodology conflict: Statcounter Global Stats, July 2026, approximately 77.9% worldwide and 75.3% across Europe for the leading assistant. Similarweb, AI Search Stats 2026, published July 2026, approximately 53% to 54% globally for May 2026. Both valid within their own definitions and not to be averaged.
  • Zero-click behaviour: SparkToro, June 2026, using Similarweb clickstream data. 68.01% of United States Google searches ending without a click over the first four months of 2026, up from 60.45% in 2024. European figure 59.7% for 2024. Zero-click is not equivalent to AI Overview coverage.
  • AI Overview coverage of tracked queries: BrightEdge Generative Parser via Search Engine Journal, 48% in February 2026 against 31% in February 2025. REPORTED, tracked-keyword sample rather than a market census.
  • German click-through impact: SISTRIX analysis of more than 100 million German keywords, AI Overviews present on approximately 20% of keywords, position-one organic click-through falling from approximately 27% to approximately 11%. REPORTED. This is Germany, not Luxembourg.
  • AI referral traffic in the European Union: SE Ranking, June 2026. 0.2689% of total web traffic in 2026 against 0.2023% in 2025, with an intra-year decline from 0.2858% in February 2026 to 0.2506% in April 2026. REPORTED.
  • AI feature eligibility depends on ordinary indexability and snippet eligibility with no additional technical requirement: Google Search Central documentation. No study establishes that structured data or a machine-readable policy file causes AI citation.
  • No published figure exists for AI Overview prevalence in Luxembourg search results, and country-level assistant platform share for Luxembourg was not verifiable from the source during research. Both treated as unavailable rather than estimated or substituted with foreign figures.
  • This article is measurement methodology, not legal, financial or investment advice.
0 Comments 0 Comments
0 Comments 0 Comments