Nine Metrics a Citation Report Must Contain
SEO

Nine Metrics a Citation Report Must Contain

Exact definitions, eligible run rules, coding method and the run log fields that make a number reproducible instead of asserted.

Ask a vendor how they measure AI visibility and you will usually get one of two answers. Either a proprietary score, or a rankings chart with the words AI visibility written above it. Both are ways of not answering the question.

The reason this happens is structural rather than dishonest. There is no standards body prescribing what a citation report should contain, and no survey establishing what buyers currently accept as proof. In the absence of a standard, whoever writes the report defines the metric, and a metric you define yourself is difficult to fail.

What follows is a set of nine metrics with exact definitions, plus the reasons each one has to be reported separately rather than blended. This is not a proprietary framework. It is the minimum a buyer should be able to demand, and every definition here is written so that a competing vendor could implement it identically.

Nine metrics, three layers

What a Citation Report Should Contain

Each layer answers a different question. Collapsing them into one score removes the ability to diagnose anything.

Layer one, presence

Are you in the answer at all, and in what capacity?

Mention rateOwned-source citation rateCitation share

Layer two, quality of presence

Is the answer helping you, and is it accurate?

Recommendation rateAnswer placementFactual accuracyCitation validity

Layer three, reliability and outcome

Would you get the same result tomorrow, and did anybody arrive?

StabilityAI referral sessions

Why nine and not one. High visibility attached to wrong facts is a failure, not a result. A brand can be mentioned frequently and recommended never. A citation can be present and irrelevant to the sentence beside it. One number cannot distinguish those cases, which is exactly why single-score reporting is popular.

Definitions written so that any vendor could implement them identically. No proprietary scoring is used.
Created by Arfadia • arfadia.com/blog

The nine, defined

Definitions matter more than names here, because the same word gets used for different things across vendors. What follows is the exact denominator for each.

Metric Exact definition Why it is not a ranking metric
Mention rateShare of eligible prompt runs in which the brand name appears anywhere in the answer textAn assistant can name a brand with no link at all, and independently of its position in any results page
Owned-source citation rateShare of eligible runs containing at least one citation to a domain the client ownsMeasures source attribution directly rather than list position
Citation shareCited URLs belonging to the client divided by all cited URLs across the fixed panelMeasures competition between sources, not competition for a slot
Recommendation rateShare of eligible runs in which the brand is explicitly recommended, shortlisted or compared favourablyA mention can be neutral or unfavourable. Being named is not being endorsed
Answer placementCoded as first recommendation, later recommendation, table inclusion, footnote only, or source onlyPosition inside a generated answer is a different variable from position in a link list
Factual accuracyShare of coded claims about the brand that match a client-approved fact baseVisibility attached to wrong facts is negative value, and no ranking metric detects it
Citation validityShare of citations that actually support the specific claim they are attached toA link can be present and irrelevant to the sentence beside it
StabilityConsistency of outcomes across repeated clean-session runs of the same prompt in the same windowOutputs are probabilistic. One favourable run is not a state of the world
AI referral sessionsSessions attributable to assistant domains or AI surfaces, where measurableReported separately, never merged with citation, because a citation often produces no click

Eligible runs, and why the word matters

Notice that most definitions above say eligible runs rather than all runs. That single qualifier is where a great deal of quiet manipulation lives.

A run is eligible if the prompt actually invoked the behaviour being measured. If a prompt returns a refusal, a clarifying question or an error, including it in the denominator drags every percentage down for reasons unrelated to your visibility. Excluding it is correct. Excluding it selectively, after seeing which way the result went, is not.

So the rule has to be written before the baseline and applied mechanically. Eligibility criteria go in the method document, the count of excluded runs appears in the report, and the reason for each exclusion is recorded. If a vendor cannot tell you how many runs were excluded and why, the percentages have no fixed denominator and are therefore not comparable across months.

The prompt panel is the contract

Everything above depends on a fixed set of questions. Get this part wrong and no amount of metric rigour rescues the programme.

A panel should be written from real buying jobs rather than from keyword tools, because the two produce different language. A buyer asking an assistant to shortlist suppliers phrases it as a task, not as a keyword string. It should be separated by destination country and by buyer role, since a procurement engineer in Germany and a technology director in Singapore are asking different questions about the same company. And it should be versioned, with inclusion rules documented, so that a later disagreement about whether a prompt belonged in the set can be settled by reading rather than arguing.

Then it gets frozen. A baseline is captured before any intervention. From that point, prompts are added only by written agreement, and additions are reported as a new cohort rather than folded into the existing series. Swapping in gentler prompts after a weak month is the single most effective way to make a failing programme look successful, and it is invisible unless the panel is disclosed.

Anatomy of one run

What Has to Be Recorded for a Result to Be Reproducible

If a reported number cannot be rebuilt from the log, it is an assertion rather than a measurement.

Prompt ID
Panel identifier and version, so the exact question can be matched to a documented set rather than reconstructed from memory.
Assistant
Which system, and which surface within it. A web-retrieval mode and a trained-knowledge mode behave differently and should never be pooled.
Model version
Recorded at run time. Model updates are the most common cause of a series breaking for reasons unrelated to any work performed.
Country
Where the query was issued from. Whether this changes source selection is unmeasured, which is exactly why it must be held constant and logged.
Language
And script, where relevant. A Devanagari prompt and a romanised prompt are two prompts, not one.
Session state
Clean session or carried context, signed in or not. Personalisation contaminates comparison if this drifts.
Timestamp
Date and time. Citation behaviour has been observed shifting sharply within weeks.
Raw response
Full answer text, retained. Non-negotiable
Cited URLs
Every URL the answer surfaced, in order, with owned and third-party coded separately.
Eligibility
Whether the run counted toward the denominator, and if not, the documented reason.

Google separately advises being wary of third-party tools that promise ranking success or claim to use internal Google metrics. A retained log is the answer to that concern.
Created by Arfadia • arfadia.com/blog

Volatility is a finding, not a problem to hide

Repeated measurement produces uncomfortable results, and the discomfort is informative. Citation share for a given source type has been observed shifting sharply inside a few weeks as retrieval behaviour changes. A brand can appear in an answer on Monday and be absent on Thursday with nothing on the site having changed.

The wrong response is to report the best run. The right response is to report the spread, and to treat a narrow spread as a result in itself. Stability is the metric that tells a client whether their visibility is a position they hold or a coincidence they caught.

This also disciplines expectations about time. Where an assistant retrieves live from the web, changes can surface relatively quickly. Where an answer draws on a model's own trained knowledge, movement takes considerably longer, because the underlying material has to be recrawled and reincorporated. Practitioner consensus puts the floor for measurable citation movement at roughly ninety days for that reason. Averaging those two situations into a single promised timeline produces a number that describes neither.

Two structural findings worth measuring against

Metrics are only useful if there is something to compare them to. Two findings from the reviewed research give a report somewhere to point, and both survived cross-checking against more than one source.

The first concerns where inside a page citations come from. Multiple analyses find that citations concentrate heavily in the opening portion of a page, with one large-scale study putting roughly 44 percent of citations in the first 30 percent of page content. Independent citation analysis reviewed for this material reached the same directional conclusion, and it matches Arfadia's own internal citation data. This has a direct reporting consequence: if owned-source citation rate is flat while your best material sits three screens down, the diagnosis may be structural rather than authority-related, and answer-first restructuring is a cheaper intervention than a public relations programme.

The second concerns baseline expectations. An India-focused analysis of 2,981 brands across 59,620 prompts between April and June 2026 reported that 19.3 percent met its high AI visibility benchmark, leaving 80.7 percent below it, with an average score around 30 out of 100. That is a single-vendor benchmark using a self-defined scoring method rather than an audited industry statistic, and it should be quoted with that caveat attached. Used carefully it still tells a client something worth knowing: in the Indian market most brands are not yet competing on citation, so a low starting score is normal rather than alarming, and the competitive bar is lower than the volume of marketing noise implies.

Neither finding is a target. Benchmarks borrowed from a different sample tell you where a market sits, not where your programme should be, and a report that leads with somebody else's average instead of your own baseline has substituted context for measurement.

Where owned content stops being enough

One metric in the table above tends to produce an uncomfortable first report, and it is worth anticipating rather than explaining afterwards.

Owned-source citation rate frequently comes back lower than clients expect, because the majority of citations in a generated answer point somewhere other than the brand's own website. Published splits differ materially depending on sample and method, which is why no single percentage appears in this article, but the direction is consistent across every analysis reviewed: third-party sources carry more citation weight than owned ones.

That has a budget implication that a rankings-shaped mental model gets wrong. In classic organic work, your own pages are the asset and off-site effort supports them. In citation work, your own pages are one source among several the system is choosing between, and trade publications, industry associations, standards bodies and reference records are competing sources rather than supporting ones.

So citation share and owned-source citation rate should be read together rather than separately. A rising citation share with a flat owned rate means third-party work is landing and your own pages are not being selected, which is a content-structure problem. A flat citation share with a rising owned rate means you are winning a smaller pool, which is usually a signal that the prompt panel is too narrow. Neither diagnosis is available from a single blended score, which is the argument for the whole framework in miniature.

Coding an answer, in practice

The definitions above assume somebody reads each response and codes it. That step is where reports either become trustworthy or quietly become fiction, and it deserves more attention than it usually gets in a scope of work.

Coding means deciding, for one response, whether the brand was mentioned, whether it was recommended, where it sat in the answer, which claims were made about it, whether those claims match the approved fact base, and whether each citation supports the sentence it accompanies. That is seven judgements on a single run, and the same run coded by two people should produce the same result or the scheme is underspecified.

So the codebook has to be written down, with worked examples of the ambiguous cases. Is a brand listed inside a table but absent from the prose a mention, a recommendation, or neither. Does a comparison that names you unfavourably count in recommendation rate as a zero or as an exclusion. Does a citation to a directory page that lists you alongside forty others count as an owned-source citation, a third-party citation, or a weak signal that deserves its own category. There are defensible answers to all of these. What is not defensible is deciding them differently each month.

A useful discipline is double-coding a sample. Take ten percent of runs, have a second person code them independently, and report the agreement rate alongside the metrics. Where two coders disagree often, the codebook is the problem rather than the coders, and the metric built on it should be treated as soft until the definition is tightened.

None of this is exotic. It is ordinary content analysis, borrowed from research methods that predate the technology by decades. The reason it feels unusual in this category is that most reporting has never been held to it.

The fact base, and who signs it

Factual accuracy is the metric clients underestimate and later care about most. Measuring it requires something that does not exist by default: an approved statement of what is true about the company.

Building it is tedious and quick. Legal entity name. Founding year. Locations. Services actually offered, and services deliberately not offered. Certifications with issuing body and scope. Leadership names and roles. Any figure the company is willing to stand behind publicly, with its source. The document is short, it is signed off by someone with authority, and it is versioned.

Two things then become possible. Coded claims can be checked against a reference rather than against a coder's recollection, which is what makes accuracy a metric rather than an impression. And when an assistant states something wrong, there is a correction workflow with a target: update the owned sources, correct the third-party records the answer was drawing on, and re-measure the specific prompt rather than waiting for a monthly cycle.

This is also where entity consistency stops being an abstract recommendation. Where your own site, your company records, your professional profiles and industry directories disagree about a basic fact, a system reconciling them will pick one, and you do not control which. The fact base makes those disagreements visible as a to-do list.

Two failure patterns worth naming

The first is the improving-panel problem. A programme starts with thirty prompts, adds ten easier ones in month four, and reports the aggregate. Every metric improves. Nothing about the company's visibility changed. This is why additions have to be reported as a separate cohort with their own baseline, and why the panel version belongs on the front page of the report rather than in an appendix.

The second is surface pooling. An assistant that retrieves live from the web and the same assistant answering from trained knowledge are different measurement conditions, and pooling them produces a number that moves for reasons nobody can explain afterwards. Keep them separate in the log and separate in the report, even when it makes the summary less tidy. A programme that cannot say which mechanism produced a result cannot say what to do next.

Both failures share a shape. They make the number easier to read and harder to act on. That trade is almost always the wrong one, because the point of measurement is not to produce a chart. It is to tell you which of the things you did last month is worth doing again.

What a monthly report should look like

Short, and boring in a specific way. Nine metrics against the previous period and against the dated baseline, with the number of eligible and excluded runs stated. The panel version. Any model version changes observed during the window, flagged as a confounder rather than buried.

Then the change log: which pages were modified, which third-party placements landed, and on what dates, so that movement can be associated with actions rather than asserted to follow from them. This section is what separates a report from a dashboard. A dashboard shows you a line. A change log lets you argue about causation with evidence.

And an explicit statement of what did not move, with a hypothesis. Programmes that only report improvements are selecting their own evidence, which is the same failure as swapping the prompts, executed more politely.

Applying this to a vendor, including us

Four questions settle most evaluations. Can I see the prompt panel. What is the baseline date. Are mention, citation and recommendation defined separately, with denominators stated. And can I have the raw responses for any month I choose.

A vendor who answers all four is measurable, which is not the same as good but is a prerequisite for it. A vendor who cannot answer the first is selling a report you have no way to check. A vendor who merges mention and citation has removed your ability to tell which one moved, and therefore which work was responsible.

Publishing these definitions is not an act of generosity. It is the only way the category becomes purchasable, and it is a standard we expect to be held to as readily as anybody else.


Frequently Asked Questions


Why report nine metrics instead of one visibility score?

Because a single score cannot distinguish between cases that require different responses. High visibility attached to wrong facts is a failure, not a result. A brand can be mentioned frequently and recommended never. A citation can be present and irrelevant to the sentence it sits beside. Each of those is invisible inside an aggregate number, and an aggregate you define yourself is difficult to fail.


What does eligible runs mean, and why does it matter?

A run is eligible if the prompt actually invoked the behaviour being measured. Refusals, clarifying questions and errors should be excluded from the denominator, because including them drags percentages down for reasons unrelated to visibility. The rule has to be written before the baseline and applied mechanically. If a vendor cannot state how many runs were excluded and why, the percentages have no fixed denominator and cannot be compared across months.


What is the difference between mention rate and citation rate?

Mention rate is the share of eligible runs in which the brand name appears anywhere in the answer text. Owned-source citation rate is the share of eligible runs containing at least one citation to a domain the client owns. An assistant can name a brand with no link at all, so the two move independently. Merging them into one figure hides which of the two actually changed.


How should a prompt panel be built and governed?

Written from real buying jobs rather than keyword tools, since a buyer asking an assistant to shortlist suppliers phrases it as a task rather than a keyword string. Separated by destination country and buyer role. Versioned, with inclusion rules documented. Then frozen, with a baseline captured before any intervention. After that, prompts are added only by written agreement and reported as a new cohort rather than folded into the existing series.


Why does model version need to be recorded on every run?

Because model updates are the most common cause of a measurement series breaking for reasons unrelated to any work performed. Without a version recorded at run time, a drop looks like a performance failure when it may be a platform change. The same applies to assistant surface, since a web-retrieval mode and a trained-knowledge mode behave differently and should never be pooled.


Is volatility in citation results a sign something is wrong?

No, it is a normal property of the systems and should be reported rather than smoothed. Citation share for a given source type has been observed shifting sharply within weeks as retrieval behaviour changes. The correct response is to repeat prompts in clean sessions and report the spread. A narrow spread is itself a result, because it tells a client whether their visibility is a position they hold or a coincidence they caught.


How long before citation movement is measurable?

It depends which mechanism the answer uses. Where an assistant retrieves live from the web, changes can surface relatively quickly. Where an answer draws on a model's trained knowledge, movement takes considerably longer because material has to be recrawled and reincorporated. Practitioner consensus puts the floor for measurable movement at roughly ninety days, and anything shorter delivers an audit rather than a result. Those two situations should be reported separately rather than averaged into one promised date.


What should a monthly citation report contain?

Nine metrics against the previous period and the dated baseline, with eligible and excluded run counts stated. The panel version. Any model version changes observed in the window, flagged as a confounder. A change log recording which pages were modified and which third-party placements landed, with dates, so movement can be associated with actions. And an explicit statement of what did not move, with a hypothesis, because reporting only improvements is a form of selecting your own evidence.


How do we test whether a vendor is measurable?

Four questions. Can we see the prompt panel. What is the baseline date. Are mention, citation and recommendation defined separately with denominators stated. Can we have the raw responses for any month we choose. A vendor answering all four is measurable, which is a prerequisite for being good rather than proof of it. Google separately advises being wary of third-party tools that promise ranking success or claim to use internal Google metrics, and a retained run log is the direct answer to that concern.

Sources & References:

  • Metric definitions, layer structure, eligibility rules and run-log fields are practitioner standards drawn from the cross-validated research behind Arfadia's Pune pages. They are published in implementable form so that any vendor could apply them identically, and so that Arfadia can be held to them.
  • Google, guidance on optimising for generative AI features in Google Search, published May 2026, advising site owners to be wary of third-party tools that promise ranking success or claim to use internal Google metrics, and stating that meeting all stated requirements does not mean a page will be crawled, indexed or served.
  • Google Search Central, AI features and your website. Eligibility as a supporting link requires that a page be indexed and eligible to be shown with a snippet, with no additional technical requirements and no special schema.org markup for these features.
  • Google, announcement of a Generative AI performance report in Search Console, 3 June 2026, rolling out to a subset of properties. Traffic-side reporting does not substitute for citation measurement, since a citation frequently produces no click.
  • Volatility of citation share by source type within short windows: documented in published citation-index analyses reviewed for this material. No single percentage is stated here, because any such figure describes one prompt set at one moment rather than a general property.
  • Time to measurable citation movement: practitioner consensus reported across independent research passes places the floor at approximately ninety days where trained-knowledge retrieval is involved, on the basis that material must be recrawled and reincorporated. This is practitioner guidance rather than a controlled measurement, and no guaranteed timeline is offered.
  • No standards body prescribing GEO or AEO metrics was located for India, and no survey establishing what Indian buyers currently accept as evidence was located. The definitions here are offered to fill that gap transparently rather than to claim authority.
  • Position of citations within a page: multiple analyses reviewed for this material find citations concentrating in the opening portion of a page, with one large-scale citation study reporting approximately 44 percent of citations drawn from the first 30 percent of page content. Corroborated directionally by independent citation analysis and consistent with Arfadia internal citation data.
  • India brand AI-visibility benchmark: an India-focused analysis for Q1 FY2026-27 examined 2,981 brands across 59,620 prompts between April and June 2026, reporting 19.3 percent meeting its high AI visibility benchmark and an average score of approximately 30 out of 100. Single-vendor benchmark, self-defined scoring method, reported rather than independently audited.
  • Owned versus third-party citation share: published splits differ materially by sample and method and no single figure is stated here. The consistent directional finding across all reviewed analyses is that the majority of citations in a generated answer point to sources other than the brand's own website.
  • No pricing figures, competitor names, ranking guarantees or citation guarantees appear in this article. Model outputs are probabilistic and no outcome is promised.
0 Comments 0 Comments
0 Comments 0 Comments