A page can hold position one in Google and never once be named inside an AI answer. A page sitting on the third page of results can be cited repeatedly. Both are ordinary. Neither is a paradox.
Ranking and citation are separate observations produced by separate systems. Almost every monthly report in circulation blurs them anyway, usually into a single line called AI visibility.
That blur is not always dishonest. Often it reflects genuine confusion, or a reporting tool that offers one combined number because a combined number fits on a slide. Either way, the person reading the report loses the ability to tell which thing actually moved, which is the only reason to read a report in the first place.
What follows is a practical guide to reading one. Written for the Philippine market specifically, because the evidence conventions here have a particular shape: the working proof standard is a screenshot, and no provider in this market or in Singapore was found publishing a measurement method a client could re-run.
The four observations that keep getting merged
Start by separating what is being counted. Four distinct things, and a report that does not distinguish them is not a report you can act on.
Rank is a position in a list of links, for a given query, device, location and moment. Old, well understood, reasonably stable to measure.
Presence is whether an AI Overview or generated answer appeared at all for that query. A feature observation about the results page. Says nothing about who was in it.
Mention is whether the generated answer named your brand. Genuinely valuable, particularly in a shortlist context, and frequently the outcome that moves first.
Citation is whether the answer attributed information to your source, usually with a link. This is the outcome most GEO work is actually aimed at, and it is the hardest of the four to fake because it leaves a URL behind.
You can move mention substantially without moving citation. The reverse happens too. A report folding all four into one score cannot tell you which, which means it cannot tell you whether the work is working.
What Your Report Should Be Counting Separately
Any two of these can move in opposite directions in the same month. A blended figure hides which.
Rank
Position in a list of links, for a specified query, device, location and moment. Well understood and reasonably stable to measure. Not evidence about AI answers in either direction.
Presence
Whether a generated answer appeared for that query at all. A feature observation about the results page. Tells you the surface exists for that query, nothing about who occupied it.
Mention
Whether the generated answer named your brand. Real value, especially in shortlist questions, and usually the first outcome to move. Being named is not being sourced.
Citation
Whether the answer attributed information to your source, usually with a link. Hardest of the four to fake, because it leaves a URL behind that can be checked afterwards.
The tell to watch for
A single metric labelled AI visibility, AI readiness or a similar composite, presented without the arithmetic that produced it. Ask what goes into it and with what weights. If the answer is vague you are being shown a rating rather than a measurement, and a rating cannot be re-run by you or by anyone auditing the engagement later.
Sources: Cross-validation of eight independent AI research passes on Philippine SEO and GEO, August 2026 • Practitioner consensus on ranking-versus-citation separation across four research models
Created by Arfadia • arfadia.com/blog
Ten questions that separate measurement from marketing
Read them as procurement questions. None requires technical knowledge to ask, and the quality of the answers tells you most of what you need to know.
What exactly was measured, and what is the definition? Not the metric name. The definition. Share of voice calculated how, against which competitor set, across which prompts.
What is the query or prompt list, and when was it frozen? A list that changes between reporting periods produces a comparison that means nothing. Ask for the list and the change log.
Which engine produced each figure? Engines diverge sharply in which sources they prefer. Blending them into one number is the single most common way an unfavourable result gets averaged away.
How many times was each prompt run? Generative answers vary between identical prompts. One run per prompt is a sample of one and cannot distinguish a stable result from variance.
What was the location and session state? A logged-in account with months of history is partly measuring its own history. Philippine IP and clean sessions, or the numbers describe the tester rather than the market.
What is the model and date for each observation? Results shift after model updates. Without a date stamp you cannot tell a real decline from a platform change.
Can I see the full answer text, not a crop? A cropped screenshot removes the context that would let you judge whether the mention was favourable, hedged, or third behind two competitors.
How many runs were invalid, and why? Refusals, timeouts, empty results. A report with no failure count has either been extraordinarily lucky or has quietly discarded the inconvenient observations.
What is the baseline, and when was it taken? Taken before work started, or reconstructed afterwards. Reconstructed baselines drift toward flattering starting points, generally without anyone intending it.
Can I reproduce any single figure myself? This is the question that settles it. If the answer involves a proprietary score nobody outside the agency can recompute, you are being asked to trust rather than verify.
Reading the common report patterns
| What the report says | What it may actually mean | The question to ask |
|---|---|---|
| "AI visibility up 40 percent" | A composite score moved. Possibly mention only, possibly on one engine, possibly on an expanded prompt list | What is in the composite, with what weights, and did the prompt list change? |
| "Now cited by ChatGPT" plus a screenshot | One favourable answer at one moment, possibly not reproducible an hour later | How many runs of that prompt, over what period, and how many returned the citation? |
| "Ranking first for our main keyword" | A genuine SEO outcome, and no information at all about AI answers | What is the citation rate for that same query, reported separately? |
| "Share of voice 35 percent" | Meaningless without the competitor set and prompt list behind it | Share of what, against whom, across which prompts, on which engine? |
| "AI readiness score: 82 out of 100" | A vendor-defined rating, usually of your site rather than of any AI behaviour | Does this measure anything an AI engine did, or only what our site contains? |
| "AI Overviews appear on 40 percent of your keywords" | Possibly a real observation on your panel, possibly a global average borrowed and relabelled | Was this measured on our query list from Philippine sessions, or taken from a published study? |
| "Traffic from AI is growing rapidly" | Often true and often tiny. Percentage growth on a base of forty sessions is not a channel | What is the absolute session count, not the growth rate? |
| "Impressions are up strongly" | Possibly true and possibly worthless, if clicks did not follow | What happened to click-through rate over the same period, at the same average position? |
The AI-traffic row deserves emphasis, because it is the most seductive one. AI-referred traffic frequently shows enormous percentage growth from a base near zero, and most AI answers carry no click at all, which means low referred traffic is the expected condition rather than a failure. Track it. Do not weight it heavily. Report the absolute number.
The one Philippine dataset, and how to read it properly
Most of this article is about scepticism. This section is the opposite, because there is one Philippine study worth knowing and it gets overlooked in favour of borrowed global figures.
LeapOut Digital published "Ranked But Bypassed" in June 2026. GA4 and Search Console data for 11 brands, 7 of them Philippine, tracked from January 2023 through to June 2026. Not a survey. Not a vendor case study. Observational analysis of real analytics from real properties.
The headline finding: across roughly 59 million search impressions over twelve months, fewer than 1.3 million visits resulted. Somewhere around 97 to 98 percent of people who saw these brands in Google did not click through to them.
Position-level detail is sharper. Philippine brands sitting at average positions 8.6 and 9.1 recorded click-through rates of 1.9 percent and 2.3 percent, against a pre-AI benchmark of roughly 6 to 9 percent for those positions. Compression of well over half, on Philippine properties, measured rather than modelled.
Here is what makes it worth reading rather than just quoting. The authors mark their own confidence levels. The impression-to-visit finding is self-rated high confidence, and they state explicitly that the decline cannot be attributed to AI alone. Their claim that Philippine brands trail Western markets by 18 to 24 months is rated by them as a hypothesis, not a finding.
Anyone citing that 18 to 24 month line as established fact has misread the study they are citing. That happens a lot, and it is a useful test of whether the person quoting a study has actually opened it.
One more finding from the same work, and it is the most practically useful. The authors describe what they call the Branded Content Principle: brand-name-plus-task queries, of the "[brand] pricing" or "[brand] setup guide" variety, are far more defensible against AI interception than generic category queries. Generic category queries, in their phrasing, belong to whoever the AI decides to cite. That is a content strategy implication hiding inside a measurement study, and it argues for building depth on questions only your brand can genuinely own.
The measurement problem hiding in your own analytics
Two findings from that same dataset are really about instrumentation rather than performance, and both are worth checking on your own property this week.
First, unassigned traffic grew twenty to thirty times from 2023 baselines for several of the Philippine brands tracked. Twenty to thirty times. When traffic that large moves into a bucket labelled unassigned, every channel report built on top of it becomes less reliable, and the drop you are attributing to one channel may be a classification failure rather than a real decline.
Second, the GA4 "AI Assistant" channel appeared for five Philippine brands for the first time in the first half of 2026, at volumes between one and 89 users per brand. Tiny numbers. The signal is not the volume, it is that the classification now exists and has started firing.
Which produces a practical instruction. Take a baseline of your unassigned traffic and your AI Assistant channel now, while the numbers are small enough to be uninteresting. In twelve months, when somebody asks how much your AI-referred traffic has grown, you will either have a starting point or you will have an argument.
Why this market makes the problem worse
Two Philippine-specific conditions raise the stakes on everything above.
The first is that the local evidence norm is a screenshot. Across the research reviewed for this article, the marketable proof offered by providers in this market consists overwhelmingly of before-and-after prompt screenshots and self-assessed readiness scores. There is no established local convention of independent verification, and no independently audited citation-gain result was located for any provider operating here or in Singapore. A buyer asking for reproducible method is asking for something the market does not routinely supply, which makes asking unusual enough to be informative all by itself.
The second is that the most important benchmark barely exists. Whether AI engines cite Philippine sources or default to US and global English ones for Philippine queries has never been measured. On AI Overview prevalence there is exactly one Philippine estimate, Ahrefs placing the Philippines at 29.1 percent of keywords across 108 million AI Overview queries, and other research holds that no Philippine prevalence figure should be presented as Philippine-specific at all. So nobody can tell you what a good citation rate looks like here, because there is no distribution to compare against.
Anyone presenting a target citation rate for the Philippines is presenting an invented number, however reasonable it sounds coming out of their mouth.
The practical consequence: in the absence of external benchmarks, your own baseline is the only reference point that exists. Which makes taking one before any work starts considerably more important here than in markets where published norms are available.
The Minimum Evidence Pack for a Monthly Report
None of this is exotic. It is what makes a number checkable rather than merely believable.
Frozen registry plus change log
The query and prompt list as locked before work began, with every subsequent change dated and explained. Without it, period-on-period comparison is arithmetic on shifting ground.
Full answer capture
The complete generated response, not a crop. The surrounding text is what tells you whether the brand was recommended, hedged, or listed third behind two competitors.
Engine, model, date, location, session
Five attributes on every observation. Together they let you separate a genuine change in your position from a platform update or a testing artefact.
Invalid runs, with reasons
Refusals, timeouts, empty results, rate limits. A report with no failure count has either been very lucky or has dropped the inconvenient observations without saying so.
Unassigned traffic, tracked deliberately
Unassigned traffic grew twenty to thirty times from 2023 baselines for several Philippine brands in the one local dataset available. A channel report sitting on top of that is less reliable than it looks, so the unassigned line belongs in the report rather than hidden behind it.
And rankings in a separate appendix
Ranking data is useful and should absolutely be reported. It belongs in its own clearly labelled section, never merged into a citation figure and never used as a proxy for one. If a report cannot show you rank and citation as two independent series, it cannot tell you which of them the work moved, and that is the entire purpose of reading it.
Sources: Recommended reporting standards synthesised from four independent GEO research passes, August 2026 • LeapOut Digital, "Ranked But Bypassed", June 2026, for the unassigned-traffic finding • Observed Philippine market evidence conventions across reviewed provider material
Created by Arfadia • arfadia.com/blog
What a good answer sounds like
When you ask the ten questions, you are not looking for perfect numbers. You are listening for whether the person can describe their own method without reaching for adjectives.
A good answer names the definition, admits where the sample is thin, distinguishes what was measured from what was inferred, and says plainly when something is not known. A weak answer restates the metric name more confidently, or explains that the methodology is proprietary.
Proprietary is not automatically a red flag. A proprietary tool producing a figure you can independently spot-check is fine. A proprietary score you cannot recompute or verify against anything is a rating, and ratings are opinions with numbers attached.
The other good sign, and it is easy to miss: someone who volunteers a caveat you did not ask for. LeapOut's authors marking their own 18 to 24 month claim as a hypothesis is exactly that behaviour. A vendor who does the same with their own results is telling you something about how they will report a bad month.
This is the standard our SEO service for the Philippines and GEO service for the Philippines are both built to meet, which is why every figure we hand over carries its own definition, geography, source and date. Tessar Napitupulu sets out the underlying measurement discipline in Cited or Silent.
Frequently Asked Questions
What is the difference between a mention and a citation?
A mention means the generated answer named your brand. A citation means the answer attributed information to your source, usually with a link you can follow. Both matter and they move independently: a programme can raise mention rate substantially without raising citation rate, or the reverse. They belong in separate columns of any report, because merging them hides which one changed.
Can a page rank first and never be cited by an AI answer?
Yes, and the reverse is equally common. Ranking systems and generative answer retrieval use different signals, so a first-position page can go uncited while a page far down the results is cited repeatedly. The only valid test is inspecting the generated answers and the URLs they cite. Inferring citation from organic position is not a shortcut, it is a different measurement.
What should I ask when a report shows AI visibility up 40 percent?
Three things. What goes into the composite and with what weights. Whether the query or prompt list changed between periods. And which engine or engines produced the figure. Composite scores are where unfavourable results get averaged away, usually without anyone intending it, and none of the three questions requires technical knowledge to ask.
Is a screenshot of an AI answer acceptable as proof?
Not on its own. A screenshot captures one answer at one moment with no record of the exact query, no repeat runs, no engine or model date, and no way to distinguish a stable result from ordinary answer variance. Generative outputs differ between identical prompts, so a single favourable capture cannot establish anything. Ask how many runs of that prompt were performed and how many returned the same result.
Has anyone actually measured zero-click behaviour in the Philippines?
Once. LeapOut Digital's "Ranked But Bypassed", published June 2026, tracked GA4 and Search Console data for 11 brands, 7 of them Philippine, from January 2023 to June 2026. Roughly 59 million impressions produced fewer than 1.3 million visits over twelve months, so around 97 to 98 percent did not click. Philippine brands at average positions 8.6 and 9.1 recorded click-through of 1.9 percent and 2.3 percent against a 6 to 9 percent pre-AI benchmark. The authors state the decline cannot be attributed to AI alone, and they rate their own claim that Philippine brands trail Western markets by 18 to 24 months as a hypothesis rather than a finding.
What is the Branded Content Principle?
A finding from the same LeapOut study. Brand-name-plus-task queries, of the "[brand] pricing" or "[brand] setup guide" variety, are far more defensible against AI interception than generic category queries, which in the authors' phrasing belong to whoever the AI decides to cite. The practical implication is to build depth on questions only your brand can genuinely own, rather than competing for generic category answers where the engine picks the source.
Why is our unassigned traffic growing, and does it matter?
It matters a great deal for reading any report. In the LeapOut dataset, unassigned traffic grew twenty to thirty times from 2023 baselines for several Philippine brands. When traffic that large moves into an unclassified bucket, every channel report built on top of it becomes less reliable, and an apparent decline in one channel may be a classification failure rather than a real loss. Track the unassigned line explicitly rather than letting it sit behind the channel summary.
What is a good AI citation rate for a Philippine business?
Nobody can tell you, and that is the honest answer. Whether AI engines cite Philippine sources or default to global ones for Philippine queries has never been measured, and on prevalence there is one disputed Philippine estimate rather than a distribution. Any specific target citation rate offered for this market is invented. Your own baseline, taken before work begins, is the only reference point that exists.
Should AI-referred traffic be a headline metric?
No. Track it, but do not weight it heavily. Most AI answers carry no click at all, so low referred traffic is the expected condition rather than evidence of failure. In the one Philippine dataset available, the GA4 AI Assistant channel appeared for five brands at between one and 89 users each. Percentage growth on a base that small looks dramatic and means very little, so report the absolute session count alongside any growth figure.
Is a proprietary measurement methodology a warning sign?
Not by itself. A proprietary tool producing figures you can independently spot-check is perfectly reasonable. The warning sign is a proprietary score that cannot be recomputed or verified against anything outside the agency. That is a rating rather than a measurement, and a rating cannot be re-run by you or by anyone auditing the engagement later.
Sources & References:
- Ranking and citation are separate observations produced by separate systems. This separation was stated independently in all four GEO research passes conducted for the Philippine cycle in August 2026, including the observation that a top-ranking page can go uncited while a lower-ranking page is cited repeatedly.
- Philippine zero-click measurement: LeapOut Digital, "Ranked But Bypassed", Marvin Ortiz, published 18 June 2026. GA4 and Search Console data for 11 brands, 7 of them Philippine, January 2023 to 15 June 2026. Approximately 59 million search impressions over twelve months producing fewer than 1.3 million visits, roughly 97 to 98 percent not clicking through, self-rated High Confidence with the authors' explicit caveat that this cannot be attributed to AI alone. Philippine brands at average positions 8.6 and 9.1 recording click-through of 1.9 percent and 2.3 percent against a 6 to 9 percent pre-AI benchmark. Unassigned traffic growing 20 to 30 times from 2023 baselines for several Philippine brands. The GA4 AI Assistant channel appearing for 5 Philippine brands for the first time in H1 2026 at 1 to 89 users per brand. The study's claim that Philippine brands trail Western markets by 18 to 24 months is rated by its own authors as a hypothesis, not a finding. The Branded Content Principle, holding that brand-name-plus-task queries are more defensible against AI interception than generic category queries, is also drawn from this study.
- Philippine market evidence conventions: across reviewed provider material, the predominant proof offered to buyers consists of before-and-after prompt screenshots and self-assessed readiness scores. No established local convention of independent verification was identified.
- No independently audited citation-gain result was located for any provider operating in the Philippines or Singapore across eight independent research passes. All performance claims reviewed were self-published.
- Whether AI engines cite Philippine sources rather than US or global English sources for Philippine queries is UNAVAILABLE, which means no external distribution exists against which a Philippine citation rate could be benchmarked.
- Philippine AI Overview prevalence: Ahrefs analysis of its Brand Radar database of 108 million AI Overview queries places the Philippines at 29.1 percent of keywords, level with Mexico. REPORTED, single dataset, method-dependent, and other research reviewed for this cycle states that no Philippine prevalence figure should be presented as Philippine-specific. Both positions recorded rather than reconciled.
- Generative answers vary between identical prompts, so single-run observations cannot distinguish a stable result from answer variance. Repeat runs with logged engine and model version are the minimum control for this.
- Most AI answers carry no click, so AI-referred traffic is expected to be small in absolute terms. Percentage growth figures on very small bases should be read alongside the absolute session count.