Google's own coverage documentation lists Tonga among the countries where AI Overviews are available. That is a documented fact from a primary source, and it is the single most commonly misused fact in proposals written for small markets.
Because here is what it does not say. It does not say how often an AI Overview appears when someone searches something about Tonga. It does not say what gets cited when one does appear. It does not say whether the feature behaves the same way for a query typed in Nuku'alofa as for the same query typed in Auckland. Every one of those is a separate question, and none of them has a published answer.
Availability is a fact about a feature. Frequency is a fact about a market. Converting the first into the second is the most common piece of dishonesty in this industry, and it usually happens by implication rather than by outright claim.
How the conversion happens
Nobody writes "AI Overviews trigger 40% of the time in Tonga." That would be checkable and it would be wrong.
What happens instead is a sequence. The proposal states that AI Overviews are available in Tonga, which is true. It then cites a global trigger rate from a study conducted somewhere else, which is also true in its own context. It places the two adjacent, in consecutive sentences, and lets the reader do the arithmetic. Nothing false has been written. A false impression has been created.
Watch for the same pattern with adoption figures. During research for this project we found three separate statistics about AI use in travel planning, all real, all from named sources, and all measuring different things: roughly 30% of United States travellers using AI extensively for trip planning in 2025, up from 13% the previous year; nearly 40% of United States travellers using generative AI for travel research; and 56% of United States travellers using AI to plan, book or accompany at least one trip. Those are three different behaviours with three different thresholds. Presented as a range, they suggest a measurement. They are not a measurement, they are three studies asking three questions, and three of the four research passes for this project flagged the merging risk independently.
None of them measured Tonga. All of them measured the United States.
What Is Documented, and What Is Simply Absent
Everything on the left can be checked against a primary source today. Everything on the right has not been published by anyone, and no amount of confident phrasing changes that.
Documented
Tonga is listed in Google's own AI Overviews coverage documentation
Travel carries the highest citation share of nine tracked sectors, 22.6% against a 6.8% average, on United States desktop data for May 2026
Clicks on traditional results occurred on 8% of visits containing an AI summary against 15% without, in an observed panel of 900 United States adults
No provider is verifiably headquartered in Tonga offering this work
Absent from every source reviewed
How often AI Overviews trigger for Tonga queries
What gets cited when they do
Whether behaviour differs between a query typed in Tonga and the same query typed in Auckland
Any Tonga-specific citation rate, share of voice or benchmark of any kind
Search volume for lea faka-Tonga queries
Sources: Google Search Help coverage documentation • Similarweb 2026 generative AI benchmark report, United States desktop, May 2026, modelled estimates • Pew Research Center, observed panel of 900 United States adults, March 2025 • Cross-validation across four independent research passes, 2026
Created by Arfadia • arfadia.com/blog
The one strong sector statistic, and its three conditions
There is a genuinely useful figure here, and it deserves to be used properly rather than stripped for parts.
Similarweb's 2026 generative AI benchmark report found that travel carries the highest citation share of the nine sectors it tracked: 22.6% of ChatGPT answers in the travel category contained at least one citation, against a cross-sector average of 6.8%. Retail came second at 13.5%, then sports at 10.7%, finance at 8.0%, technology at 6.6%, health at 6.3%, media and entertainment at 6.0% and education at 4.8%. The overall average had risen from somewhere around 1.3% to 1.6% in June 2025.
Three conditions travel with that number, always.
The geography is the United States, on desktop, for May 2026. Not global. Not New Zealand. Certainly not Tonga.
These are modelled estimates, not measured platform data. Similarweb's own disclaimer describes its data as estimates and extrapolations from third-party sources, without warranty of accuracy. Similarweb also sells AI visibility tracking products. That does not make the figure wrong, and the report is one of the better public datasets available. It does mean the correct confidence marker is reported rather than verified.
It measures whether any citation appeared, not who won it. This is the condition most often dropped and it is the most consequential. A 22.6% citation share in travel tells you the citation layer in that category is comparatively wide. It tells you nothing whatsoever about whether a small operator can get into it. Those are different claims and only the first one has evidence behind it.
Used with all three conditions attached, the figure supports a real argument: travel is a category where assistants cite sources more often than they do elsewhere, so publishing citable material has a better prior in travel than in most industries. That is a reasonable basis for doing the work. It is not a forecast.
Why your KPI might be measuring someone else's release notes
Here is the finding that should decide how a Tonga engagement is reported.
Following a ChatGPT update on 7 May 2026, the share of AI referral traffic landing on homepages rather than deeper pages rose from roughly a quarter to nearly sixty percent within about three weeks. No brand did anything. No content changed. A platform shipped an update and a widely tracked metric moved by more than thirty points.
Sit with what that means for reporting. If your success measure is built on referral behaviour, then a substantial share of your reported movement is platform behaviour, and you cannot separate the two after the fact. A client looking at a chart showing homepage referrals doubling has no way to know whether that was your work or a release on 7 May.
The same report noted that around 65% of cited URLs sat two or three folders deep, while 58.8% of AI referral traffic landed on homepages. Cited pages and landing pages are not the same set. Optimising the second while reporting it as evidence about the first is a category error, and it is easy to make accidentally.
The response is not to give up on measurement. It is to measure something that platform updates cannot silently rewrite.
A fixed prompt panel, and why we state the denominator
The method is unglamorous and it works.
Write a fixed set of prompts that a real customer would actually type, in the language and phrasing they would use. Run them across the assistants that matter, on a stated schedule, from a stated location. Record whether the business appears in the citation set, as a count out of the panel size. Re-run the identical panel at the same interval. Report the change in counts.
That last point about the denominator matters more than it sounds. There is no industry standard setting a minimum panel size for this kind of reporting, and we checked. One research pass proposed a fixed benchmark bank of a hundred prompts and presented it as a standard, which it is not, and no source could corroborate it. Another suggested twenty to thirty. A third stated plainly that no authoritative standard exists.
So we state our panel size rather than implying a norm. Twelve out of forty prompts is a reportable finding. Thirty percent, without the forty, is a number that could mean anything, and on a small market it usually means very little.
| Metric | Stable against platform updates? | Use it as |
|---|---|---|
| Citations out of a fixed prompt panel, stated size | Largely yes, if the panel and method are unchanged | Primary measure |
| Per-platform citation counts, reported separately | Yes, and it exposes divergence a blended score would hide | Primary measure |
| Factual accuracy of what the assistant says about you | Yes, and it is often the fastest thing to improve | Primary measure |
| Enquiries, booking requests, booking value | Yes. It is your own ledger | Outcome measure |
| AI referral traffic volume | No. Moved by more than thirty points in three weeks on a single platform update | Context only, never a target |
| Blended cross-platform visibility score | No, and it conceals the gap you are paying to find | Avoid |
| Percentage change without a denominator | Meaningless at low volumes regardless of platform behaviour | Never |
The percentage problem, stated once and properly
Two citations becoming four is a 100% increase. It is also two citations.
On a market of 104,000 people with roughly 60,700 internet users, low absolute numbers are the normal condition rather than a sign of failure. Percentage reporting on those bases produces charts that look dramatic and mean almost nothing, and clients work this out eventually. When they do, everything else in the report becomes suspect too.
Every independent research pass conducted for this project reached this conclusion separately, across both the search and the answer-engine tracks. That kind of unanimity is unusual enough to be worth acting on. Report counts. Where a percentage genuinely helps, publish the numerator, the denominator and the measurement window beside it, every single time.
Building a Baseline Where No Benchmark Exists
No industry standard governs any of this. So the standard has to be stated, held constant, and disclosed in every report.
Write prompts a real customer would type
In their phrasing, from their market, at their point of decision. Not brand-name lookups, which flatter the report and predict nothing.
Fix the panel size and disclose it
No authoritative standard sets a minimum. One research source proposed a hundred as a standard, which it is not; another suggested twenty to thirty; a third confirmed no standard exists. State yours instead of implying a norm.
Record the location and date of every run
Assistant behaviour varies by location. A baseline gathered from one country and a re-test from another is two different measurements wearing the same label.
Report each platform separately, never blended
A business can appear in one assistant and be absent from another on the identical prompt. A single blended score hides precisely the gap the engagement is meant to find.
Track accuracy as well as presence
Being cited with wrong opening hours or a wrong season date is worse than absence. Accuracy is frequently the fastest metric to improve and the most commercially useful.
Re-run identically, and change one thing at a time
If the panel changes between runs, the comparison is void. Platform updates will still move things underneath you, which is why the panel stays fixed and the notes record what shipped when.
Sources: Sampling standard absence confirmed across independent research passes for this project • Platform instability from Similarweb 2026 generative AI benchmark reporting, 7 May 2026 ChatGPT update • Method is Arfadia practice, documented since 2023
Created by Arfadia • arfadia.com/blog
What a first engagement honestly looks like
Measurement first, recommendations second. In that order, because reversing it means guessing.
The opening deliverable is a baseline: a fixed prompt panel, run across the assistants, from the origin markets that matter, with the result recorded as counts and the accuracy of each mention noted. That produces a document showing where the business currently sits, which nobody has ever produced for it before, because no such baseline exists for any Tongan business.
What comes next depends on what the baseline shows, and the honest possibilities include outcomes that reduce the scope of work. If the business is already accurately represented in most of the panel, the remaining work is maintenance rather than a programme. If it is absent everywhere and the reason is that the assistants have no source to draw on, the answer is publication rather than optimisation. If it appears but with wrong details, the fix is entity correction, which is fast and cheap.
None of those conclusions can be reached in advance. That is the point of measuring first, and it is why we will not write a strategy document for a Tongan business before we have run the panel.
Tessar Napitupulu sets out the measurement discipline behind answer-engine work, including why absolute counts beat rates on small bases, in Cited or Silent.
Frequently Asked Questions
Are AI Overviews available in Tonga?
Yes. Google's own coverage documentation lists Tonga among the countries where the feature is available, which is a documented fact from a primary source. That is a statement about availability only. How often it triggers for Tonga queries, and what it cites when it does, is absent from every source reviewed for this article.
What is the difference between availability and frequency?
Availability is a fact about a feature being switched on in a country. Frequency is a fact about how often it actually appears for real queries in that market. Proposals commonly state the first, cite a global trigger rate measured somewhere else in the next sentence, and let the reader infer the second. Nothing false is written and a false impression is created.
Is the 22.6% travel citation figure reliable?
It is one of the better public datasets available, with three conditions that must travel with it. The geography is United States desktop data for May 2026, not global and not Tonga. The figures are modelled estimates extrapolated from third-party sources, published by a vendor that also sells AI visibility tracking, so the correct marker is reported rather than verified. And it measures whether any citation appeared in a category, not whether a particular business can win one.
Why not report AI referral traffic?
Because it moves on platform updates you do not control. Following a ChatGPT update on 7 May 2026, the share of AI referral traffic landing on homepages rather than deeper pages rose from roughly a quarter to nearly sixty percent within about three weeks, with no action taken by any brand. Reporting built on referral behaviour cannot separate your work from a release note after the fact.
How many prompts should a tracking panel contain?
No authoritative standard exists, and we checked. One research source proposed a fixed bank of a hundred prompts and presented it as a standard, which no source could corroborate. Another suggested twenty to thirty. A third confirmed plainly that no industry minimum has been established. The workable answer is to state your panel size in every report rather than implying a norm, so that twelve out of forty is reportable and thirty percent on its own is not.
Why report counts instead of percentages?
Because two citations becoming four is a 100% increase and also two citations. On a market with roughly 60,700 internet users, low absolute numbers are the normal condition rather than a sign of failure, and percentage reporting on those bases produces charts that look dramatic and mean very little. Every independent research pass conducted for this project reached the same conclusion separately across both tracks.
Should visibility across platforms be combined into one score?
No. A business can appear in one assistant and be entirely absent from another on the identical prompt, and a blended score conceals exactly the gap the engagement is meant to find. Report each platform separately, with its own count and its own denominator.
What does a first engagement produce?
A baseline before any recommendations: a fixed prompt panel run across the assistants from the relevant origin markets, with results recorded as counts and the factual accuracy of each mention noted. No such baseline currently exists for any Tongan business. What follows depends on what it shows, and the honest possibilities include findings that reduce the scope of work, such as the business already being accurately represented and needing maintenance rather than a programme.
Sources & References:
- AI Overviews availability: Google Search Help coverage documentation listing Tonga among supported countries. Verified August 2026. Trigger frequency and citation composition for Tonga queries are absent from all sources reviewed for this article.
- Sector citation rates: Similarweb, 2026 generative AI benchmark report, published July 2026. Travel 22.6%, retail 13.5%, sports 10.7%, finance 8.0%, technology 6.6%, health 6.3%, media and entertainment 6.0%, education 4.8%, against a nine-sector average of 6.8%. United States desktop data for May 2026. The cross-sector average had risen from approximately 1.3% to 1.6% in June 2025. Similarweb's own disclaimer describes its data as estimates and extrapolations from third-party sources without warranty of accuracy; Similarweb also sells AI visibility tracking products. Marked REPORTED, not VERIFIED. The measure captures whether any citation appears in a category, not which businesses receive them.
- Platform instability: following a ChatGPT update on 7 May 2026, the share of AI referral traffic landing on homepages rather than deeper pages rose from approximately 25% to nearly 60% within about three weeks with no brand action. The same reporting notes approximately 65% of cited URLs sitting two or three folders deep while 58.8% of AI referral traffic landed on homepages.
- Click behaviour: Pew Research Center, observed panel of 900 United States adults, data collected March 2025. Clicks on traditional search results occurred on 8% of visits containing an AI summary against 15% of visits without one. Links within summaries were clicked on 1% of visits. Wider ranges of 47% to 61% circulating in industry commentary derive from secondary aggregators and are not used here.
- Travel AI adoption figures, all United States and all measuring different behaviours: approximately 30% using AI extensively for trip planning in 2025, up from 13% the prior year, per Skift US Travel Tracker via McKinsey; nearly 40% using generative AI for travel research, per Phocuswright November 2025; 56% using AI to plan, book or accompany at least one trip, per Phocuswright March 2026. These measure different thresholds and are not a range. Three of four independent research passes flagged the merging risk separately.
- Sampling standards: no authoritative standard setting a minimum prompt panel size for answer-engine reporting was identified. One research source proposed a fixed hundred-prompt benchmark bank and presented it as a standard, which no other source corroborated; another recommended twenty to thirty; a third confirmed no industry standard exists.
- Absolute-count reporting on small bases was independently recommended by every research pass conducted for this project, across both the search and answer-engine tracks.
- Market context: Tonga population near 104,000 with approximately 60,700 internet users, per Kepios and DataReportal, Digital 2026: Tonga. No provider verifiably headquartered in Tonga offering answer-engine or generative-engine optimisation was identified.
- The measurement method described here is Arfadia practice, documented since 2023, rather than an industry standard. It is presented as a stated method precisely because no standard exists to appeal to.