Measuring GEO When One Mandate Is Worth Millions
SEO

Measuring GEO When One Mandate Is Worth Millions

Four tiers of visibility measurement, and the one calculation we refuse to run because the causal path behind it cannot be observed.

A reinsurance platform in Hamilton might sign three new cedants in a good year. A captive manager might onboard a dozen. A fund administrator competing for institutional mandates might win two that matter and lose four that also mattered.

Now try to report on generative engine optimisation against that. The usual measurement apparatus, sessions, conversion rate, cost per acquisition, return on ad spend, was built for markets where outcomes arrive weekly and in volume. Here they arrive quarterly, in ones, after months of relationship-building that happens nowhere a tracking script can see.

This is the honest measurement problem in Bermuda financial services, and it is not solved by finding a cleverer attribution model. It is solved by deciding, in advance, which things can be measured and which cannot, and then refusing to blur the line between them. What follows is how we do that, including the specific numbers we will not report and why.

Citation and ranking are not the same measurement

Start with the distinction everything else rests on.

A ranking is a position in a list of links. A citation is a model naming you as a source inside an answer it generated, often with no click at all. These are produced by different mechanisms, and they move independently of each other.

In most sectors they at least correlate loosely. In finance they barely do. One tracked analysis of AI answer citations found that only around 11 per cent of citations in finance also appeared in the organic top ten, the lowest overlap of any industry it measured, and that roughly two thirds of finance citations came from pages sitting outside the top hundred entirely. Across regulated categories generally, health, finance and legal together, the overlap ran around 17 per cent.

That figure comes from a single vendor analysis, it is global rather than Bermuda-specific, and we report it as such. But even discounted heavily, the implication holds. A rank-tracking dashboard showing steady positions can sit alongside complete invisibility in generated answers, and neither number would look wrong.

Which means: if you are buying GEO and receiving ranking reports, you are not receiving a measurement of what you bought.

Four tiers, one of them prohibited

What Can Be Measured, and What Must Not Be Claimed

Every tier below is reportable except the last. The last one is where most GEO reporting in regulated finance quietly goes wrong.

Tier 1, directly observable

Citation presence per prompt, per engine, per date

Verbatim answer text captured and archived

Factual accuracy of statements made about your entity

Non-brand impressions segmented by searcher country

Tier 2, comparative within a fixed panel

Share of visible answers against named alternatives

Movement over time within a versioned, frozen panel

Cross-engine coverage breadth

Tier 3, reportable with method stated

Qualified enquiries with source captured at point of contact

Self-reported discovery channel from enquiry forms

Pipeline influence, reported separately and never multiplied

Tier 4, do not report

Return on investment calculated from citation counts

Revenue attributed to an AI answer appearing

Mandate value multiplied by visibility percentage

Any forecast of citations or placements won

Tier 4 items are excluded because no measurable causal path exists between an AI citation and an institutional mandate. They are omitted as a matter of method, not caution.

The four tiers, and why the fourth one is empty

Tier one is what a model actually did. Citation presence for a named prompt on a named engine on a named date, with the verbatim output captured. Whether the statements made about your entity were factually correct. Non-brand impressions from Search Console segmented by searcher country. All of this is observable, and all of it can be archived so a client can audit the claim months later.

Tier two is comparative, and it is only valid inside a fixed panel. Share of visible answers against named alternatives, movement over time, breadth of coverage across engines. The measure is legitimate. Its integrity depends entirely on the panel being frozen and versioned before measurement begins, which brings us to the most common quiet manipulation in this field.

Share of voice has a denominator you chose. If the panel is assembled after you know which prompts you perform well on, the number goes up without anything improving. The control is boring and effective: freeze the panel, version it, deliberately include prompts where you currently do badly, and publish the panel composition alongside the figure. A share of voice number reported without its panel is not a metric, it is a claim.

Tier three is reportable with the method stated in the same breath. Qualified enquiries with the discovery channel captured at the point of contact. Self-reported attribution, which is imperfect and still the best available. Pipeline influence, meaning the observation that a visibility programme was running while a mandate progressed, reported as exactly that and nothing more.

Tier four is where we stop. No return on investment calculated from citation counts. No revenue attributed to an answer appearing. No mandate value multiplied by a visibility percentage. No forecast of citations or placements.

The reason is not caution. It is that the causal path does not exist in measurable form. A risk manager reads a generated answer in March, mentions the domicile to a colleague in April, meets your team at a conference in June, and signs in October after three calls that left no digital trace. Which of those is the citation worth? There is no defensible answer, and multiplying anyway produces a figure that looks rigorous precisely because it has decimal places.

In a market where your report is read by a compliance function, that failure mode is worse than useless. It is a document with your name on it making a claim that cannot be supported.

The surface moves faster than the reporting cycle

There is a second constraint, and it changes what a GEO engagement should even consist of.

Analysis of roughly 600,000 citation events during early 2026 found only around 11 per cent overlap between the domains ChatGPT and Perplexity cite for comparable queries. It also found approximately half of cited domains changing from one month to the next.

Read those together. Measuring one engine measures roughly a tenth of the surface. And whatever you measured has half-rotated by the following month.

Which means a one-off GEO audit, however thorough, describes a state that no longer holds by the time it is presented. Not because the work was poor. Because the thing being measured moves at that speed.

Practitioner sources generally suggest sixty to ninety days minimum before changes to your material affect what models retrieve, since sources have to be recrawled and reprocessed. Combine that with monthly churn and the conclusion is unavoidable: monitoring is the deliverable, not a report. Anything under about ninety days is an audit, which is a legitimate purchase, but it should be named honestly rather than sold as a programme.

Why a one-off audit expires

The Answer Surface Does Not Hold Still

Three findings from an analysis of roughly 600,000 citation events during January and February 2026. Together they determine what a GEO deliverable has to be.

~11%

Overlap between domains ChatGPT and Perplexity cite for comparable queries

~50%

Of cited domains changing from one month to the next

60-90

Days typically needed before changes to source material affect retrieval

What follows, operationally

Monitoring is the deliverable. A single audit describes a state that has already moved.

Every engine is tested, because measuring one measures roughly a tenth of the surface.

Movement must be separated from churn, which requires a baseline and a fixed cadence.

Engagements under ninety days produce an audit, and should be sold as one.

Citation overlap and domain churn figures per Similarweb analysis, January to February 2026, REPORTED. Time-to-effect range per general practitioner sources, REPORTED. None of these figures is Bermuda-specific, and no Bermuda-specific citation study was located in any source reviewed as at August 2026.

Who in this market is already using these systems

A reasonable objection to everything above: is any of this actually happening in Bermuda, or is it a global trend being sold into a market that has not caught up?

Most attempts to answer that reach for global vendor surveys. There is a better source, and it is the regulator.

The Bermuda Monetary Authority surveyed its own insurance sector on artificial intelligence and machine learning adoption, publishing the results in November 2022 from fieldwork conducted in late 2021. Across all respondents, 38 per cent were using AI or machine learning systems. Among property and casualty insurers specifically, 23.7 per cent. Among insurance groups, 68 per cent. Of the 62 per cent not yet using these systems, 23 per cent planned to adopt within five years.

The split is the finding, not the headline. The Authority attributed the gap between groups and smaller commercial insurers to the scale of the groups' international operations, and that is precisely the same division that determines who your search audience is. The internationally operating entities are simultaneously the larger buyers and the earlier adopters.

Two honest caveats. The fieldwork is from late 2021, so adoption has almost certainly moved, probably substantially. And measuring whether an insurer uses machine learning in underwriting is not the same as measuring whether a broker consults a generative model when comparing domiciles. Those are different behaviours.

Still, it is the only Bermuda-specific institutional AI adoption data in existence that we have located, it comes from the supervisor rather than from a company selling AI services, and it establishes that this market was not starting from zero four years ago.

Regulator survey, not vendor research

Who In This Market Already Uses These Systems

The Bermuda Monetary Authority surveyed its own insurance sector. The split between large and small is the finding, not the headline number.

Insurance groupsusing AI or ML systems
68%
All respondentsacross the surveyed market
38%
Property and casualtyinsurers specifically
23.7%
Planning adoptionwithin five years, of the 62% not yet using
23%

The Authority attributed the gap between groups and smaller commercial insurers to the scale of the groups' international operations. That is the same split that determines who your search audience is: the internationally operating entities are both the larger buyers and the earlier adopters.

Bermuda Monetary Authority, Bermuda Insurance Sector Artificial Intelligence and Machine Learning Survey, 2022 Report, published November 2022, based on a survey conducted in late 2021. VERIFIED against the regulator publication and independent trade press reporting. Note the age: this is a 2021 field date and adoption has almost certainly moved since. It is used here as a documented Bermuda baseline, not as a current reading.

The regulator has since gone further. In July 2025 the Authority published a discussion paper on the responsible use of artificial intelligence in Bermuda's financial services sector, setting out an outcome-based approach to supervision. Regulators publish discussion papers about things their supervised entities are doing, not about things they might do one day.

Global adoption figures point the same way, though they should be labelled clearly as global rather than local. Industry research published in mid-2025 reported that roughly 90 per cent of surveyed insurers were at least evaluating generative AI, with around 55 per cent in early or full adoption, drawn from a US insurance executive sample. Separate tracking reported an 87 per cent year-on-year increase in publicly disclosed AI deployment across insurers into early 2026. A governance survey covering Bermuda directors, reported locally in mid-2026, found 62 per cent agreeing that AI investments were yielding positive results.

None of that measures whether an AI answer influenced a domicile decision. We have said repeatedly that no such measurement exists. What it measures is whether the population you are trying to reach is comfortable operating through these systems, and the answer to that question is not in doubt.

Why the structural work rests on one peer-reviewed study

Almost every quantified claim in the generative optimisation category comes from a company selling the service. That is worth saying out loud, because it applies to most of the figures in circulation, and a compliance function reading your proposal will apply exactly that discount.

There is one substantial exception, and our method is anchored to it.

Research presented at the ACM SIGKDD conference in 2024 tested structural interventions on source visibility inside generated answers, and found that adding statistics, cited quotations and direct-answer structuring raised visibility by roughly 30 to 40 per cent within its experimental benchmark. Independent authors, an established peer-reviewed venue, published methodology.

Its limits should travel with it. The benchmark was experimental rather than a live commercial engine, the models tested are not the models running today, and external validity for a Bermuda reinsurance query in 2026 has not been demonstrated by anyone. We are not going to pretend otherwise.

But it is the difference between a method grounded in something testable and a method grounded in vendor assertion. When a client asks why we structure content the way we do, the answer names a study, a venue and a year, and then names what the study does not establish. That combination is rare enough in this category to be a differentiator on its own.

Panel design for a market of a few hundred entities

Bermuda's addressable set is unusually small and unusually well documented, and that is an advantage for panel construction rather than a limitation.

There were 1,210 insurers and reinsurers on the register at the end of 2025, 608 captives, 21 licensed fund administrators, 766 registered funds and 49 operating digital asset providers. Those entities can be enumerated. Their buyers can be characterised: more than 70 per cent of long-term reserves originate in the United States, with Asia second and Europe including the UK below 5 per cent.

So build the panel from the decision, not from keyword volume, which does not exist here anyway. The prompts institutional buyers actually run cluster into a small number of shapes: comparative domicile questions, requirement and threshold questions, process and timeline questions, and provider-selection questions that only arrive after the domicile question is settled. Each of those shapes needs representation, and provider-selection prompts should be the smallest group rather than the largest, because they are the last question asked and the least frequently searched.

Include prompts you currently lose. This is worth stating twice. A panel composed only of questions where the client already appears will show excellent numbers and teach nobody anything.

Version everything. Each prompt carries the date it entered the panel. When a prompt is retired, it is marked retired rather than deleted, so a client comparing this quarter to last can see exactly what changed in the instrument as well as in the result.

Accuracy is the metric nobody reports

One measure deserves separating out, because it is the one clients most often discover they needed after the fact.

Models say things about your entity. Sometimes those things are wrong: an outdated licence class, a former parent company, a merged entity treated as still separate, a regulatory status that changed two years ago. For a regulated firm this is not a marketing inconvenience, it is a factual misstatement about a supervised business circulating in a channel you do not control.

It is also the only GEO metric that can move backwards for reasons entirely outside your content programme, which is precisely why it should be reported as a standing line item rather than mentioned when convenient. Errors found this period, errors corrected, errors outstanding and the source producing each one.

Correction is usually an entity-consistency problem rather than a content problem. Models reconcile a firm against whatever sources they can resolve, so inaccuracies typically trace to your legal entity name appearing differently across your own site, the regulator register, trade body listings and reference sources. Fixing the inconsistency addresses the cause. Publishing another page does not.

And no agency can compel a model to change its output. What can be done is to remove the material the model is reasoning from and monitor whether the answer follows.

Measure What it genuinely tells you How it gets abused Control
Citation presenceWhether a named source appeared for a specific prompt on a specific date and engineReported as a running total with no denominator, so growth looks automaticAlways report as a rate against a fixed panel, with the panel size shown
Share of voiceRelative visibility against named alternatives within a defined prompt setPanel quietly reweighted toward prompts the client already winsFreeze and version the panel before measuring, include prompts you lose
Answer accuracyWhether what a model says about your entity is factually correctIgnored entirely, because it is the only metric that can move backwardsReport errors found and corrected as a standing line item
Non-brand impressionsReal, unmodelled reach by searcher country from Search ConsoleAggregated globally, hiding that growth came from the wrong countriesSegment by country every time, never report a single global figure
Qualified enquiriesActual contacts, with the discovery channel captured at the point of contactAttributed to AI on the basis of timing coincidence aloneAsk the enquirer, record the answer verbatim, and accept that many will not know
Pipeline influenceThat a visibility programme was running while a mandate progressedConverted into a return figure by multiplying against mandate valueReport separately, state the attribution method beside it, never multiply

What a defensible report looks like

Concretely, for a Bermuda platform, a quarterly report that survives scrutiny contains: the versioned prompt panel with its composition and change history; citation rate per engine against that panel with dated raw captures archived; share of visible answers with the panel shown beside it; an accuracy line covering errors found, corrected and outstanding; non-brand impressions segmented by searcher country; qualified enquiries with self-reported discovery channel; and a separate, unmultiplied note on pipeline activity during the period with the attribution method stated.

It does not contain a return figure. It does not contain a forecast. It does not contain a number whose provenance cannot be traced in one step.

That report is shorter and less impressive than the alternative. It is also the version that can be handed to a chief risk officer without anyone having to explain how the headline was calculated, and in this market that property is worth more than the headline would have been.

Tessar Napitupulu writes about measurement discipline in AI visibility programmes, and about the difference between what is observed and what is inferred, in Cited or Silent, available as a free gated edition, with retailer editions on Amazon, Google Play and Apple Books.


Frequently Asked Questions


What is the difference between a ranking and a citation?

A ranking is your position in a list of links a person may or may not click. A citation is a model naming you as a source inside a generated answer, frequently with no click at all. They are produced by different mechanisms and they move independently. A page can rank first and never be cited, and a page well outside the first results can be cited repeatedly. In financial content the gap is unusually wide: one tracked analysis found only around 11 per cent of AI answer citations in finance also appeared in the organic top ten, the lowest overlap of any industry measured, with roughly two thirds coming from pages outside the top hundred.


Can you guarantee our firm will be cited in ChatGPT or Google AI Overviews?

No, and any agency offering that guarantee is describing something outside its control. The engines control retrieval and generation, their source preferences shift, and analysis of roughly 600,000 citation events found around half of cited domains changing month to month. Citation is also not purchasable: platform operators state that product results in generated answers are selected independently of advertising and partnership arrangements, with sponsored placements labelled separately. What can be committed to is method and deliverables, meaning a versioned prompt panel, a fixed testing cadence, dated raw captures and correction of inaccuracies within your control.


Why do you refuse to report a return on investment figure for GEO?

Because the causal chain cannot be measured, and manufacturing one would be dishonest to a reader whose compliance function will check it. There is no observable path from a citation appearing in an AI answer to a treaty placement or a fund mandate signed months later, often after several offline conversations. Multiplying a citation count by an average mandate value produces a number that looks rigorous and means nothing. We report visibility measures that are real, report pipeline separately with the attribution method stated alongside it, and leave the two unmultiplied.


How many prompts should be monitored, and how often?

Industry practice reported across research sources ranges roughly from twenty to a hundred prompts, with some practitioners suggesting thirty to fifty tested across five or more engines. We treat that as context rather than as a standard, and we set panel size from the actual decision structure of your buyers rather than from a benchmark. What matters more than the count is that the panel is versioned, that every prompt carries the date it entered the panel, and that testing happens on a fixed schedule rather than when someone remembers.


Why monitor multiple engines instead of the largest one?

Because they do not agree. Analysis of approximately 600,000 citation events during early 2026 found only around 11 per cent overlap between the domains ChatGPT and Perplexity cite for comparable queries. Optimising for one engine therefore leaves most of the answer surface unmeasured. The same analysis found roughly half of cited domains rotating month to month, which means a single point-in-time audit describes a state that has already changed by the time it is presented.


How long before GEO work produces measurable movement?

Sources in this field generally suggest a minimum of sixty to ninety days, because models need to recrawl and reprocess sources before changes to your material can affect what they retrieve. Engagements shorter than about ninety days tend to deliver an audit rather than a result, which is a reasonable thing to buy if an audit is what you want, but it should be named as such. For a market where the buying cycle itself runs in quarters rather than weeks, this is rarely the binding constraint.


What should we do if a model states something inaccurate about our entity?

Document it first, with the engine, the exact query, the date and the verbatim output, because an undated screenshot is not evidence of anything. Inaccuracies usually trace to inconsistent legal entity naming across your website, the regulator register, trade body listings and reference sources, or to outdated third-party material that is easier for a model to retrieve than your current position. Correct what is within your control, pursue correction of the third-party material where that is possible, and monitor whether the output changes. No agency can compel a model to change what it says.


Is share of voice a reliable GEO metric?

It is useful and it is fragile, and both need saying. Share of voice across a fixed prompt panel is a legitimate comparative measure of how often you appear relative to named alternatives. Its fragility is that the denominator is a panel you chose, so it can be inflated by choosing prompts you already perform well on. The controls are to version the panel, freeze it before measuring, include prompts where you currently do badly, and report the panel composition alongside the figure rather than only the figure.


Is there any Bermuda-specific evidence that this market uses AI at all?

Yes, from the regulator rather than from a vendor. The Bermuda Monetary Authority surveyed its own insurance sector and published results in November 2022, based on late-2021 fieldwork: 38 per cent of all respondents using AI or machine learning, 23.7 per cent among property and casualty insurers specifically, and 68 per cent among insurance groups, with the Authority attributing that gap to the scale of the groups' international operations. Of the 62 per cent not yet using these systems, 23 per cent planned adoption within five years. The Authority followed this with a July 2025 discussion paper on responsible AI use in the territory's financial services sector. Two caveats travel with the figures: the fieldwork is from 2021 and adoption has almost certainly moved, and using machine learning in underwriting is a different behaviour from consulting a generative model when comparing domiciles.

Sources & References:

  • Overlap between AI answer citations and organic ranking in finance: approximately 11.3 per cent of AI Overview citations in finance also ranking in the organic top ten, the lowest of any tracked industry; approximately 66 per cent of finance citations originating from pages outside the top hundred; approximately 17 per cent overlap across YMYL categories generally. MADX Digital analysis. REPORTED, single vendor source, global rather than Bermuda-specific, presented with that limitation stated in the body text.
  • Cross-engine citation divergence and domain churn: approximately 11 per cent overlap between domains cited by ChatGPT and Perplexity, and approximately 50 per cent of cited domains changing month to month, from analysis of approximately 600,000 citation events during January and February 2026. Similarweb. REPORTED.
  • Time to measurable effect of sixty to ninety days minimum, reflecting model recrawl and reprocessing cycles. REPORTED from general practitioner sources across the research reviewed; no controlled study establishing this interval was located.
  • Bermuda Monetary Authority Annual Report 2025 figures as tabled in the House of Assembly, June 2026: 1,210 insurers and reinsurers on the register at 31 December 2025; 21 fund administrator licences; 766 registered funds; 49 operating digital asset business providers. VERIFIED. Captive count of 608 at end-2025 per BMA registration statistics reported in trade press.
  • Buyer geography: more than 70 per cent of Bermuda long-term insurers' reserves originating in the United States, Asia second, Europe including the United Kingdom below 5 per cent. BMA, Bermuda Long-term Insurance Market Analysis and Stress Testing Report, December 2025. VERIFIED.
  • Prompt panel size: industry practice reported in the range of approximately twenty to one hundred prompts, with some practitioners suggesting thirty to fifty tested across five or more engines. REPORTED as context. This is not presented as a standard, and no deliverable commitment to a specific prompt count is implied.
  • Platform position on paid placement: operators of major generative search products state that product results within generated answers are selected independently of advertising and partnership arrangements, with sponsored placements labelled separately. REPORTED, and subject to change by the platforms concerned.
  • Bermuda Monetary Authority, Bermuda Insurance Sector Artificial Intelligence and Machine Learning Survey, 2022 Report, published November 2022 from fieldwork conducted in late 2021: 38 per cent of all respondents using AI or machine learning systems; 23.7 per cent among property and casualty insurers; 68 per cent among insurance groups, a gap the Authority attributed to the scale of their international operations; and 23 per cent of the 62 per cent of non-adopters planning adoption within five years. VERIFIED against the regulator publication and corroborated by independent trade press reporting. The 2021 field date is stated in the body text because adoption is likely to have moved since.
  • Bermuda Monetary Authority discussion paper on the responsible use of artificial intelligence in Bermuda's financial services sector, published July 2025, setting out an outcome-based supervisory approach. VERIFIED.
  • Global insurance sector AI adoption context, labelled as global rather than Bermuda-specific throughout: industry research published mid-2025 reporting approximately 90 per cent of surveyed insurers evaluating generative AI with approximately 55 per cent in early or full adoption, from a United States insurance executive sample; separate tracking reporting an 87 per cent year-on-year increase in publicly disclosed AI deployment into early 2026. REPORTED.
  • Governance survey covering Bermuda directors, reported locally in mid-2026, finding 62 per cent agreeing that AI investments were yielding positive results. REPORTED, regional sample.
  • Generative engine optimisation structural findings: research presented at the ACM SIGKDD conference in 2024 reporting source visibility improvements of approximately 30 to 40 per cent from statistics, cited quotations and direct-answer structuring within its experimental benchmark. Peer-reviewed, independent authorship. External validity for current commercial models and for Bermuda-specific queries has not been demonstrated, and that limitation is stated in the body text.
  • No published measurement of AI citation shares for Bermuda domicile queries, and no study of AI system use in captive or reinsurance domicile selection, was located in any of the eight research documents reviewed for this article, nor in subsequent verification, as at August 2026. Figures cited above are global or sector-general and are labelled as such throughout.
  • The four-tier measurement framework and the exclusion of return-on-investment calculation from GEO reporting are methodological positions taken by the author, not findings from a cited study.
0 Comments 0 Comments
0 Comments 0 Comments