The Irish AI Citation Gap Nobody Has Measured Yet
SEO

The Irish AI Citation Gap Nobody Has Measured Yet

No published study compares local and British domains in AI answers. Here is the one direct test that exists, and how to run your own.

The first thing an Irish business wants to know about AI search is straightforward. When someone asks ChatGPT or Google for a recommendation in our category, does the answer cite Irish sources or British ones?

There is no published measurement of that. Four independent research passes were run for this article across different model families, and all four returned the same verdict: no controlled study measures the share of Irish versus British domains cited for Ireland-intent commercial prompts across any major assistant. If someone quotes you a precise percentage, they constructed it.

What does exist is one direct test, run in August 2026, and it points the opposite way to the claim most agencies repeat. That is the awkward part, and it is where this piece starts.

What the Only Direct Test Found

Nine commercial prompts were put through a retrieval layer, each one explicitly naming Ireland: payroll software, CRM, accounting, HR, cybersecurity consulting, business energy, digital marketing agency, business insurance broker and e-commerce platform. Every grounding URL the system returned was collected and classified.

Nine Ireland-qualified commercial prompts, August 2026
52 grounding URLs returned, classified by domain
The popular claim is that British content crowds out Irish businesses in AI answers. On explicitly Ireland-qualified prompts, this run found no British commercial domains at all.
30Irish domains
22Other, mostly global
0UK
Irish domains, 57.7% Non-Irish, 42.3% British commercial domains, zero

What this test did not do

One engine. One run. No control for IP geolocation. And, critically, it did not test the same prompts with the word Ireland removed. That unqualified condition is the one real buyers are in most of the time, and it stayed untested.

Direct retrieval test conducted 18 August 2026 as part of the research compiled for this article. Reported here with its limitations attached because it is the entire published evidence base on the question, and because it contradicts the prevailing claim rather than supporting it.

A separate exploratory audit of six explicitly commercial Irish phrases produced a consistent picture: results were dominated by Irish agencies, Irish directories and Ireland-specific pages. Clear British displacement did not appear there either.

Two tests, both pointing the same way, both on prompts that name the country. So the honest headline is not "British content dominates Irish AI answers." It is closer to "when the prompt says Ireland, the systems appear to listen, and nobody has checked what happens when it does not."

Why That Test Does Not Settle It

It would be convenient to stop there and tell Irish clients the problem is overstated. That would be wrong, for four reasons worth stating precisely.

A single run on a single engine is an observation, not a measurement. Assistants are probabilistic. The same prompt can return different grounding sets on consecutive runs, and different engines use different retrieval architectures. Nine prompts through one system on one day establishes a direction, not a rate.

There was no geolocation control. The test relied on the prompt text carrying the Irish signal rather than on the request originating from an Irish IP address. Those are different conditions and they may produce different behaviour.

The unqualified prompt was never tested. This is the gap that matters commercially. When an Irish buyer types "best payroll software for a small business" without adding "in Ireland," what happens? Nobody has published an answer. That is not a minor omission; it is probably the majority of real queries.

And the structural pressure toward British sources is real regardless of what one test showed. That pressure comes from language, and it is measurable from a different direction entirely.

The Language Shield Ireland Does Not Have

Research into how AI search engines cite sources across languages and countries produced a finding that explains Ireland's position better than any Ireland-specific study currently does.

When a question is asked in French, roughly 16% of cited sources are French domains. Ask the same question in English, and that drops to about 1.4%. In Japanese, around 26% of citations go to Japanese domains, against roughly 1% for English-language equivalents. In Mexico, about 96% of cited sources were Spanish-language.

So the local language acts as a filter. Asking in French pulls French sources into the answer, because the corpus the system searches in French is overwhelmingly French. That is a structural protection, and it is granted automatically by the act of not speaking English.

MarketLocal-language protectionWhat happens on the English-language query
FranceStrong. Roughly 16% of citations to French domains on French-language queriesFalls to around 1.4%. The protection is real but only applies in French.
JapanStrong. Around 26% of citations to Japanese domains on Japanese queriesFalls to roughly 1%. Same pattern, larger gap.
MexicoVery strong. Around 96% of cited sources in SpanishNot the primary query condition for most local buyers.
IrelandNone available. The commercial query language is EnglishThis is the only condition Ireland has. There is no second language to retreat into.

Ahrefs measured the scale of that pool directly. Across 108 million AI Overview queries analysed in November 2025, English accounted for 52.75% of every AI Overview generated worldwide. Spanish, the next largest, took 11.48%. Portuguese 5.01%. Japanese 4.37%. German 3.82%.

More than half of all AI Overviews in the world are in the language Irish businesses compete in. Ireland is inside the largest, most crowded corpus that exists, with no linguistic boundary marking where Irish content ends and British, American, Australian or Canadian content begins.

That is the structural argument, and it is genuinely strong. It is also not the same thing as a measurement, and the difference matters. A well-supported hypothesis about mechanism is not evidence about magnitude.

Why the Basque Study Is the Wrong Analogy

The study most often reached for here examined territorial bias in AI citations for Basque-related topics. It found that only 15.5% of citations went to local or regional sources, while 67.4% went to national-level media. Published in Hipertext.net in 2025 by Sarrionandia, Peña-Fernández and Pérez-Dasilva.

It is a real finding and it does demonstrate territorial bias exists. But applied to Ireland it understates the problem, and it is worth being clear why.

The Basque case measures regional media losing out to national media inside a single country. One jurisdiction, one regulator, one currency, one legal system. The gradient runs from local to national within a shared state.

Ireland's asymmetry is different in kind. A sovereign national corpus competes against a substantially larger neighbouring corpus in an identical language, across a border that separates two currencies, two tax systems, two sets of regulators and two legal orders. The content looks interchangeable to a retrieval system and is legally and commercially not interchangeable at all.

That is arguably harder than the Basque situation, not easier. A Basque reader served national Spanish media gets content from their own country. An Irish reader served British content gets content from a different jurisdiction, with the wrong regulator, the wrong currency and the wrong consumer protections, presented with no visible signal that anything is off.

Exposure Splits by Query Class

Pulling the direct test and the structural argument together produces something more useful than either alone. Exposure is not uniform. It varies by what the buyer is asking for.

Four exposure profiles, one keyword universe
Reporting these as one blended number destroys the signal

Supplier selection with an explicit Ireland qualifier

Appears to localise strongly. The one direct test returned 30 Irish domains from 52 grounding URLs and zero British commercial domains.

Tested once, direction established

Informational and how-to queries

Most exposed to the larger British corpus. Corroborated indirectly: Irish informational terms held only 1.16% click-through at positions four to ten, against 4.24% for commercial-investigation terms, and AI Overviews trigger more often on informational queries.

Inference, supported from two directions

Product comparison and best-of queries

Likely to face British media, affiliate sites and marketplaces with far larger content volumes and stronger domain histories.

Inference, not measured

Regulated topics: legal, financial, medical

Likely to hold Irish sources, because the authoritative source genuinely is Irish. A question about Irish financial consumer protection has no correct British answer.

Inference, not measured

Unqualified prompts from users physically in Ireland

Completely unknown. Not tested by anyone, and probably the largest share of real queries.

Unmeasured
Confidence markers are deliberate. One profile has a single direct observation. One has indirect corroboration from click data. Two are reasoned inference. One is a blank. Any proposal that assigns a single displacement percentage across all five is not measuring anything.

Why Domain Extension Coding Produces the Wrong Answer

Most attempts to measure this classify citations by top-level domain. Count the .ie results, count the .co.uk results, report a ratio. It is the obvious method and it is misleading in both directions.

A .com can be an Irish publisher. Plenty of established Irish businesses, media outlets and professional bodies operate on .com domains for historical or commercial reasons. Counting those as non-Irish undercounts Irish representation.

A .ie can be a reseller, a lead-generation site, or a foreign company's Irish landing page with no Irish operation behind it. Counting those as Irish overcounts it. Worse, it can score a competitor's thin Irish microsite as a local win.

The classification that actually answers the buyer's question is publisher jurisdiction: where is this organisation established, which regulator governs it, whose consumer law applies to a transaction with it. That requires looking at the entity, not the string after the final dot. It is slower. It is also the only version of the number that means anything.

The same logic applies to how the monitoring itself runs. Simulating an Irish locale through an API parameter is not the same as querying from an Irish-geolocated browser, and source substitution is precisely the behaviour a simulated locale is most likely to hide.

Selection Is Not Absorption

One further distinction, because it is where most citation reporting quietly overstates results.

Selection means a source was retrieved and displayed in the citation list. Absorption means the content of that source actually shaped the claims, wording or evidence in the generated answer. These come apart constantly. A page can sit in the citation list contributing nothing substantive. An answer can lean heavily on material it never links.

There is a third failure mode that damages brands most quietly: being cited accurately as a source while being described inaccurately as a business. Wrong service list, wrong market, wrong regulatory status. The citation looks like a win in a dashboard and reads like a liability to a buyer.

Reporting selection alone flatters the numbers. A defensible report separates brand mention, linked citation, recommendation, description accuracy and absorption, and tracks stability across repeated runs. That separation is central to how we structure GEO measurement for Irish clients, and to the general method in our core generative engine optimisation practice.

How to Build a Panel That Actually Measures This

Six design decisions, in order of how much they affect the result.

Split the panel in two halves. Half the prompts name Ireland explicitly. Half do not, asked from an Irish-geolocated browser. Comparing those two halves on your own category is the measurement the market is missing, and you can produce it for yourself in a fortnight.

Freeze and version the prompt set. Every prompt tagged with intent, persona, funnel stage, language and qualifier state. Once you start editing prompts mid-programme, run-to-run comparison stops being valid and you will not notice.

Code citations by publisher jurisdiction, not extension. Irish, British, global, competitor, regulator, media, directory, user-generated. Slower, correct.

Run from real Irish-geolocated browsers. Not API locale parameters. Record which assistant, which mode, which date, and whether a web search was triggered at all, because sometimes it is not.

Repeat runs and record the variance. Citation stability across reruns is itself a finding. A source cited in one run of five is not the same result as one cited in five of five, and a single screenshot cannot tell them apart.

Preserve the failures. Runs where you were absent, described wrongly, or where the assistant returned nothing useful. Deleting those turns a measurement programme into a highlight reel.

What Would Settle the Question

Worth naming what an actual answer would require, since the field keeps producing approximations of it.

A representative sample of Irish commercial intents, several hundred prompts at minimum, spanning categories. Both qualifier conditions. Multiple assistants, tested in parallel on the same days. Repeated runs to establish variance. Requests originating from Irish IP addresses. Citations classified by publisher jurisdiction rather than domain string. And publication of the prompt list so the work can be reproduced.

Nobody has done that. It is not technically difficult, it is just unglamorous and expensive, and it produces a number that might disappoint whoever funded it. The academic work that does exist on optimising for generative engines, including the GEO paper presented at ACM SIGKDD in 2024, demonstrated visibility improvements of up to 40% under experimental conditions across roughly 10,000 queries. That is real evidence that content structure influences citation. It is not evidence about Irish jurisdiction mix, and it should not be presented as such.

Until the Irish study exists, the defensible position is the one this article has tried to hold: name the mechanism confidently, name the magnitude as unmeasured, and build the instrument on your own prompt set rather than borrowing a percentage from a market that is not yours. The organic side of the same problem follows the same discipline.

The argument that citation and ranking are separate disciplines requiring separate instruments is developed at length in Cited or Silent. The gated edition is free to download, and our own cross-market measurement work is published in the AI Citation Rate Report 2026.

Written by Tessar Napitupulu, Founder and CEO of PT Arfadia Digital Indonesia, a member of the Forbes Agency Council, and author of Found Before They Search and Cited or Silent. Arfadia has documented its generative engine optimisation practice since 2023 and works from Jakarta, Bandung and Bali.




Frequently Asked Questions


Has anyone measured whether AI answers cite Irish or British sources?

No controlled, published study measures this. Four independent research passes across different model families were run for this article and all four returned the same verdict on the Irish versus British citation split for Ireland-intent commercial prompts: unavailable. The nearest evidence is a single retrieval test of nine Ireland-qualified commercial prompts in August 2026, which returned 30 Irish domains from 52 grounding URLs and no British commercial domains at all. One engine, one run, and it did not test the unqualified version of the same prompts.


Does British content dominate Irish AI answers?

The claim is widely repeated and the only direct evidence available points the other way. On prompts that explicitly name Ireland, Irish domains made up 57.7% of grounding URLs in the one test that exists, with zero British commercial domains. A separate audit of six commercial Irish phrases also found Irish agencies, directories and Ireland-specific pages dominating. Neither test examined what happens when the prompt omits the country, which is probably the majority of real queries, so the honest answer is that displacement is structurally plausible and quantitatively unmeasured.


Why does Ireland face a harder problem than France or Japan?

Because English gives it no protective boundary. Research into citation behaviour across languages found roughly 16% of citations going to French domains on French-language queries, dropping to about 1.4% on English-language ones, and around 26% to Japanese domains on Japanese queries against roughly 1% in English. The local language acts as a filter. Ireland's commercial query language is English, and Ahrefs found English accounts for 52.75% of all AI Overviews worldwide from a 108 million query dataset. Ireland competes inside the largest corpus that exists with no second language to retreat into.


Is the Basque territorial bias study a good analogy for Ireland?

It is useful but it understates Ireland's problem. The study, published in Hipertext.net in 2025 by Sarrionandia, Peña-Fernández and Pérez-Dasilva, found only 15.5% of citations going to local or regional sources for Basque topics against 67.4% to national media. That measures regional versus national media inside one country, with one regulator, one currency and one legal system. Ireland's asymmetry is a sovereign national corpus competing against a much larger neighbouring corpus in an identical language across a border separating two currencies, two tax systems and two legal orders. The content looks interchangeable to a retrieval system and is not interchangeable in law.


Should we classify citations by domain extension?

No, and this is the most common methodological error in Irish citation reporting. A .com can be an established Irish publisher, so counting it as non-Irish undercounts Irish representation. A .ie can be a reseller, a lead-generation site or a foreign company's Irish landing page with no Irish operation behind it, so counting it as Irish overcounts. The classification that answers the buyer's question is publisher jurisdiction: where the organisation is established, which regulator governs it, whose consumer law applies. That requires examining the entity rather than the string after the final dot.


What is the difference between citation selection and absorption?

Selection means a source was retrieved and displayed in the citation list. Absorption means the content of that source actually shaped the claims, wording or evidence in the answer. They separate constantly: a page can appear in the citation list while contributing nothing substantive, and an answer can lean heavily on material it never links. There is a third failure mode worth tracking separately, which is being cited accurately as a source while being described inaccurately as a business. That looks like a win in a dashboard and reads like a liability to a buyer.


Can we simulate an Irish location through an API parameter?

You can, but it undermines the measurement. Source substitution is precisely the behaviour a simulated locale is most likely to hide, because the retrieval layer may treat a locale parameter differently from a genuine request origin. Monitoring should run from real Irish-geolocated browsers, recording which assistant, which mode, which date, and whether a web search was triggered at all, because sometimes it is not.


Which types of Irish query are most at risk?

Exposure splits by query class. Supplier-selection queries carrying an explicit Ireland or Dublin qualifier appear to localise strongly, based on one direct test. Informational and how-to queries look most exposed, and that inference is supported from two directions: Irish informational terms held only 1.16% click-through at positions four to ten against 4.24% for commercial-investigation terms, and AI Overviews trigger more often on informational queries. Product comparison queries face British media, affiliates and marketplaces with larger content volumes. Regulated topics probably hold Irish sources because the authoritative source genuinely is Irish. Unqualified prompts from Irish users are unmeasured.


How long does it take to measure this for our own business?

A first baseline is achievable in a fortnight. Build a frozen prompt panel split into two halves, one naming Ireland explicitly and one not, run it across the major assistants from Irish-geolocated browsers, repeat the runs to establish variance, and classify every citation by publisher jurisdiction. The result is specific to your category, which makes it more useful than any national average would be, and you own the dataset regardless of who runs the programme afterwards.


Does the academic research on generative engine optimisation prove citation can be improved?

It demonstrates that content structure influences citation under experimental conditions. The GEO paper presented at ACM SIGKDD in 2024, arXiv reference 2311.09735, reported visibility improvements of up to 40% across a benchmark of roughly 10,000 queries. That is genuine evidence that structural intervention works. It is not evidence about Irish jurisdiction mix, not a client guarantee, and not transferable to a specific market without testing. Treat it as proof of mechanism rather than proof of outcome.

Sources & References:

  • Direct retrieval test conducted 18 August 2026 for this research programme. Nine commercial prompts, each explicitly naming Ireland, covering payroll software, CRM, accounting, HR, cybersecurity consulting, business energy, digital marketing agency, business insurance broker and e-commerce platform. 52 grounding URLs collected: 30 Irish domains, 22 non-Irish, zero British commercial domains. Limitations stated by the researcher: one engine, one run, no IP geolocation control, and no test of the same prompts with the Ireland qualifier removed.
  • Separate exploratory audit of six explicitly commercial Irish search phrases, finding results dominated by Irish agencies, Irish directories and Ireland-specific pages, with no clear British displacement observed.
  • Weglot, "How AI Search Engines Cite Sources by Language and Country", published August 2026. Methodology: 400 questions across 5 topic categories and 3 levels of specificity, written in English, French, Spanish and Japanese, put to 4 AI systems across 5 countries, producing more than 16,000 answers with every cited source logged. French-language queries sent 16% of citations to .fr domains against 1.4% for English queries, with French government .gouv.fr domains rising from near-zero in English to 1.25% in French. Japanese-language queries sent 26% of citations to .jp domains against approximately 1% in English, with .go.jp and .ac.jp at 4.26% in Japanese against 0.13% in English. Spanish-language queries in Spain sent 7% of citations to .es domains against approximately 1% in English. A separate Weglot study of the Mexican market found 96% of AI Overview citations came from Spanish sources. Ireland and the Irish language were not among the cases tested.
  • Ahrefs Brand Radar, "Which Countries Have the Most AI Overviews", published 4 November 2025. 108 million AI Overview queries analysed. English accounted for 52.75% of all AI Overviews, Spanish 11.48%, Portuguese 5.01%, Japanese 4.37%, German 3.82%. Ireland showed AI Overviews on 12.00% of tracked queries, ranked 44th of 50 countries.
  • Sarrionandia, Peña-Fernández and Pérez-Dasilva, study of territorial bias in AI citation for Basque-related topics, Hipertext.net, 2025. 15.5% of citations to local or regional sources; 67.4% to national-level media. Included here with an explicit statement of why the analogy understates Ireland's structural position.
  • Friday.ie, Irish search click-through analysis, published 6 July 2026. Commercial-investigation terms held 4.24% click-through rate at positions four to ten against 1.16% for informational terms, used here as indirect corroboration of the query-class exposure split rather than as citation evidence.
  • Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, "GEO: Generative Engine Optimization", ACM SIGKDD 2024, arXiv:2311.09735. Visibility improvements of up to 40% under experimental conditions across a benchmark of approximately 10,000 queries. Cited as evidence of mechanism, not of market-specific outcome.
  • Share of AI citations for Ireland-intent commercial prompts going to Irish versus British publishers: UNAVAILABLE. No controlled published study located across four independent research passes.
  • Citation behaviour for prompts issued without an Ireland qualifier from Irish IP addresses: UNAVAILABLE. Not tested by any source located.
  • Citation behaviour for prompts written in Irish: UNAVAILABLE. Flagged as unmeasured by all four research passes.
0 Comments 0 Comments
0 Comments 0 Comments