Here is a question a Luxembourg marketing director can reasonably ask, and which nobody can currently answer with a measurement: when someone in Luxembourg asks an AI assistant to recommend a fund administrator, a private clinic or a fiduciary, which language does the assistant answer in, and which language are the sources it cites written in?
Four separate research passes went looking for that measurement while this article was being prepared. None of them found it. There is no published, controlled study of assistant language behaviour for queries originating in Luxembourg across Luxembourgish, French, German and English.
That absence matters more here than almost anywhere. This is a country with three administrative languages, English as the effective working language of its largest industry, 46.6% of residents holding foreign nationality across 180 nationalities, and a workforce nearly half of which commutes in from three neighbouring countries. Getting the language decision wrong is not a rounding error in the content plan. It is the content plan.
What follows is everything that is actually known, sorted by how much weight it can carry, followed by the test that would settle the rest.
What Google's own documentation already settles
Start with the part that is not in dispute, because it is published by the platform itself and takes about ten seconds to verify.
Luxembourg appears on Google's official supported-country lists for both AI Overviews and AI Mode. That is worth stating precisely, because for a long period it was genuinely unclear. The European expansion of AI Overviews in March 2025 covered nine countries, and Luxembourg was not among them. France did not receive either feature until 22 July 2026, after a prolonged dispute over publisher compensation. Neighbouring launch dates were never transferable to Luxembourg, and no Luxembourg-specific announcement was ever made. The supported-country list is what resolves it. So this is not a wait-and-see market.
Now read the other column on the same two pages, the one listing supported languages. Luxembourgish is not on it. Not for AI Overviews, and not for AI Mode.
What is on it makes the omission harder to shrug off. Romansh is supported, and Romansh is Switzerland's fourth official language with roughly 60,000 speakers. Faroese is supported, with around 70,000. Maltese, Western Frisian, Irish, Welsh, Basque, Galician, Wolof and Kinyarwanda are all supported. Luxembourgish has something in the order of 400,000 speakers and is the national language of a European Union member state.
So the practical answer to "do we need Luxembourgish content to be visible in Google's AI features" is currently no, and the reason is not that Luxembourgish underperforms. It is that the language is not processed by those features at all. Language support lists change without notice, which is why every claim built on this needs a date attached and a quarterly re-check.
What peer-reviewed research says about Luxembourgish and language models
The platform documentation tells you about product coverage. It does not tell you why. For that there is academic work, and it is unusually specific for a language this size.
Research published in Findings of the ACL: EACL 2026 by Lothritz, Cabot and Bernardy tested large language models against Luxembourgish language proficiency examinations. The results split by model scale. Frontier models performed strongly on the exams. Smaller models did not, and their failures were characterised rather than merely counted: mistranslated words, code-switching into German or French mid-output, incorrect noun gender, and violations of the Eifeler Regel, the phonological rule governing when a final n is dropped in Luxembourgish.
That last detail is the one worth holding onto. A model that violates the Eifeler Regel is not making a stylistic choice. It is producing text a Luxembourgish reader will immediately register as wrong, in the same way an English reader registers a misplaced apostrophe. Failure is visible to the audience, not just to a benchmark.
Around that paper sits a small research ecosystem. Generation models LuxGen and LuxT5 were presented at VarDial 2025. LuxemBERT, published at LREC 2022, was developed in collaboration between a Luxembourg bank and the University of Luxembourg's interdisciplinary security and trust centre. More recent benchmark work includes LuxIT, LexNeo-Bench and ltzGLUE. The recurring theme across all of it is data scarcity: Luxembourgish is a low-resource language, UNESCO classifies it as vulnerable, and the corpora available for training and evaluation are small compared with the languages models handle fluently.
Two honest qualifications. First, this research concerns language capability, not citation behaviour. A model can write acceptable Luxembourgish and still never cite a Luxembourgish source. Second, capability is improving, and the frontier-model results show the direction. Nothing here supports a claim that Luxembourgish will remain permanently unusable.
What Is Known Per Language, and What Is Not
The honest position differs by language, so a single multilingual strategy statement covering all four would have to be wrong about at least three of them.
Best supported, and the retrieval default
Supported across AI features. Research on multilingual retrieval shows configurations that include English tend to cite English documents more often. Working language of finance, funds and the EU institutional cluster.
Well supported, largest workplace footprint
Supported. Used at work by 69.2% of employed residents and the main working language of 56% of companies. Sole language of legislative drafting. Also the language of most cross-border commuters.
Well supported, demand-led rather than default
Supported. Used at work by 29.5% of employed residents, main working language of 6% of companies, yet dominant in the print press. Justified where genuine DACH or eastern-border business exists.
Not a supported language for Google AI features
Absent from both supported-language lists as at August 2026. Characterised in peer-reviewed work as low-resource with documented generation failure modes. Main language of 48.9% of residents and spoken by 61.2%, so real for brand and civic purposes.
Language statistics from STATEC. Support status from Google published language lists, checked August 2026 and subject to change without notice.
What multilingual retrieval research says about citation, as opposed to fluency
Capability research asks whether a model can write a language. A different body of work asks a question closer to ours: when a system retrieves documents to ground an answer, does the language of those documents skew?
It does. Work published at ACL 2026 by Wang and colleagues examines language bias in multilingual retrieval-augmented generation and documents bias in reranking both towards English and towards the language of the query. A separate benchmark study in Findings of ACL 2025 on cross-lingual robustness found that configurations including English cited English-language documents more frequently than configurations restricted to the languages actually relevant to the question.
Both are peer-reviewed, which puts them above anything else in this article for reliability. Both also come with the same limitation, and it is a large one: they are experimental systems evaluated on multi-language benchmarks. Neither is a measurement of a commercial assistant answering a Luxembourg buyer. The direction is informative. The magnitude for this market is not established.
Two vendor studies, labelled as vendor studies
Beyond the academic work there are commercial analyses with much larger sample sizes and much weaker independence. They are worth reading and worth discounting.
One monitoring vendor analysed more than ten million prompts and more than twenty million query fan-outs, the internal sub-queries a system generates while assembling an answer. Roughly 78% of non-English prompts triggered at least one English-language fan-out, and around 43% of all fan-outs arising from non-English prompts were executed in English. The reported pattern is that the first fan-out usually stays in the user's language while English enters later in the chain. If that holds, a French-language query about a Luxembourg service may still be partly resolved against English sources, without the user ever seeing that happen.
A second study examined roughly 12,000 pages across 15 languages and reported that non-English content received about 34% fewer direct citations on average, with Western European languages including French and German showing a smaller gap, and approximately 31% of citations still going to English-language sources even when the query was in a local language. Luxembourgish was not among the 15 languages tested.
Both figures are single-vendor, commercially interested, and not reproducible by a reader. Use them to form a hypothesis. Do not put them in a board pack as findings.
So what is actually missing
Assemble everything above and a specific shaped hole remains. We know Luxembourg is covered by the features. We know Luxembourgish is not a supported language for them. We know retrieval systems in experimental settings lean towards English and towards the query language. We have two vendor estimates suggesting English pulls citation share even from non-English queries.
What we do not know, for this market, is any of the following.
Which language an assistant actually answers in when a query originates on a Luxembourg network, tested across the four candidate languages. Which source language is cited most often for Luxembourg-origin prompts. Whether publishing a French version rather than an English one changes citation likelihood at all, in either direction. Whether adding a Luxembourgish version earns any citations on any platform. And which assistants Luxembourg finance professionals actually use when researching a vendor, since no sector survey exists and country-level platform share for Luxembourg alone is thin enough that we decline to present European or global share as a Luxembourg number.
Five open questions, all commercially consequential, none answered by any of the eight research sources reviewed for this cycle. That is not a gap to paper over with a confident recommendation. It is a research brief.
It is worth setting them out alongside the nearest available evidence, because in each case something exists and in each case it falls short in a specific way.
| Open question | Closest available evidence | Why it falls short |
|---|---|---|
| Which language does an assistant answer in for a Luxembourg-origin query? | Platform documentation confirming that location can influence relevance | Confirms a mechanism exists, says nothing about the outcome across four candidate languages |
| Which source language is cited most for those queries? | Peer-reviewed multilingual retrieval studies showing bias towards English and the query language | Experimental systems on multi-language benchmarks, not commercial assistants answering Luxembourg buyers |
| Does a French version rather than an English one change citation likelihood? | Vendor study reporting non-English content receiving about 34% fewer direct citations | Single vendor, commercially interested, not reproducible, and averaged across 15 unrelated languages |
| Would a Luxembourgish version earn citations anywhere? | Absence of Luxembourgish from Google's AI feature language lists, plus peer-reviewed low-resource findings | Answers the Google case and the capability case. Says nothing about other platforms, and never measured for any language pairing |
| Which assistants do Luxembourg finance professionals use for vendor research? | European and global platform share figures from panel measurement | Not Luxembourg, and usage share does not indicate which platform produces commercial discovery in a sector |
Read down the third column and a pattern emerges. Every one of these questions has evidence adjacent to it, and every piece of that evidence is either about a different geography, a different system, or a different question. Adjacent evidence presented as an answer is how a plan acquires false confidence.
The Measurement That Would Settle It
Hold these constant, vary those, record all of it. The output is a dataset the client keeps, not a score the agency owns.
Hold constant
Clean accounts with no personalisation history. Interface language and preferred-language settings recorded. Same prompt wording within each language. Same competitor list. Same product surface, since a chat interface and a search AI mode are not the same system.
Vary deliberately
Prompt language across the four candidates. Location condition, treating approximate and permitted precise location as separate cases. Platform. Day of observation. Intent type, keeping transactional, comparison, reputational and regulated-finance prompts apart.
Record for every run
Timestamp, engine, surface, full answer text, every cited URL, the language the answer was written in, the language of each cited source, and the source type: regulator, government, media, association, directory, owned site or academic publication.
Repeat before concluding
Every prompt run more than once on separate days, because cited-source sets overlap only about 34% to 42% between consecutive days. A single run cannot distinguish a real pattern from ordinary movement.
Volatility figures from Schulte, Measuring Visibility in AI Search, April 2026. REPORTED, preprint.
What the result would change, and what it would not
Suppose the test comes back showing that French-language queries from Luxembourg overwhelmingly cite English-language sources. That would argue for putting the strongest, most citable material in English first, with French versions serving the human reader and the ranking system rather than the citation system. Suppose the opposite. Then French becomes the priority surface for the same content.
Either outcome reallocates budget. Neither outcome guarantees a citation, because the systems producing these answers belong to third parties and change their source selection between sessions. Any provider promising otherwise should be asked to put the definition of citation and the measurement method into the contract.
What to do while the question is open
Four things hold regardless of how the language test resolves, which makes them the sensible place to start.
Publish claims in a form that survives extraction. Self-contained passages that keep their meaning when lifted out of the page, a direct answer in the opening lines of each section, question-shaped headings, explicit geographic scope, and a named source with a date on every factual claim. This helps a human reader and a retrieval system for the same underlying reason.
Keep one identity across every language version. Consistent legal name, address, executive identities and service descriptions across the trade register and major business profiles, with language-appropriate descriptions and the correct inLanguage value on each variant. Small-market entities are thin in knowledge graphs, so third-party corroboration does work that on-page changes cannot.
Do not sell structured data as a citation mechanism. Google states that eligibility for its AI features rests on ordinary indexability and snippet eligibility, with no additional technical requirement, and no study establishes that adding schema causes a citation. Implement it because it clarifies entities. Say that, and nothing more.
Finally, write down what you do not know. A content plan that names its own open questions is more defensible in front of a compliance committee than one that presents a language decision as settled science. In this market, that turns out to be a commercial advantage rather than a weakness.
Frequently Asked Questions
Are Google's AI features available in Luxembourg?
Yes. Luxembourg appears on Google's official supported-country lists for both AI Overviews and AI Mode. This is worth confirming from the source because it was unclear for a long period: the March 2025 European expansion covered nine countries and Luxembourg was not one of them, France received both features only on 22 July 2026, and no Luxembourg-specific announcement was ever made. Neighbouring-country launch dates were never transferable. The supported-country list resolves it, which means content published for this market can surface inside AI-generated results today.
Is Luxembourgish supported by AI Overviews or AI Mode?
Not as at August 2026, on Google's own published supported-language lists. The comparison makes the omission notable rather than obscure: Romansh with roughly 60,000 speakers is supported, as are Faroese, Maltese, Western Frisian, Irish, Welsh, Basque and Galician, while Luxembourgish with around 400,000 speakers and national-language status in an EU member state is not. Support lists change without notice, so this needs a date attached whenever it is quoted and a periodic re-check at the source.
Does that mean a Luxembourgish version of our site is pointless?
No, it means the justification has to be honest. A Luxembourgish version can serve brand, civic, municipal and cultural purposes, and for some organisations speaking the national language is part of the identity rather than a traffic tactic. What it should not be sold as is an AI visibility play, because the language is not currently processed by Google's AI features at all. Build it with clear eyes about what it will and will not do, and revisit the position when support status changes.
Which language will an AI assistant cite for a Luxembourg query?
Nobody has published a controlled measurement, and four independent research passes for this article failed to find one. What exists is directional: peer-reviewed work at ACL 2026 documents language bias in multilingual retrieval towards English and towards the query language, a Findings of ACL 2025 benchmark found English-inclusive configurations cite English documents more often, and two vendor studies suggest English pulls citation share even from non-English prompts. All of that is experimental or commercially interested, and none of it measures Luxembourg. This is exactly why the first deliverable should be the measurement rather than a recommendation resting on an assumption.
Why does the research on Luxembourgish language models matter to marketing?
Because it explains the mechanism rather than just the outcome. Work in Findings of the ACL: EACL 2026 tested models against Luxembourgish proficiency examinations and found frontier models performing well while smaller models failed in characterised ways, including code-switching into German or French, incorrect noun gender and violations of the Eifeler Regel. Those failures are visible to a Luxembourgish reader, not merely to a benchmark. Combined with UNESCO classifying the language as vulnerable and the training corpora being small, it tells you the constraint is data scarcity rather than an arbitrary product decision, and that the position may improve.
Should we trust the vendor studies on cross-language citation?
Read them, form hypotheses from them, and do not present them as findings. One analysed more than ten million prompts and reported that roughly 78% of non-English prompts triggered at least one English fan-out, with around 43% of resulting fan-outs executed in English. Another examined about 12,000 pages across 15 languages and reported non-English content receiving roughly 34% fewer direct citations, with approximately 31% of citations still going to English sources for local-language queries, and Luxembourgish not among the languages tested. Both are single-vendor, commercially interested and not reproducible by a reader.
What can we do now, before the language question is resolved?
Four things that hold under any outcome. Structure claims so they survive extraction: self-contained passages, a direct answer early in each section, question-shaped headings, explicit geographic scope, and a named source with a date on every factual claim. Keep one consistent entity identity across all language versions, including register entries and business profiles, with the correct inLanguage value on each. Implement structured data for entity clarity while stating plainly that it does not cause citations. And document your open questions, which in a market with a compliance culture is a commercial advantage rather than an admission.
Sources & References:
- Luxembourg listed on Google's official supported-country lists for both AI Overviews and AI Mode. Luxembourgish absent from the supported-language lists for both features, while Romansh, Faroese, Maltese, Western Frisian, Irish, Welsh, Basque, Galician, Wolof and Kinyarwanda are present. Google Search support documentation, checked August 2026. Support lists change without notice.
- AI Overviews European expansion of March 2025 covered nine countries and did not include Luxembourg. Google official announcement. France received AI Overviews and AI Mode on 22 July 2026, following a publisher-compensation dispute.
- Luxembourgish language proficiency exam testing of large language models, including documented failure modes covering mistranslated words, code-switching, incorrect noun gender and Eifeler Regel violations: Lothritz, Cabot and Bernardy, Findings of the ACL: EACL 2026. Peer-reviewed.
- Supporting Luxembourgish model and benchmark work: LuxGen and LuxT5, VarDial 2025. LuxemBERT, LREC 2022, developed with a Luxembourg bank and the University of Luxembourg interdisciplinary centre. LuxIT, LexNeo-Bench and ltzGLUE benchmarks. Approximately 400,000 speakers, classified vulnerable by UNESCO.
- Language bias in multilingual retrieval-augmented generation, including reranking bias towards English and towards the query language: Wang and colleagues, ACL 2026. Peer-reviewed.
- Cross-lingual robustness benchmark finding that English-inclusive configurations cite English-language documents more frequently than configurations restricted to relevant languages: Findings of ACL 2025. Peer-reviewed.
- Query fan-out language analysis across more than ten million prompts and more than twenty million fan-outs, with approximately 78% of non-English prompts triggering at least one English fan-out and around 43% of resulting fan-outs executed in English: Peec.ai, February 2026. REPORTED, single vendor with commercial interest.
- Cross-language citation study across approximately 12,000 pages and 15 languages, reporting non-English content receiving about 34% fewer direct citations and approximately 31% of citations going to English sources for local-language queries, with Luxembourgish not tested: WhatsMyGeoScore, August 2026. REPORTED, single vendor.
- Answer volatility, with cited-source set overlap of approximately 34% to 42% between consecutive observation days over a window of about 45 days: Schulte, Measuring Visibility in AI Search, April 2026. REPORTED, preprint.
- Language statistics: STATEC. Workplace language use from the 2021 census basis, French 69.2%, Luxembourgish 54.4%, English 40.0%, German 29.5%, multiple responses permitted. Main company working language French 56%, Luxembourgish 20%, English 18%, German 6%. Luxembourgish main language 48.9% and spoken 61.2%, linguistic diversity data updated January 2026.
- Population and nationality composition: STATEC, published May 2026. 690,959 residents at 1 January 2026, foreign nationals 322,050 or 46.6%, across 180 nationalities. Cross-border employment from STATEC Regards 02/26.
- AI feature eligibility depends on ordinary indexability and snippet eligibility with no additional technical requirement: Google Search Central documentation. No study establishes that structured data causes AI citation.
- No controlled public measurement exists of assistant answer language or cited-source language for queries originating in Luxembourg, and no sector survey of assistant usage by Luxembourg finance professionals was found. Both stated as unavailable rather than estimated.
- This article is search and AI visibility analysis, not legal advice.