Why a Swiss German Prompt Returns a German Answer
SEO

Why a Swiss German Prompt Returns a German Answer

No written dialect standard means retrieval pools Switzerland with a neighbour ten times larger. The mechanism is documented, the magnitude is not.

Ask an assistant, in German, which suppliers it would recommend for a business in Switzerland. Read the sources it cites. In a meaningful share of cases, several of them will be German rather than Swiss, the pricing context will be European rather than Swiss, and the legal framing will point at German statutes that do not apply in Bern.

This is not a bug that anyone has fixed, and it is not a conspiracy either. It follows from three structural facts about Switzerland that stack on top of each other, and the mechanism behind it has been documented in academic work. What has not been documented is how often it happens. That distinction runs through this whole article, and any agency that blurs it is selling you a number it does not have.

What follows is the mechanism, the academic evidence for it, why Switzerland is unusually exposed, what remains unmeasured, and how to test it for your own category rather than accepting an industry claim.

Three facts that stack

Take them in order, because the third one only matters because of the first two.

Swiss German has no standardised written form. The Alemannic dialects used in everyday speech across German-speaking Switzerland have no agreed orthography, no dictionary of record, no version taught as correct writing. So when a Swiss company publishes anything, it publishes in Swiss Standard German.

Swiss Standard German is close to Germany-German. Close enough that a retrieval system treats them as one language. They differ in real ways, and those differences matter commercially, as covered separately in this series. But a language identifier looks at a Swiss page and a German page and returns the same answer: this is German.

The German-language web is roughly ten times larger. Germany plus Austria plus the German-language corpus accumulated over decades of publishing dwarfs what eight and a half million Swiss residents have produced, and a large share of Swiss output is in French or Italian rather than German.

Put those together. A Swiss user asks a question in German. The system resolves the language as German. It then searches a German-language pool in which Swiss material is a minority by a wide margin. Nothing malicious happens at any step. The outcome is still that Swiss sources compete against a much larger neighbour on a playing field that does not know the border exists.

What the academic evidence actually says

There is one study worth citing here, and it is worth citing precisely rather than loosely.

Chirkova and colleagues published work in October 2024 examining multilingual retrieval-augmented generation across 49 languages. The finding relevant to Switzerland is that models systematically prefer information written in the same language as the query, and that when same-language material is thin they fall back on higher-resource languages rather than on lower-resource ones. That is a general property of multilingual retrieval, established across a large language sample, not a claim about Switzerland.

Apply it to the three stacked facts and you get a prediction rather than a measurement. Same-language preference means a German-language query pulls German-language sources. Since Swiss Standard German is classified as German, Swiss sources are inside that pool but heavily outnumbered. And the higher-resource fallback tendency points in the same direction, toward the larger German corpus.

Related work from September 2025 examined cross-language stability in retrieval and found that the same question asked in different languages does not return equivalent source sets. That reinforces the point without quantifying the Swiss case either.

So the mechanism has named academic support. The Swiss magnitude does not. All four of our independent research passes reported the same thing: no published study measures how often German-domain sources displace Swiss ones in AI answers about Swiss topics. One of those passes described the phenomenon at length and then admitted, in its own words, that it had found no study measuring the percentage. We treat that admission as the honest part of the document and the description as a hypothesis.

Mechanism, not magnitude

How a Swiss Question Ends Up With German Sources

Four steps, none of them a malfunction, and a result that still disadvantages Swiss publishers.

A Swiss user types the question

Not in dialect, because dialect has no written standard. In Swiss Standard German, which is what Swiss people actually write and type.

The system resolves the language as German

Correctly, by any technical standard. Swiss Standard German is German. The Swiss variety markers, including ss for ß and Helvetisms, do not create a separate language classification.

Retrieval draws from the German-language pool

Chirkova and colleagues, across 49 languages in October 2024, found models systematically prefer same-language material and fall back on higher-resource languages when it is thin. Both tendencies point at the same pool.

Swiss material is a minority inside that pool

A German-language web roughly ten times larger, accumulated over decades, against the German-language output of eight and a half million residents, a large share of whom publish in French or Italian.

And here is what is not known

No published study measures how often this displacement actually occurs, for any Swiss category, on any engine. Four independent research passes returned the same answer. The mechanism has named academic support. The Swiss magnitude does not exist as a figure, and constructing one would be inventing evidence rather than reporting it.

Sources: Chirkova et al., multilingual retrieval-augmented generation across 49 languages, October 2024, arXiv • Related cross-language retrieval stability work, September 2025 • Swiss German written-standard status cross-validated across four independent research passes
Created by Arfadia • arfadia.com/blog

Why Switzerland is more exposed than its neighbours

Austria shares the same broad problem and carries it more lightly. Austrian Standard German is also a national variety inside a German-dominated pool, and Austria is also much smaller than Germany. The difference is that Austria has one national language, so all of its output goes into that one pool.

Switzerland splits its production four ways. German-language Swiss output is not the full national output, it is roughly the portion produced by the German-speaking regions. French-language Swiss material sits in a pool dominated by France. Italian-language Swiss material sits in a pool dominated by Italy. Romansh has almost no pool at all.

So Switzerland is not one small country competing against one large neighbour. It is three small language regions each competing against a large neighbour, simultaneously, with a fourth that barely registers. That structural position is the reason this article exists as a Swiss article rather than a general one.

It is also, not coincidentally, why Apertus was built. The Swiss federal institutions behind it stated the rationale plainly: a country with four national languages found itself dependent on models that were not designed for its linguistic diversity. Apertus explicitly includes Swiss German and Romansh in its training data. That is the same structural problem approached from the model side rather than the content side, and it is a useful confirmation that the problem is real enough for two federal technical institutes to spend supercomputer time on it.

What people get wrong when they try to fix this

Three responses are common and only one of them works.

Translating harder. Producing more German-language content does not address the classification step. If the material still reads as generically German, with no Swiss anchoring, it enters the same oversubscribed pool as everything else and competes on volume against a much larger corpus. Volume is the one competition Switzerland cannot win.

Buying a country signal and stopping there. A .ch domain, a Swiss address in the footer, a CHF price list. These help, and they are necessary, and they are not sufficient on their own. A retrieval system working from a passage of text is looking at the passage, not at the domain registrar record.

Anchoring the content itself. This is the approach that engages the actual mechanism. If same-language preference is the driver, then the useful move is making the Swiss identity of the material legible inside the text that gets retrieved. Canton and city named explicitly. CHF as the currency in every commercial expression. Swiss statutes cited by article number rather than described generically. Swiss regulators named. Swiss orthography with ss throughout. Helvetisms used as primary vocabulary rather than as parenthetical alternatives.

The honest framing for that third approach matters. It is a reasoned response to a documented mechanism, not a technique with a measured effect size. Nobody has run the controlled study. We present it as a hypothesis worth testing on your own tracked prompts, and we design the measurement so you can see whether it worked for you.

How to test it for your own category

Here is the part that turns an unmeasurable industry claim into a measurable client-specific number.

Build a tracked prompt set for your category, in Swiss Standard German, phrased the way a Swiss buyer would actually ask. Run each prompt repeatedly, across the engines that matter in Switzerland, and record every cited source. Then classify each cited source by country of origin. That gives you a Swiss-versus-foreign source ratio for your own category, on named engines, on stated dates, with a stated run count.

Do the same in French, separately, with the France comparison in mind rather than Germany. Do it in Italian if Ticino matters to you.

Now you have a baseline that is real. Apply Swiss source anchoring to your content, wait for a reasonable indexing and retrieval window, and re-run the same prompt set the same number of times. The change in the source-origin ratio is your answer, for your category, and it is worth more than any market-wide percentage would be even if one existed.

What you can claim Evidence status How to say it honestly
Models prefer same-language sourcesVerified, named study across 49 languagesState it with the study, the language count and the date. It is a general property, not a Swiss finding
Swiss German has no written standardVerified across four research passesState it flatly. This one is not contested by anybody
German sources appear in Swiss answersObservable per prompt, per engine, per dateShow the captured run with a date and engine named. One observation is an observation, not a rate
How often it happens across SwitzerlandUnmeasured, four passes returned nothingSay it is unmeasured. Offer a per-category baseline instead of a market figure
That Swiss anchoring fixes itReasoned from mechanism, effect size unmeasuredPresent as a hypothesis with a designed test, and report the result whichever way it goes
A specific percentage improvementNo basis whatsoeverDo not make this claim. It cannot be supported from any source we located
Inside the retrieved text

Six Signals That Make Swiss Content Legibly Swiss

The domain and the footer address are outside the passage. These are inside it.

Canton and city, named

Zürich, Geneva, Basel, Bern, Lausanne, Lugano and the relevant canton, written into the body text rather than left to the contact page.

CHF in every commercial expression

Not EUR, not converted, not omitted. Currency is one of the fastest signals that a passage is about a different market.

Swiss statutes by article number

The Code of Obligations, the revised FADP, FinSA where relevant. Article numbers are specific, checkable and unmistakably Swiss.

Swiss regulators named

FINMA, Swissmedic, the federal data protection commissioner. Naming the regulator locates the content in a jurisdiction.

ss throughout, never ß

Swiss orthography in the body copy, headings, alt text and structured data, not only in the visible text a human skims.

Helvetisms as primary terms

Offerte, Spital, Natel, Velo, parkieren used as the head vocabulary rather than in brackets after the German equivalent.

And measure it rather than assuming it

Build the tracked prompt set, record the country of origin of every cited source across repeated runs, apply the anchoring, then re-run the same set the same number of times. The change in the source-origin ratio is your evidence. If it does not move, that is also a result, and it belongs in the report.

Approach reasoned from Chirkova et al., October 2024, on same-language retrieval preference. Effect size for Swiss source anchoring is unmeasured and presented as a testable hypothesis.
Created by Arfadia • arfadia.com/blog

The part that is uncomfortable to say out loud

An agency that promises to fix cross-border source displacement in Switzerland is promising something nobody has measured. That includes us, which is why this article describes a test rather than a result.

The reason to say it anyway is that the alternative is worse. Across four research passes we watched one model describe this phenomenon in confident detail, give it a name, and then concede in its own limitations section that it had found no study measuring the magnitude. That pattern, a vivid mechanism dressed as a finding, is how invented statistics enter client decks. A buyer cannot easily tell the difference between a mechanism and a measurement when both are presented in the same tone of voice.

So the useful commitment is narrower and more checkable. We will name the mechanism and its source. We will not attach a Swiss percentage to it. We will build you a baseline you can verify yourself, and we will report the re-test whichever direction it goes.

Tessar Napitupulu covers how retrieval systems select sources, and how to measure AI visibility without manufacturing evidence, in Cited or Silent, available as a free gated edition, with retailer editions on Amazon, Google Play and Apple Books.


Frequently Asked Questions


Why would an AI answer about a Swiss company cite a German source?

Because three structural facts stack. Swiss German has no standardised written form, so Swiss companies publish in Swiss Standard German. Swiss Standard German is classified as German by any language identifier, so Swiss and German material sit in one retrieval pool. That pool is dominated by a German-language web roughly ten times larger, accumulated over decades, while a substantial share of Swiss output is in French or Italian rather than German. Chirkova and colleagues, across 49 languages in October 2024, found models systematically prefer same-language material and fall back on higher-resource languages when it is thin. Both tendencies point toward the larger German corpus.


Is there a measurement of how often this happens?

No, and this is where most claims in this area fall apart. Four independent research passes found no published study measuring how often German-domain sources displace Swiss ones in AI answers about Swiss topics, for any category, on any engine. The mechanism has named academic support. The Swiss magnitude does not exist as a figure. What can be measured is your own category: build a tracked prompt set in Swiss Standard German, run it repeatedly across the relevant engines, and classify the country of origin of every cited source. That produces a real baseline with a stated run count, engines and dates.


Does a .ch domain solve the problem?

It helps and it is not sufficient alone. A country-code domain, a Swiss address and CHF pricing are useful signals, but a retrieval system working from a passage of text is evaluating the passage rather than the registrar record. The signals that operate inside the retrieved text are different: canton and city named in the body copy, CHF in every commercial expression, Swiss statutes cited by article number, Swiss regulators named, ss rather than the Eszett throughout including in alt text and structured data, and Helvetisms used as primary vocabulary rather than as bracketed alternatives.


Is Switzerland worse affected than Austria?

Structurally yes, because Austria has one national language and Switzerland splits its output four ways. Austrian Standard German competes inside a German-dominated pool, which is the same broad problem, but all Austrian output goes into that single pool. Swiss German-language material is only the portion produced by the German-speaking regions. French-language Swiss material competes in a pool dominated by France, Italian-language Swiss material in a pool dominated by Italy, and Romansh has almost no pool at all. Switzerland is therefore three small language regions each facing a large neighbour at the same time.


Does Apertus fix this?

Not directly, though its existence confirms the problem is taken seriously. The Swiss federal institutions behind Apertus stated the rationale explicitly: a country with four national languages was dependent on models not designed for its linguistic diversity. Apertus includes Swiss German and Romansh in its training data. That addresses the issue from the model side rather than the content side. It is a foundation model distributed through Swisscom, Hugging Face and the Public AI network rather than a consumer answer engine with a citation surface a brand can appear in, so it does not change where your content gets cited today.


Should we publish in Swiss German dialect to differentiate?

No, because there is no stable written form to publish in. Swiss German has no agreed orthography and no version taught as correct writing, so dialect text would be idiosyncratic, inconsistent between authors, and impossible for a Swiss user to match by typing. The differentiation available is Swiss Standard German done properly, with ss throughout, Swiss institutional vocabulary and Helvetisms as head terms. That is a written variety that Swiss readers recognise and Swiss users type, which is the combination that matters.


How long before source anchoring shows an effect?

Unknown, and we will not invent a timeline for something whose effect size is unmeasured. The dependencies are real: content has to be published, crawled and indexed, retrieval indexes have to refresh, and assistant outputs vary between runs regardless. What we can commit to is measurement design rather than a schedule. Baseline the tracked prompt set with a stated run count before changes, allow a defined window, then re-run the identical set the same number of times and compare source-origin ratios. If the ratio has not moved, the report says so.

Sources & References:

  • Chirkova et al., multilingual retrieval-augmented generation study across 49 languages, October 2024, arXiv. Finding: models systematically prefer information in the same language as the query, and favour higher-resource languages when same-language information is absent or thin. General property established across a large language sample, not a Switzerland-specific finding.
  • Related work on cross-language retrieval stability, September 2025, arXiv. Finding: the same question asked in different languages does not return equivalent source sets.
  • Swiss German has no standardised written form; Swiss publishing occurs in Swiss Standard German. Cross-validated across four independent research passes.
  • Apertus: released 2 September 2025 by EPFL, ETH Zürich and the Swiss National Supercomputing Centre. 8-billion and 70-billion parameter sizes, 15 trillion training tokens, more than 1,000 languages, approximately 40% non-English training data, explicitly including Swiss German and Romansh. Trained on the Alps supercomputer in Lugano. Distributed via Swisscom, Hugging Face and the Public AI network. Institutional sources state more than 1,000 languages; a figure of 1,811 appearing in some secondary coverage is not used here.
  • Reported as unavailable after four independent research passes: any published measurement of how often German-domain sources displace Swiss ones in AI answers about Swiss topics; any measured effect size for Swiss source anchoring on citation outcomes; any Swiss AI Overview trigger rate.
  • Noted during cross-validation: one research pass described cross-border source displacement in detail, named the phenomenon, and then stated in its own limitations that it had found no study measuring the magnitude. The description is treated here as a hypothesis and the admission as the reliable part of the document.
  • Federal Statistical Office language data and Swiss Standard German orthography conventions are covered in the companion article on Swiss keyword research in this series.
  • This article is orientation on retrieval behaviour and measurement design, not legal advice. Swiss legal, tax and regulatory questions should be reviewed by qualified Swiss advisers.
0 Comments 0 Comments
0 Comments 0 Comments