Here is a number: forty per cent Arabic AI visibility. It appears on a slide, the client nods, the retainer renews.
Now here are the questions that number cannot survive. Across how many prompts? Chosen by whom? In which register of Arabic? Tested on which engines, on which dates, across how many runs? Does visibility mean the brand was named, or that a page was cited as a source? And when the prompt was Arabic, was the cited document Arabic or English?
Most Arabic visibility figures in circulation collapse under the third question. Almost all of them collapse under the last one, which is a shame, because the last one is the only question whose answer tells you what to do next.
This is about how to build a measurement you can defend, in a market where no independent benchmark exists to check anyone against. Including us.
Three things that are not the same thing
Start here, because conflating these is how honest people produce misleading reports.
Ranking is where a page sits in a list of links. It is measured against a keyword and a location. It has been the industry's currency for two decades.
Citation is whether an answer engine used your page as a source when it generated a response, and attributed it. This is a different event with a different mechanism. A page can rank first and never be cited. A page sitting deep in the results can be cited repeatedly. Retrieval for generation does not read a ranked list top to bottom, so rank position cannot be used as evidence of citation performance.
Mention is whether your brand name appeared in the generated text. That is not citation. A model can name a company from parametric memory without attributing any source at all, which means no link, no referral, and no page of yours involved.
The commercial difference matters. A mention builds awareness in a channel you cannot measure through your analytics. A citation potentially sends traffic and definitely signals that your content is the material the system reached for. Reports that add them together produce a bigger number and a less useful one, and the direction of the error is always flattering, which is why it happens.
One more distinction that gets skipped. Being cited is not the same as being recommended. An engine can cite your page while framing your category as risky, or while naming a competitor as the better option and using your page to support a factual point. Citation count on its own does not tell you which of those happened.
Four paths, and why one number destroys them
In a bilingual market this is the measurement decision that matters most, and it is the one almost nobody makes.
A question and its answer each have a language. That gives four combinations, and they behave differently enough that averaging them is not simplification, it is deletion.
An Arabic question answered from an Arabic source. An Arabic question answered from an English source. An English question answered from an English source. An English question answered from an Arabic source.
The reason to separate them is measured rather than theoretical. Work presented at ACL ArabicNLP in 2025 examining retrieval over mixed Arabic and English corpora found substantial losses when the language of the query and the language of the supporting document diverged, ranging from roughly thirteen to forty-two per cent depending on the embedding model and the domain tested, with legal-domain material at the worse end and travel-domain material at the better end.
So consider what a blended figure does to that. Suppose you test a hundred prompts, half Arabic and half English, and report forty per cent visibility. If most of your wins came from the English-question, English-source path, the number is describing your English performance while wearing a bilingual label. Your Arabic gap is invisible, and it is invisible in precisely the place where a fix would have the largest effect.
Now report the same test as four numbers. Suddenly you can see that Arabic questions are being answered from English documents, which tells you the specific thing to do: publish that answer in Arabic. That is a diagnostic. Forty per cent is a mood.
The fourth path deserves separate mention because it is the one people forget entirely. In Qatar a large share of the professional audience queries in English about topics documented locally, sometimes in Arabic. English question, Arabic source. It carries the same mismatch penalty in the opposite direction, and nobody measures it because nobody thinks of it.
Five Measurements Worth Reporting
The test for any citation metric is whether a bad result tells you what to change. If it does not, it is a score rather than a measurement.
Cited-answer rate, per engine and per language path
The share of panel prompts where an engine names the brand and attributes a source. Never blended across engines, never blended across the four paths.
Answers: where are we actually being used as a sourceSource-language split on Arabic prompts
For every citation earned on an Arabic question, whether the cited document was Arabic or English. The single most actionable number in a bilingual programme.
Answers: what do we publish next, and in which languageEntity accuracy in both languages
Whether the engine describes the business correctly: name, services, location, scope. Checked separately in Arabic and English, because they diverge.
Answers: is being visible currently helping or hurtingFraming when named
Neutral, positive, or hedged. A cautious mention is a different commercial outcome from a recommendation, and citation counts cannot distinguish them.
Answers: what is the buyer actually reading about usReproducibility data
Prompt, engine, language, date, run number and raw answer text, exported so the client's own team can re-run the test without the agency present.
Answers: can we verify any of this ourselvesReproducibility is the metric that makes the other four trustworthy. An agency unwilling to hand over the raw answers is asking to be believed rather than checked.
The prompt panel is the instrument
Everything above depends on one artefact, and it has to be built before any baseline is taken.
A prompt panel is a fixed, written set of buyer questions. Fixed is the operative word. If the questions change between reports, the numbers are not comparable, and a vendor who quietly swaps in easier prompts can show improvement without anything improving.
Four properties make a panel defensible.
Written down and agreed before the baseline. Argue about the panel at the start. It is a cheap argument then and an impossible one later.
Covering three language variants. Modern Standard Arabic in the register your category actually uses, Gulf dialect where buyers speak conversationally, and English for the substantial non-Arabic-speaking commercial audience. These are three sets of questions, not one set translated twice.
Weighted toward commercial intent, not brand. A panel dominated by questions containing your own name will produce impressive numbers that mean nothing, because a model asked about your brand will discuss your brand. The useful questions are the ones a buyer asks before they know you exist.
Sized honestly. Twenty prompts is a probe. A few hundred starts to be a measurement. Whatever the number is, it goes in the report, because a percentage without a denominator is a decoration.
On dialect, one caution worth building in. The main published Arabic dialect benchmark covers five dialects and Qatari is not among them, which means machine handling of Qatari dialect specifically has not been systematically measured by anyone. Treat dialectal prompts as a genuine test rather than an assumed-working channel, and watch what the results tell you.
Runs, dates, and the thing about determinism
Answer engines are not calculators. Ask the same question twice and you can get different answers, different sources, and your brand present in one and absent from the other.
Which means a single-day snapshot is not a result. It is one sample of a distribution.
Three practices follow. Run each prompt multiple times rather than once. Use isolated sessions so that earlier queries in the same session do not shape later answers, which is a common and rarely disclosed contamination in vendor testing. And record the date range, because a measurement from one week and a measurement from three months later are describing different systems.
That last point is not pedantry. Google brought its Arabic answer features to the region and to Arabic across 2025, adding Modern Standard Arabic to its more advanced answer mode in October 2025 as part of an expansion covering thirty-eight new languages. A baseline taken before that and a measurement taken after it are not comparable, and any report that spans the change without noting it is describing two different worlds as one trend.
There is a related detail worth knowing about how these systems work, because it changes what a prompt panel should contain. Google has described its answer mode using a technique that runs multiple background searches from a single user question and aggregates the results. It also reports that in markets where the feature is live, users submit queries two to three times longer than traditional search inputs. So panel prompts written like keywords are testing the wrong behaviour. They should be written like the long, conversational questions buyers actually type.
What the evidence says actually moves citation
Measurement is only useful if something can be done with the result, so it is worth being specific about what the research supports.
The peer-reviewed study on generative engine optimisation presented at ACM SIGKDD in 2024 tested content-side changes against citation outcomes directly. The findings were clear in direction and in ranking. Named expert attribution produced the largest single lift, just under forty-one per cent. Statistics paired with a named source produced a substantial lift of around thirty-one per cent. Inline citations produced a lift of around twenty-eight per cent. Keyword stuffing reduced citation rates by roughly eight per cent.
Read the list again and notice what is not on it. No technical trick. No markup lever. No volume play. Every positive item is an editorial decision about credibility: quote a named human, source your numbers, cite your references.
Arabic-specific remediation work presented at ICNLSP in 2025 pointed the same way from a different angle, reporting improvements of up to eight points in precision at five, around eleven per cent in answering accuracy, and correct citation rates reaching as high as ninety-five per cent under the configurations tested. Correct citation in Arabic responds to structure and pipeline design rather than being a fixed ceiling.
And there is a retrieval finding that explains why structure matters so much. Work on Arabic lexical retrieval presented at ABJADNLP in 2025 reported recall in the top five results of 0.88 alongside a mean reciprocal rank of 0.48. The right document usually gets into the candidate set. It frequently is not at the top of it. Being findable is largely solved. Being the passage the system reaches for first is not, and that is decided at the level of the individual extractable unit.
Practical consequence for content: build in self-contained units. One claim per section, the answer in the opening sentences, a named source attached to every figure, and a heading shaped like the question it answers. Sections need to survive being lifted out of the page, because that is what happens to them. We cover the retrieval mechanics behind this in our generative engine optimisation practice for Qatar.
What to refuse to report, and why
A measurement practice is defined as much by its refusals as by its metrics.
| Commonly reported | Why it should not be | What to report instead |
|---|---|---|
| One blended Arabic visibility percentage | Averages four language paths with measured performance differences of up to forty-two per cent between them, deleting the only diagnostic that indicates what to publish next. | Four separate rates, one per language path, per engine. |
| Mentions counted as citations | A model can name a brand from memory with no source attributed, which produces no referral and involves none of your content. Combining the two always inflates in the flattering direction. | Attributed citations and unattributed mentions as separate lines. |
| Ranking data as citation evidence | Retrieval for generation does not read a ranked list in order. A first-place page may never be cited and a deeply ranked page may be cited repeatedly. | Citation measured directly, with ranking reported separately as its own outcome. |
| Industry-wide prevalence figures | Published trigger-rate figures for AI answers come from English and largely United States keyword sets, vary widely by provider and definition, and are not disaggregated by language. | Your own panel results, with the panel disclosed. |
| A single snapshot as a trend | Outputs vary between runs of the same prompt, so one sample of a distribution is not a position. | Multiple runs, isolated sessions, date range stated. |
| Only the prompts that performed | The failed prompts are the working document. Removing them turns a measurement into a brochure and removes the reason to run it. | The full panel, wins and misses together. |
That last row is the one to press hardest on when evaluating any agency, ours included. Ask to see the prompts where the brand did not appear. A vendor measuring something will hand them over, because the misses are where the next quarter's work comes from. A vendor selling something will change the subject.
Six Questions and What the Answers Tell You
None of these require technical knowledge to ask. All of them are difficult to answer if there is no method underneath.
Can I see the full prompt panel?
Good sign: a written document, agreed in advance, with a stated count and language breakdown.
Warning sign: the panel is described rather than shown, or it changes between reports.
How many of those prompts contain our brand name?
Good sign: a minority, with the rest written as questions a buyer asks before knowing you exist.
Warning sign: most of them, which guarantees a flattering result that means nothing.
When the prompt was Arabic, was the cited document Arabic or English?
Good sign: a straight answer with a split, because they tracked it.
Warning sign: confusion about why it matters, which usually means one blended figure.
How many runs, and were sessions isolated?
Good sign: multiple runs, isolated sessions, date range disclosed.
Warning sign: a single test date presented as a position.
How exactly is this percentage calculated?
Good sign: a formula, and an acknowledgement that other tools define it differently.
Warning sign: a metric name treated as if it were an industry standard.
Can I see the prompts where we did not appear?
Good sign: handed over without hesitation, because that is where the next work comes from.
Warning sign: any answer other than yes.
Different tools compute share-of-voice style metrics differently, so two vendors can report different numbers about the same brand in the same week without either being dishonest. That is exactly why the calculation method belongs in the report.
The gap that makes all of this necessary
One structural fact sits under this entire article, and it is the reason methodology has to substitute for benchmarking.
No independent study measures how often answer engines cite Arabic sources for Arabic questions in the Qatari market. None measures Arabic against English answer quality inside Qatar. That was checked across the research assembled for this work and the gap remains open. Separately, no published measurement compares the trigger rate of AI answers for Arabic queries against English ones in any geography, which means the prevalence figures in wide circulation are all drawn from English and largely United States keyword sets and should never be restated as Arabic numbers.
The temptation is to treat that as a marketing problem and fill it with something. Resist it. The absence has a practical consequence that is more useful than a borrowed statistic: it means every Arabic visibility figure you will ever be shown is a vendor measurement, and the only way to evaluate one is to inspect its method.
Which is a sentence that applies to us as much as to anyone else, and is the reason this article is a list of questions rather than a list of our results. For the wider technical and regulatory context around this work, our team has documented the full picture for this market.
The short version
Separate ranking from citation from mention, because they are three different events and merging them always flatters. Measure four language paths rather than one blended number, because the mismatch between question language and source language carries a measured cost of up to forty-two per cent and the split is the only thing that tells you what to publish next. Build a fixed prompt panel, agreed in writing before the baseline, weighted toward buyer questions rather than brand questions, with its count and language breakdown stated in every report.
Run each prompt multiple times in isolated sessions and disclose the date range, because outputs vary and platforms change. Act on the evidence that is actually published: named expert attribution, sourced statistics and inline citations all measurably raised citation rates in peer-reviewed testing, and keyword stuffing lowered them. Report entity accuracy and framing alongside citation rate, because a confidently wrong description is worse than an absence.
And keep the failed prompts in the report. They are the reason to run the measurement at all.
Frequently asked questions
What is the difference between a citation and a mention, and does it matter commercially?
A citation means an answer engine used your page as a source and attributed it. A mention means your brand name appeared in the generated text, which a model can do from memory with no source attributed at all. The commercial difference is real: a citation potentially sends a referral and confirms that your content was the material the system reached for, while a mention builds awareness in a channel your analytics cannot see. Reports that combine them produce a larger number and a less useful one, and the error always runs in the flattering direction, which is why the practice persists. A third distinction gets skipped too: being cited is not being recommended, since an engine can cite your page while naming a competitor as the better option.
Why can rankings not be used as evidence of AI citation performance?
Because retrieval for generation does not work by reading a ranked list from the top. A page can hold the first organic position and never be cited in a generated answer, and a page sitting well down the results can be cited repeatedly. They are two different selection mechanisms measuring two different things. Reporting rankings as evidence of citation skips a step, and the step it skips is the entire question being asked. Rankings remain a legitimate outcome to measure and report, just as their own line rather than as a proxy.
Why insist on four language paths instead of one Arabic visibility figure?
Because the four behave differently by a margin large enough that averaging them removes the diagnostic. Work presented at ACL ArabicNLP in 2025 measured losses of roughly thirteen to forty-two per cent when the language of the query and the language of the supporting document diverged, depending on embedding model and domain. A blended figure can therefore be carried almost entirely by English-question, English-source performance while wearing a bilingual label, leaving the Arabic gap invisible in exactly the place where fixing it would matter most. Reported as four numbers, the same test tells you whether Arabic questions are being answered from English documents, which is a specific instruction to publish that answer in Arabic.
How large should a prompt panel be, and who should choose the prompts?
Large enough that the percentage has a meaningful denominator, and chosen jointly, in writing, before any baseline is taken. Twenty prompts is a probe rather than a measurement. A few hundred starts to be defensible. Whatever the number, it belongs in every report, because a percentage without a denominator is decoration. Composition matters as much as size: a panel dominated by prompts containing your own brand name will produce impressive results that mean nothing, since a model asked about your brand will discuss your brand. The valuable prompts are the ones a buyer asks before they know you exist.
How many times should each prompt be run?
More than once, in isolated sessions, with the date range recorded. Answer engines are not deterministic, so the same prompt can return different answers and different sources on different runs, which makes a single-day snapshot one sample of a distribution rather than a position. Session isolation matters because earlier queries in a session can shape later answers, and that contamination is common in vendor testing and rarely disclosed. Date range matters because the platforms themselves move: Arabic support in Google's more advanced answer mode arrived in October 2025 as part of a thirty-eight language expansion, so measurements taken either side of that are describing different systems.
What content changes are actually proven to increase citation rates?
The peer-reviewed study on generative engine optimisation presented at ACM SIGKDD in 2024 tested content changes against citation outcomes and found named expert attribution produced the largest single lift at just under forty-one per cent, statistics paired with a named source produced around thirty-one per cent, inline citations around twenty-eight per cent, and keyword stuffing reduced citation rates by roughly eight per cent. Every positive item on that list is an editorial decision about credibility rather than a technical trick. Arabic-specific work at ICNLSP in 2025 pointed the same direction, reporting gains of up to eight points in precision at five, around eleven per cent in answering accuracy, and correct citation rates as high as ninety-five per cent.
Sources & References
Aggarwal et al., GEO: Generative Engine Optimization, ACM SIGKDD 2024, for the measured citation lifts from named expert attribution, sourced statistics and inline citations, and for the measured reduction from keyword stuffing.
Amiraz et al., The Cross-Lingual Cost: Retrieval Biases in RAG over Arabic-English Corpora, ACL ArabicNLP 2025, for the losses measured when query language and document language diverge. Al-Rasheed et al., ABJADNLP 2025, for Arabic lexical retrieval recall and mean reciprocal rank. Enhancing Arabic retrieval-augmented generation, ICNLSP 2025, for the measured improvements in precision, answering accuracy and correct citation rate.
Google MENA product announcements for the arrival of Arabic answer features in the region across 2025 and the addition of Modern Standard Arabic to its advanced answer mode in October 2025 across thirty-eight new languages, and for the described query expansion technique and the reported increase in query length. DialectalArabicMMLU, 2025, for the five-dialect coverage that does not include Qatari dialect.
Published AI answer trigger-rate figures from commercial measurement providers were reviewed during research. They are drawn from English and largely United States keyword sets, differ substantially between providers by definition and measurement window, and are not disaggregated by language, so they are not restated here as Arabic or Qatari figures.