Most Arabic keyword lists for the Qatari market were never researched. They were translated. Someone built the English list first, ran it through a translation step, handed it to a writer, and the writer produced pages targeting phrases that are grammatically correct and commercially useless. The terms read fine. Nobody types them.
That is the single most expensive mistake in this market, and it is invisible from the inside. A translated keyword list does not throw an error. It produces content, the content gets published, and eight months later the traffic still is not there while everyone argues about backlinks.
What follows is how the work actually has to be done: why Arabic splits into two registers that behave differently in search, why the same word arrives at a search engine in three or four spellings, what one recent benchmark reveals about Qatari dialect specifically, and where an honest practitioner has to say the data does not exist.
Arabic is two languages wearing one name
Arabic is diglossic. That is not a stylistic observation, it is a structural fact with direct consequences for keyword targeting, and it is where most strategies go wrong on day one.
Modern Standard Arabic, often called Fus'ha, is the written register. It carries formal, institutional, legal, procurement and business-to-business language across the entire Arab world. It is the register Qatari ministries write in, the register tender documents use, the register a board paper appears in. Nobody grows up speaking it at home.
Gulf Arabic is what people actually speak. Linguistically it is a defined dialect group, not slang and not a corruption of the standard language. The academic corpus work covering it spans Bahrain, Qatar, the United Arab Emirates, Kuwait and the eastern part of Saudi Arabia, which tells you something useful straight away: the register that matters commercially in Doha is shared across a specific set of neighbours and is not interchangeable with Egyptian or Levantine Arabic.
Here is where it bites. Search intent splits along the register line, and it splits predictably.
A procurement officer researching facility management contractors writes in Modern Standard Arabic, because that is the vocabulary the tender itself will use and because the query is professional. A resident looking for a dermatology clinic on a Thursday evening types the way they speak. Same country, same search engine, two different vocabularies. Target one and you lose the other.
The frequently cited example among Gulf practitioners is the word for car, which diverges sharply between Egyptian and Gulf usage. Vocabulary divergence of that kind is not an edge case. It runs through everyday commercial categories, and a keyword that performs in Cairo can be the wrong word entirely in Doha.
Four Decisions That Have To Be Made First
Each one is made once, at the keyword stage. Getting any of them wrong is expensive to undo after the content exists.
Which register for this page
Modern Standard Arabic for institutional, legal, procurement and business-to-business intent. Gulf dialect for conversational, consumer and local intent. Decided per page, per intent, never once for the whole site.
Which spellings count as targets
Hamza omission on alef, yeh against alef maksura, teh marbuta against heh. Each variant is a separate string arriving at a search engine, not a typo to be corrected in the copy.
Which terminology the market uses
Industry vocabulary checked against how Qatari commercial and tender documents phrase it, not against a dictionary equivalent. A literal translation of an industry term is frequently not the industry term.
Which claims you refuse to make
Some things are genuinely unmeasured for Qatar, including the share of searches using Latin-script Arabic. Building a workstream on an unmeasured assumption is how budget disappears quietly.
Register split follows the documented diglossia of Arabic. Orthographic variants follow standard Arabic informal writing patterns. The gaps are stated as gaps rather than filled with estimates.
The spelling problem nobody puts in the budget
Written Arabic tolerates variation that written English does not, and Qatari searchers use that tolerance constantly. Three patterns matter enough to plan around.
The hamza gets dropped. The glottal marker that sits on an alef is routinely omitted in informal typing. The word is still understood by a human reader. To a search engine it is a different string.
Yeh and alef maksura get swapped. Two characters that look similar and sound similar in final position. Users pick whichever their keyboard habit produces, and both appear in real query logs.
Teh marbuta and heh alternate. The feminine ending is written one way in careful text and another way in fast typing, and both reach the search box.
Now put those together. A three-word Arabic commercial phrase where each word carries one possible variant can arrive at a search engine in several distinct forms. Your keyword set contains one of them. The textbook one.
This is not exotic linguistics. It is the Arabic equivalent of ignoring the difference between organise and organize, except that it stacks multiplicatively across a phrase rather than appearing once. And unlike the English case, most keyword tools will not normalise it for you helpfully. They will report volume for the string you asked about, and stay silent about the three you did not.
The practical fix is unglamorous. Enumerate the variants per target term at the research stage, check each one against live results rather than trusting a single tool figure, and decide which get their own targeting and which get folded into on-page coverage. It takes hours. It is also the difference between a keyword set that reflects the market and one that reflects a dictionary.
What one benchmark tells us about Qatari dialect, by leaving it out
Here is a finding worth sitting with. In 2025 a dialect-focused Arabic benchmark was published covering five Arabic dialects, with three thousand question-and-answer pairs per dialect, fifteen thousand items in total, spanning thirty-two domains. Serious work.
Qatari dialect is not one of the five.
That absence is more useful than a number would have been. It tells you that when a language model, an answer engine or a keyword tool handles Qatari dialect, its performance on that specific variety has not been systematically measured by anyone. Not badly measured. Not measured.
Two consequences follow directly. First, treat any vendor claim about dialect performance in Qatar as untested rather than false, and ask what it was tested against. Second, when dialect targeting matters for a category, validate it inside that category with real queries rather than importing a general assumption about Gulf Arabic.
It also has a bearing on where AI search is going, which we will come back to.
Latin-script Arabic, and a number we will not quote you
Arabizi, sometimes called Franco-Arabic or Arabish, is Arabic written in Latin characters with numerals standing in for sounds that Latin script lacks. It is well documented across the Gulf, particularly among younger and bilingual users, and academic work on Kuwaiti youth has described the emergence of a genuine digraphia, two writing systems coexisting for one language. Code-switching within a single query is common too, the sort of thing where a product name in English sits beside a place name in Arabic.
So the behaviour is real. The question a client actually asks is different: how much of Qatar's search volume does it account for?
Nobody has published that figure. A claim circulating in Gulf marketing blogs puts the share of Arabic product searches using local dialect rather than Modern Standard Arabic at over sixty-seven per cent, attributed loosely to a search trends source. It could not be traced to any primary study during research for this article. A separate figure about Arabic-language search share in a neighbouring market has the same problem. Both get repeated because they are convenient.
We are not going to hand you either one. What we do instead is treat Latin-script and mixed-language queries as a hypothesis to be tested against your own search console data, where the answer is specific to your category and actually knowable, rather than as a plank of an initial strategy justified by a statistic with no parent.
If an agency quotes you a percentage for Arabizi search share in Qatar, ask for the study. The request is usually the end of the conversation.
Keyword tools were built English first
Every mainstream keyword tool was designed around English and extended outward. That extension is uneven for Arabic, and practitioners across the Gulf report the same pattern: reported volumes for Arabic terms need a manual check against real search results rather than blind trust.
Several things compound it. Arabic morphology is root-and-pattern based, so a single root generates a large family of related forms, and tools differ in how aggressively they group or split those forms. Diacritics, which carry meaning but are usually omitted in commercial text, introduce another axis of variation. Add the orthographic variants above and the same commercial concept can be spread across a dozen tool entries, each showing a fraction of the real demand.
The workaround is not a better tool. It is a different method: use tools to generate candidates and establish rough relative ordering, then validate the shortlist manually against live results, and involve someone who reads Arabic natively in the judgement about which forms real buyers use. That is slower than exporting a spreadsheet. It is also the only version that produces a list worth writing against.
Mapping register to intent, in practice
The abstract principle becomes operational when you assign register at the page level. This is roughly how the decision runs across common Qatari commercial categories.
| Page type | Primary register | Why |
|---|---|---|
| Tender and procurement pages | Modern Standard Arabic | Bids to Qatari government entities normally go in Arabic unless the tender document states otherwise, and the vocabulary of the tender is the vocabulary of the search. |
| Corporate and capability pages | Modern Standard Arabic, with English in parallel | Institutional credibility sits in the formal register. English runs alongside because a large share of the commercial audience in Qatar does not read Arabic. |
| Service and product detail pages | Mixed, decided per category | Depends on whether the buyer is professional or consumer. This is the row where the wrong blanket decision does most damage. |
| Frequently asked questions | Gulf dialect phrasing, standard orthography | The question is asked the way people speak. The answer can still be written cleanly. |
| Local and near-me intent | Gulf dialect | Conversational and voice-adjacent queries lean dialectal, and the comparative words people reach for differ from the formal equivalents. |
| Editorial and explanatory content | Modern Standard Arabic | Travels across the Gulf rather than only Qatar, and is the register answer engines currently support best. |
Notice that no row says both. Every page gets a primary register and a reason. Pages that try to serve both registers at once usually serve neither convincingly, and they read to a native speaker like a document assembled by committee.
Where AI search changes the calculation
Two platform facts change how a keyword set should be built, and both are recent enough that most existing Arabic keyword strategies predate them.
Google brought AI Overviews to the Middle East and North Africa and to Arabic in May 2025, as part of an expansion reaching more than two hundred countries and territories and more than forty languages. Then in October 2025 it began rolling out AI Mode in thirty-eight new languages, Modern Standard Arabic among them, powered by Gemini 2.5.
Read that second one carefully. The Arabic that Google announced for AI Mode is Modern Standard Arabic specifically. Not Gulf dialect.
Combine it with the benchmark gap from earlier, where Qatari dialect is absent from the main published dialect evaluation, and a pattern emerges that has real strategic weight. The formal register is where answer engines are strongest and best supported. The dialectal register is where a large share of consumer intent lives and where machine support is least measured. Those are not the same page, and they are increasingly not the same channel.
There is a second finding worth knowing here, from peer-reviewed work presented at ACL ArabicNLP in 2025. When the language of a question and the language of the document holding the answer do not match, retrieval and end-to-end accuracy both drop, by margins ranging from roughly thirteen to forty-two per cent depending on the embedding model and the domain tested. The loss follows the mismatch, not the language. Which means keyword research in this market is no longer only about which words to target. It is also about making sure an answer exists in the language the question will be asked in. We go into that finding properly in our work on generative engine optimisation for Qatar, because it changes measurement as much as it changes content.
A sequence that works
Putting it together, the order matters as much as the components.
Start with intent inventory rather than keywords. List the actual decisions your buyers make and who makes them. Procurement officer, department head, consumer, expatriate professional. Assign register to each before touching a tool.
Then build Arabic candidates in Arabic. Not from the English list. Involve someone who reads and writes the language, working from how the category is actually described in Qatari commercial documents and media.
Enumerate orthographic variants per candidate term. Decide which get standalone targeting and which get covered on-page.
Validate the shortlist against live results. Tool volume is a starting hypothesis for Arabic, not a finding.
Build the English set as its own set, in parallel, with its own intent map. It is not a translation of the Arabic and the Arabic is not a translation of it.
Then, and only then, map terms to pages. One primary register per page, one intent per page.
None of this is fast. It is, though, the version that produces a keyword set describing the Qatari market rather than describing an English spreadsheet that passed through a translation step. If you want the wider technical and regulatory context that surrounds this work, our team has documented the full picture for the Qatari market, including the right-to-left engineering that a bilingual build quietly depends on.
What Is Actually Documented, and What Is Not
The right-hand column is not a weakness in the method. Stating it is the method.
Documented
Not measured
Anything in the right-hand column should be treated as a hypothesis to test against first-party data, not as a gap to be filled with a borrowed statistic.
The short version
Arabic keyword research for Qatar is two research projects, not one translated project. Register gets decided per page and per intent, because the formal and spoken varieties are not synonyms and they attract different queries. Orthographic variants get enumerated deliberately, because each one arrives at a search engine as a separate string. Tool volume for Arabic gets treated as a hypothesis rather than a finding, and gets validated against live results by someone who reads the language.
And the gaps get named. Qatari dialect is absent from the main dialect benchmark. No published measurement of the Latin-script share of Qatari search volume was found during research for this article, and the same applies to the split between Arabic and English. Any strategy that quietly fills those gaps with borrowed numbers is building on nothing, and it will be an expensive nothing by the time anyone notices.
Frequently asked questions
Can I not just translate my English keyword list into Arabic?
No, and this is the mistake that costs the most. Translation produces terms that are grammatically correct but were never search terms, because the words a buyer types in Arabic are not the words a translator picks for an English concept. Volume also gets inherited from the English list, which means the resulting plan reflects English demand rather than Qatari demand. Arabic candidates have to be generated in Arabic, from how the category is described in Qatari commercial writing and media, then validated against live results.
Should our Arabic content use Modern Standard Arabic or Gulf dialect?
Both, assigned per page rather than chosen once for the site. Modern Standard Arabic belongs on institutional, legal, procurement and business-to-business pages, because that is the register Qatari organisations write in and the register that tender vocabulary comes from. Gulf dialect belongs on frequently asked questions, consumer pages and local intent, because that is how people speak and search conversationally. A page that tries to do both usually convinces nobody, and reads to a native speaker as inconsistent.
How many spelling variants do we actually need to target?
It depends on the term, which is why enumeration happens per keyword rather than as a global rule. Three patterns account for most of it: hamza omission on an alef, yeh substituted for alef maksura, and teh marbuta alternating with heh. A multi-word phrase where several words each carry a variant can reach a search engine in several distinct forms. Some variants earn standalone targeting, others are better covered within on-page content. The decision is commercial, but it has to be made knowingly rather than by default.
Is it worth targeting Arabizi or Latin-script Arabic queries?
Test it, do not assume it. The behaviour is genuinely documented across the Gulf, including academic work describing two writing systems coexisting among younger users, and mixed-language queries do occur. What does not exist is a published measurement of how much of Qatar's search volume it represents. Percentages circulating in Gulf marketing content could not be traced to any primary study. So the honest approach is to look for it in your own search console data, where the answer is specific to your category and actually verifiable, rather than committing budget on the strength of a borrowed figure.
Why does the absence of Qatari dialect from a benchmark matter to my keyword plan?
Because it changes what you can claim and what you should verify. A 2025 dialect benchmark covering five Arabic dialects across fifteen thousand items does not include Qatari. That means performance on Qatari dialect specifically has not been systematically measured, by anyone, for language models or the tools built on them. It is not evidence that machine handling of Qatari dialect is poor. It is evidence that nobody knows, which is a different and more actionable fact: validate dialect targeting inside your own category with real queries, and treat any vendor claim about dialect performance in Qatar as untested until they show you what it was tested against.
Do keyword tools report Arabic search volume accurately?
Treat their Arabic figures as directional rather than authoritative. The tools were built English first and extended outward, and Arabic makes that extension harder: root-and-pattern morphology spreads one concept across many related forms, omitted diacritics add another axis of variation, and orthographic variants split demand further. The same commercial concept can appear across a dozen separate tool entries, each showing part of the real picture. Use tools to generate candidates and rough ordering, then validate the shortlist manually against live results with native input.
Sources & References
Google MENA, product announcement on the rollout of AI Mode in thirty-eight new languages including Modern Standard Arabic, October 2025. Google, announcement of AI Overviews expansion to the Middle East and North Africa and to Arabic, May 2025.
Amiraz et al., The Cross-Lingual Cost: Retrieval Biases in RAG over Arabic-English Corpora, ACL ArabicNLP 2025, for the measured effect of a mismatch between query language and document language.
DialectalArabicMMLU, 2025, for the five-dialect coverage and the absence of Qatari dialect. Gulf Arabic dialect-group definition from published corpus work covering Bahrain, Qatar, the United Arab Emirates, Kuwait and eastern Saudi Arabia.
Academic work on Arabizi and emerging digraphia among Gulf youth. Claims about dialect share of Arabic product searches and about Arabic search share in neighbouring markets were checked during research and could not be traced to a primary source, and are therefore reported as untraceable rather than repeated as fact.