Most markets have one set of answer engines to think about. A brand works out where its buyers ask questions, tests across ChatGPT and Gemini and a handful of others, and that is the surface.
Qatar has that surface plus a sovereign one. The country commissioned, built and now maintains its own Arabic model family, running inside its own borders on its own hardware, deliberately structured as a closed environment so that data does not leave. No global agency playbook mentions it, because no global agency playbook was written for a country that did this.
What follows is what the stack actually contains, what an independent benchmark found when it finally measured the model against twenty-five others, and which of those findings a Qatari organisation should act on. Some of it is more impressive than the marketing suggests. Some of it is considerably less.
What Fanar actually is
Fanar is built by Qatar Computing Research Institute at Hamad Bin Khalifa University, with the Ministry of Communications and Information Technology. It is not a single model. It is a family, and it has a version history worth stating carefully rather than as a snapshot.
The first release arrived in December 2024 at nine billion parameters. The second, Fanar 2.0, was announced on 9 December 2025 at the opening of the second World Summit AI in Doha, at twenty-seven billion parameters. The institute's executive director confirmed at that same summit that work on a third version had already begun, targeted for December 2026.
Read that last sentence and then be careful with any content, including ours, that describes a current version. This is a moving target on an annual cadence, which is exactly why we describe it as a family with a history rather than as a product with a specification.
The technical details of Fanar 2.0 come from the institute's own published work. The core model is continually pre-trained from a Gemma-3-27B backbone, which is a Google model, and that is worth noting plainly: sovereign does not mean built from nothing. It means the training, the data pipelines and the deployment infrastructure were designed and operated within the institute. The work ran on 256 NVIDIA H100 graphics processors.
The efficiency claim is the part that impressed us most. According to the institute, Fanar 2.0 delivers improvements across every benchmark it reports despite using roughly eight times fewer pre-training tokens than the first version. The stated gains are specific: Arabic knowledge up 9.1 points, language up 7.3, dialects up 3.5, and English capability up around 7. The strategy behind that was data quality over volume, plus targeted continual pre-training and model merging.
One honest wrinkle. The institute's publication page describes a curated corpus of around 120 billion high-quality tokens across three data recipes. The model card for the same release describes roughly 166 billion Arabic and English tokens. Two figures from the same publisher for what appears to be the same thing. We are not going to pick one and present it as fact, and if a vendor quotes you a precise token count for this model, ask which of the two they are using and why.
Beyond the core language model, the family is broader than most people realise. It includes a bilingual moderation filter for Arabic safety and cultural alignment at four billion parameters, a speech family that added long-form Arabic speech recognition capable of handling hours-long audio with speaker changes, a vision family for Arabic-aware image and video understanding, a component for classical Arabic poetry generation, a bilingual translation component, and an agentic tool-calling framework for multi-step workflows. It supports Modern Standard Arabic alongside Gulf, Levantine and Egyptian dialects.
And it includes one component that matters more than the rest for the benchmark discussion coming up: a multi-agent system built specifically for Islamic content.
Why a country builds its own model
This did not start with the model. It started with policy, and the policy is unusually explicit.
Qatar published a national artificial intelligence strategy in 2019, developed with Qatar Computing Research Institute and adopted through what is now the Ministry of Communications and Information Technology, announced at the country's technology summit that year. It is organised around six pillars covering talent and education, data, employment, business and economy, research, and ethics.
What makes it relevant here is that the strategy names Arabic language processing as an explicit national priority. Not as a general aspiration about technology. As a specific ambition to lead in the use of artificial intelligence for the Arabic language, citing existing institutional assets including an Arabic text-processing suite and a speech-to-text system developed for a major Arabic broadcaster.
A national artificial intelligence committee was established within the ministry through a cabinet decision in 2021. The Qatar Central Bank has since issued artificial intelligence guidance for licensed financial institutions, which matters because it means the regulated sector has a supervisory expectation attached to how these systems get used rather than a policy vacuum.
So the sequence runs: strategy names Arabic as a priority in 2019, governance structure follows in 2021, sovereign model ships in 2024, second version in 2025, sector guidance in place for financial institutions. That is a country executing a stated plan over six years, which is a different thing from a country buying technology.
The commercial consequence for anyone preparing a tender response or a board paper in Qatar: alignment with stated national priorities functions as a procurement signal here in a way it does not in private-sector-led markets. That is not a claim about how tenders are scored. It is an observation about what the buyer has already committed to publicly.
Six Years of Stated Plan, Executed in Order
The reason to read this as a timeline rather than a specification is at the bottom of it.
2019
National AI strategy names Arabic as a priority
Six pillars covering talent, data, employment, economy, research and ethics. Arabic language processing stated as an explicit national ambition rather than a general technology goal.
2021
Governance structure established
A national artificial intelligence committee created within the ministry through a cabinet decision, giving the strategy an institutional owner.
December 2024
First sovereign model ships at nine billion parameters
Built by the national computing research institute at Hamad Bin Khalifa University with the communications ministry.
December 2025
Second version at twenty-seven billion parameters
Continually pre-trained from a Gemma-3-27B backbone on 256 NVIDIA H100 processors, entirely inside the institute. Reported gains across every benchmark despite roughly eight times fewer pre-training tokens than the first version.
Announced for December 2026
Third version already in development
Confirmed by the institute's executive director at the same summit where the second version launched.
Which is why version numbers are the wrong thing to build on
Any content, strategy or vendor claim that describes a specific version as the current state has a shelf life measured in months on this cadence. Track the family and the direction, not the specification.
Sourced from the institute's own technical publication and model documentation, plus statements made at the December 2025 summit. Token count for the second version is reported inconsistently by the publisher and is not stated here as a single figure.
Then somebody independent measured it
Until recently, almost everything anyone could say about Fanar's capability came from the institute that built it. That is not an accusation, it is a normal situation for a new model, and the institute's own reporting has been detailed. But self-reported benchmark gains and independent evaluation are different categories of evidence.
In 2026 an independent benchmark arrived. Researchers at the University of Edinburgh published IslamicMMLU, an evaluation of language model performance on Islamic knowledge, and it is substantial: 10,013 multiple-choice questions across three tracks, Quran with 2,013 questions, Hadith with 4,000, and jurisprudence with 4,000, spanning twelve distinct task types. All questions were generated from native Arabic source material rather than translated from English. Twenty-six models were evaluated, with results published to a public leaderboard.
Why does this matter for a commercial audience rather than a scholarly one? Because it is the first independent measurement of Qatar's sovereign model family against frontier competitors on a domain the country's own stack was specifically built to serve. For any organisation in Islamic finance, halal commerce, or a regulated Qatari sector where religious knowledge accuracy has commercial weight, this is the closest thing to an audited capability statement that currently exists.
The overall spread across all twenty-six models ran from 39.8 per cent to 93.8 per cent, against a random-guess floor of 25 per cent for four-option questions. Here is where the Arabic-specialised models landed alongside a selection of the general-purpose field.
| Model | Overall | Quran | Hadith | Jurisprudence |
|---|---|---|---|---|
| Highest scoring model tested | 93.8% | 99.3% | 93.0% | 89.1% |
| Second highest | 92.3% | 97.3% | 91.7% | 87.9% |
| Fanar-Sadiq, the Islamic-content component | 81.6% | 94.4% | 72.9% | 77.4% |
| Fanar, the general model | 66.0% | 56.9% | 66.5% | 74.5% |
| A Gulf regional Arabic model, 70 billion parameters | 63.4% | 48.1% | 77.8% | 64.4% |
| A Saudi Arabic model, 7 billion parameters | 59.5% | 43.2% | 70.9% | 64.5% |
| Fanar-C, a 27 billion parameter variant | 54.0% | 41.9% | 60.7% | 59.5% |
| Lowest scoring baseline model | 39.8% | 32.4% | 41.6% | 45.2% |
What those numbers actually say
Four readings, and only the first one is flattering.
The specialised component performs. Fanar-Sadiq reached 81.6 per cent overall, placing eleventh of twenty-six models, ahead of several well-known general-purpose systems. On the Quran track it scored 94.4 per cent, which is competitive with the frontier models at the top of the table. For a model built by one national research institute against systems built by the largest technology companies in the world, that is a genuine result.
Its performance is uneven, sharply. The same model scored 72.9 per cent on Hadith, more than twenty points below its Quran score. The researchers noted this pattern explicitly as evidence that a single aggregate figure conceals domain-specific weaknesses. Anyone quoting the 81.6 per cent headline without the Hadith figure is giving you half a result.
The general Arabic models did not do well. Fanar at 66.0 per cent and the 27 billion parameter variant at 54.0 per cent both sat with or below general-purpose mid-tier models. The two other Gulf regional Arabic models tested, at 70 and 7 billion parameters, landed at 63.4 and 59.5 per cent. Large-scale Arabic pre-training did not deliver on this domain.
And the researchers drew the conclusion that matters most. Their finding was that Arabic pre-training alone appears insufficient for Islamic knowledge, and that domain-specific training data and retrieval augmentation matter more than language coverage does.
Sit with that last one, because it is the single most useful sentence in the paper for anyone doing commercial content work. The component that succeeded was the one with domain-specific data and a retrieval architecture around it. The models that merely had a lot of Arabic did not. Language coverage is not domain competence.
The direct translation into content strategy: publishing a large volume of Arabic material does not make you the authority a system reaches for. Publishing well-structured, well-sourced, domain-specific Arabic material with clear retrievable units is what does. That is the same conclusion the citation research reaches from a completely different direction, which is the kind of convergence worth trusting. We cover the retrieval mechanics behind it in our search practice for the Qatari market.
The bias finding nobody expected
The benchmark included something genuinely novel, and it is worth explaining because it has implications well beyond religious content.
Sunni jurisprudence recognises four schools of thought, all considered equally valid. The same question can receive different rulings in each, and none is wrong. So the researchers built 800 questions where all four answer options were correct, one per school, with the question phrased without specifying a school. A model's choice therefore reveals which tradition it defaults to when nothing constrains it.
A model with no preference would select each school roughly a quarter of the time. What they found was that preferences were model-specific rather than uniform. The Arabic-specialised models showed a moderate preference for one school at around 32 to 33 per cent, which the researchers suggested may reflect the composition of Gulf training data. One frontier model showed a notable preference for a different school at 35.6 per cent. The top-scoring model was close to uniform.
They also found a moderate negative correlation between accuracy and bias magnitude, suggesting higher-performing models tend toward more balanced selection, with exceptions.
The generalisable lesson is not about jurisprudence. It is that when a model is asked a question with multiple legitimate answers, it picks one, and its pick reflects its training data rather than a neutral survey of the options. Replace schools of jurisprudence with competing industry standards, regional business practices, or vendor categories, and the structure of the problem is identical. A model asked which approach is correct, where several are, will name one. Whether it names yours is not a neutral process.
What the Benchmark Covers, and What It Explicitly Does Not
The right-hand column comes from the paper's own limitations section. A result quoted without it is a result quoted badly.
Measured
Out of scope, per the authors
The authors also discourage using their results to market any system as an authoritative religious advisor, and note that religious guidance should involve qualified human scholars. We repeat that here because it is their position and it is the right one.
The limits the authors state themselves
An article arguing for evidence discipline should hold itself to it, so the limitations above deserve prose rather than only a panel.
The benchmark covers Sunni tradition only. The Quran and Hadith tracks use Sunni canonical sources and the jurisprudence track covers the four Sunni schools. The authors state directly that this limits applicability to roughly fifteen per cent of the global Muslim population, and that extending coverage is planned.
The jurisprudence corpus derives entirely from one encyclopedia, an authoritative multi-volume work commissioned by a major Islamic university, but a single source nonetheless. That constrains coverage of minority positions within schools, regional variation, and contemporary independent legal reasoning.
Human validation was partial. A comparative jurisprudence professor reviewed 213 sampled questions and approved 207, a rate above ninety-seven per cent, with three requiring minor revision and three rejected for factual errors. The authors acknowledge that agreement across multiple experts would strengthen the validation and that a single reviewer limits how far the quality evidence generalises.
The format measures recognition rather than production. A model that picks the right ruling from four options may not generate that ruling unprompted, and four options permit a quarter of correct answers by chance.
And the results are dated. Models were accessed by application programming interface across a window in late 2025 and early 2026. Rankings and bias patterns move as models update.
Notice what this list does not do. It does not undermine the finding. It bounds it, which is what a limitations section is for, and it gives you the questions to ask when somebody quotes one number from it at you.
What a Qatari organisation should actually do with this
Four practical positions follow.
Decide whether the sovereign surface is yours. It genuinely matters for some organisations and is secondary for others. If you are state-linked, in a regulated sector, or operating in Islamic finance or halal commerce, a national tool running as a closed environment where data does not leave the country is a real consideration and probably one your own buyers and internal teams are already using. If you are a private consumer brand, the global engines carry the overwhelming majority of discovery and the sovereign stack is a secondary surface. Either answer is defensible. Assuming without asking is not.
Set engine coverage per client, not per template. Global answer engines do most of the work in Qatar, and Google brought its Arabic answer features to the region across 2025. Adding a sovereign model to the tested set is a decision that follows from who your buyers are, not from a standard package.
Treat the domain-competence finding as the strategic one. Not the parameter count, not the version number, and not the leaderboard position. The researchers found that domain-specific data and retrieval structure beat raw language coverage. That is a statement about what makes content authoritative to a machine, and it says: depth in your category, clearly structured and properly sourced, outperforms volume in your language.
Expect the numbers to move and design for that. A third version of the sovereign model family is targeted for December 2026. The benchmark authors say their rankings may shift with model updates. Anything you publish that hangs on a specific version or a specific score has an expiry date, so anchor content to mechanisms and structures instead. For the surrounding picture on how these decisions interact with the technical and regulatory layers in this market, our team has documented the full scope.
The short version
Qatar named Arabic language processing a national priority in 2019, built the governance for it in 2021, shipped a sovereign Arabic model in December 2024 at nine billion parameters, and a second version in December 2025 at twenty-seven billion, trained inside the country on 256 processors from a Google backbone, with a third version targeted for December 2026. That is a plan executed in sequence, and it creates a retrieval surface that exists in almost no other market.
An independent benchmark from the University of Edinburgh then measured it against twenty-five other models across 10,013 questions. The Islamic-content component reached 81.6 per cent overall, eleventh of twenty-six, with a Quran score competitive with frontier models and a Hadith score more than twenty points lower. The general Arabic models scored 66.0 and 54.0 per cent, at or below general-purpose mid-tier systems.
And the authors' conclusion is the part to keep: Arabic pre-training alone was not sufficient, while domain-specific data and retrieval augmentation mattered more than language coverage. Which is a research finding that happens to describe, precisely, why volume of Arabic content is not the same thing as authority in Arabic.
Frequently asked questions
What is Fanar and who built it?
Fanar is Qatar's sovereign Arabic model family, built by Qatar Computing Research Institute at Hamad Bin Khalifa University together with the Ministry of Communications and Information Technology. The first release arrived in December 2024 at nine billion parameters. The second version was announced on 9 December 2025 at twenty-seven billion parameters, continually pre-trained from a Gemma-3-27B backbone on 256 NVIDIA H100 processors, with the training, data pipelines and deployment infrastructure operated inside the institute. A third version has been confirmed as in development, targeted for December 2026. It supports Modern Standard Arabic alongside Gulf, Levantine and Egyptian dialects, and the family includes separate components for moderation, speech, vision, translation, poetry and Islamic content.
Does sovereign mean it was built from scratch?
No, and the distinction is worth stating plainly rather than letting the word do the work. The core of the second version is continually pre-trained from a Gemma-3-27B backbone, which is a Google model. What is sovereign is the training, the data curation, the infrastructure and the deployment, all of which the institute designed and operated within Qatar, and the closed environment that keeps data from leaving. That is a meaningful form of sovereignty for an organisation with data residency requirements. It is not the same claim as building a foundation model from nothing, and any vendor conflating the two is worth questioning.
How does Fanar compare to global models on an independent test?
The clearest available answer comes from IslamicMMLU, published by researchers at the University of Edinburgh, which evaluated twenty-six models across 10,013 Arabic questions covering Quran, Hadith and jurisprudence. The Islamic-content component of the Qatari family scored 81.6 per cent overall, placing eleventh of twenty-six, ahead of several well-known general-purpose systems, with 94.4 per cent on the Quran track which is competitive with the highest scorers. Its Hadith score was 72.9 per cent, more than twenty points lower, which the researchers highlighted as evidence that aggregate scores hide domain weaknesses. The general Arabic model scored 66.0 per cent and a 27 billion parameter variant scored 54.0 per cent, both at or below general-purpose mid-tier models.
What was the most important conclusion of that benchmark?
That Arabic pre-training alone appears insufficient for domain knowledge, and that domain-specific training data and retrieval augmentation matter more than language coverage. The component that performed was the one with domain-specific data and a retrieval architecture around it. The models that simply had large-scale Arabic pre-training did not perform as well, including two other Gulf regional models at 63.4 and 59.5 per cent. For anyone producing content, the translation is direct: volume of Arabic material does not create authority, whereas depth in a specific domain, clearly structured and properly sourced, is what a retrieval system can actually use.
Should our organisation include the sovereign model in our AI visibility testing?
It depends on who your buyers are, and the honest answer is that it matters a great deal for some organisations and very little for others. If you are state-linked, operating in a regulated sector, or working in Islamic finance or halal commerce, a national tool that runs as a closed environment with data staying in-country is a real surface and likely one your buyers or internal teams already use. If you are a private consumer brand, global answer engines carry the overwhelming share of discovery and the sovereign stack is secondary. The decision should be made per client after asking, rather than included or excluded by default in a standard package.
What is the madhab bias finding and why does it matter outside religious content?
The benchmark included 800 questions where all four answer options were correct, one representing each of the four equally valid Sunni schools of jurisprudence, with the question phrased without specifying a school. A model with no preference would pick each roughly a quarter of the time. Instead preferences were model-specific: the Arabic-specialised models leaned toward one school at around 32 to 33 per cent, one frontier model leaned toward a different school at 35.6 per cent, and the top scorer was close to uniform. The generalisable point has nothing to do with jurisprudence. When a model faces a question with several legitimate answers, it picks one, and its pick reflects its training data rather than a neutral survey. Substitute competing industry standards or vendor categories and the structure is identical.
Sources & References
Abdelaal, Al Haffar, Fawzi and Magdy, IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge, University of Edinburgh, preprint 2026, for the question counts, the twenty-six model evaluation, all per-track scores, the madhab bias methodology and findings, and the stated limitations.
Qatar Computing Research Institute at Hamad Bin Khalifa University, Fanar 2.0 technical publication and model documentation, for the parameter counts, the Gemma-3-27B backbone, the 256 NVIDIA H100 training configuration, the reported benchmark gains, the token efficiency claim and the component list. Fanar Team, Fanar: An Arabic-Centric Multimodal Generative AI Platform, preprint 2025.
Qatar News Agency and reporting from the second World Summit AI in Doha, December 2025, for the launch date of the second version and the confirmation of a third version targeted for December 2026. Hamad Bin Khalifa University, for the 2019 national artificial intelligence strategy, its six pillars and its stated Arabic language processing priority.
Note on a discrepancy left open rather than resolved: the institute's publication page describes a curated corpus of around 120 billion tokens for the second version, while the model documentation for the same release describes roughly 166 billion. Both figures come from the publisher and are reported here as inconsistent rather than reconciled into one number.