Pune sits in Maharashtra. The state language is Marathi. Hindi is widely understood but Maharashtra is not a Hindi-primary state, and treating a national Hindi rollout as evidence of Hindi dominance in Pune is a mistake people make constantly. English is the working language of export-facing business. Those four sentences are all true at once, which is why the language question here is harder than it looks.
The usual way this gets resolved is by assertion. Somebody quotes a percentage for how much of Pune's commercial search happens in Marathi, and a content plan gets built on it. We went looking for the source of figures like that across four separate research passes. There is no public source that measures it. Not a weak source, not a dated source. None was located.
That finding is more useful than a number would have been, because it tells you exactly what to do: test before you build. What follows separates what is genuinely verifiable about language in this market from what is assumption wearing a percentage sign, and then gives a validation sequence that produces your own answer in a few weeks.
What We Actually Know About Language in This Market
Product availability is documented and dated. Commercial demand is not. These two get conflated in almost every pitch.
Verifiable
Dated, first-party, checkable
Not measured by anyone
No public source located
Left column: Google India product announcements, each dated, plus published Indic-language retrieval benchmarks. Right column: no public source located across four independent research passes conducted for this material.
Created by Arfadia • arfadia.com/blog
Why product availability keeps getting mistaken for demand
The confusion is easy to understand. Google adding Marathi to a flagship product is a real signal that somebody at Google modelled the opportunity and found it worth building for. It would be strange to read that as evidence of nothing.
But it is evidence about supply, not about your buyers. A platform adds a language because of total population reach across a country, and your question is narrower by several orders of magnitude: how many people, in one city, searching commercially, for your category, in that language, with budget. A national language rollout tells you almost nothing about that intersection.
The two questions get merged because merging them produces a confident sentence, and confident sentences win pitches. It is the same failure pattern as attributing a state export figure to a city. The number is real, the inference is not, and the person who eventually checks will find the gap.
| Event | What it proves | What it does not prove |
|---|---|---|
| AI Overviews launched in India, August 2024, English and Hindi | AI answers have been present on Indian results for roughly two years | How often they appear for your category, or in which language your buyers ask |
| AI Mode added Hindi, September 2025 | Hindi is supported in a flagship surface | That Hindi matters in Maharashtra, which is not a Hindi-primary state |
| AI Mode added Marathi and six others, October 2025 | Marathi queries can receive AI answers | That anybody is asking commercial questions in Marathi at scale in your category |
| Search Live added Marathi, March 2026 | Voice interaction in Marathi is supported | That voice is a channel for business-to-business procurement |
| Published Indic retrieval benchmarks show weaker performance than English | Retrieval in these languages is a genuine engineering problem | A specific accuracy figure. Reviewed benchmark numbers conflicted and are not restated here |
The code-mixing problem nobody plans for
Even if you decide to publish in Marathi, there is a second decision that content plans routinely skip: which form of Marathi.
Real users type in more than one register. Some search in Devanagari. Some type Marathi words in Latin characters. Many mix Marathi and English inside a single query, particularly for technical or commercial terms where the English word is simply the word people use. Those are three distinct query populations, and keyword research done in only one of them will systematically underestimate the other two.
This also affects measurement in a way that catches people out. A romanised Marathi query can return English pages that happen to match the transliteration, which looks like a Marathi result and is not. Before concluding that native-language content is winning, check whether the pages ranking are actually in the language, or whether they are English documents matching a string.
The practical consequence is that keyword research has to be run in both scripts, that Search Console data has to be segmented by page language rather than assumed, and that a native speaker has to look at the results page. There is no tool-only path through this.
Validating a Language Before You Publish In It
Runs in a few weeks. Produces your own answer for your own category, which is the only answer that exists.
Keyword research in both scripts, location-targeted
Run Keyword Planner with Pune location targeting, once in Devanagari and once in Latin characters, including code-mixed forms. Compare against the English set for the same intents rather than in isolation.
Segment Search Console by page language
Split existing query data by the language of the page that received the impression, using script detection rather than assumption. You may already have signal you have never looked at.
Run a paid test with matched ads and pages
Language-matched ad copy pointing at a language-matched page. Mismatched tests measure the mismatch, not the language. This is the fastest source of real demand data available.
Track outcomes by language in the CRM
Enquiry volume is not the outcome. Qualified enquiries and closed business by language of first touch is. A language can generate traffic and no pipeline.
Review the results page with a native speaker
Check whether ranking pages are genuinely in the language or English documents matching a transliterated string. A tool cannot tell you this and a non-speaker will misread it.
Then the decision point. If steps three and four show qualified demand, build properly, with native writers and named reviewers rather than machine translation. If they do not, you have spent a few weeks instead of a content budget, and you have an answer nobody else in your category has.
Validation sequence assembled from the cross-validated research behind Arfadia's Pune pages. No published source measures Marathi commercial query volume in Pune, which is why this produces your own data rather than citing someone else's.
Created by Arfadia • arfadia.com/blog
The citation question, which is separate again
Everything above concerns whether people search in a language. There is a second question that matters if AI visibility is part of your programme, and it is genuinely open.
Suppose you publish an excellent Marathi page. When someone asks a Marathi question and receives a Marathi answer, does that page get cited, or does the system draw on an English source and simply answer in Marathi. Those two outcomes look identical to the user and are completely different for you. In the first case your Marathi investment earns the citation. In the second case your English content is doing the work and the Marathi page is decoration.
Nobody has published a measurement of this. Retrieval in Indic languages is documented as weaker than in English, and Indic content is thin in training material, both of which make the second outcome plausible. Plausible is not measured. And this is a question you can answer for your own category in a fortnight by running the same intents in both languages and coding which sources the answers cite.
The practical hedge, until you have that data, is to pair rather than choose. Publish the native-script page and ensure the equivalent English content covers the same facts with the same structure. If the citation comes through English, you are covered. If it comes through Marathi, you are covered. The cost of the hedge is one extra page and it removes a bet you have no basis for making.
What the one solid India measurement actually says
Since so much of this article is about the absence of data, it is worth naming the exception, because there is one and it is good.
Researchers at MIT ran 24,000 search queries across 243 countries, generating 2.8 million AI and traditional results in 2024 and 2025. Crucially, 12,000 of those queries were identical across both years, which isolates changes in platform behaviour from changes in how people search. India came out at 54 percent more searches returning an AI result in 2025 than in 2024.
Read that carefully, because it is easy to overstate. It is a relative change between two years. It does not say that 54 percent of Indian searches show an AI answer, and any vendor quoting it that way has misread it. What it does establish is direction and magnitude, from a primary source, for India specifically, which is more than any prevalence figure in circulation can claim.
The same study produced a finding that matters more for language planning than the headline number. Across a seven-country panel including India, question-style queries returned an AI answer roughly 60 percent of the time, statements around 37 percent, and navigational queries only about 12 percent. Query shape drives exposure. So the language decision and the query-shape decision interact: a Marathi page built around question-shaped intents sits in a much higher exposure band than the same page built around brand-name lookups, regardless of language.
One more figure from the same work, offered as context rather than as part of the Pune decision. Indonesia registered a 76 percent rise over the same period, higher than India's. For a Pune company selling into Southeast Asia, that is a reason to treat the destination market's AI exposure as a separate planning input rather than assuming the home market's profile travels.
What the Indic-language research actually found
The retrieval gap gets waved at more often than it gets described, so it is worth naming the work rather than gesturing at it, and being precise about what each study does and does not establish.
IndicRAGSuite evaluates retrieval and response generation across Indian languages including Hindi and Marathi. It is a benchmark for system performance. It does not measure whether consumer assistants preferentially cite Indian websites for Indian users, and reading it as though it does is a category error.
A separate 2025 study examining cultural bias in large language models found a weighted average recall gap of around 0.26 across models, with English-origin content retrieved more reliably than Indian-origin content, and a precision gap of similar magnitude. That is a measured retrieval asymmetry in content origin rather than in language alone, which is a subtly different and arguably more uncomfortable finding.
Reviews of Marathi natural language processing point in the same direction, describing the language as under-resourced relative to English in the material these systems learn from. Perplexity-side research reviewed for this material notes Indic content sitting well below one percent of global web-crawl data.
Where we deliberately stop is at accuracy percentages for specific languages. Two of the research passes produced detailed per-language accuracy tables, and the tables disagreed with one another, with one of them internally incoherent as transcribed. A number that two sources contradict is not a number. The directional finding is solid and sufficient: retrieval in Indian languages is measurably weaker than in English, and content origin appears to matter alongside language.
Adoption, and why it does not answer the language question
People reach for adoption statistics when the language data runs out, and the two are less connected than they look.
The platform-level figure is large: 100 million weekly active ChatGPT users in India, disclosed by the platform in February 2026. The population-level figure is more modest: the Microsoft AI Economy Institute measured India's diffusion rate, the share of population estimated to be using generative AI tools, at 14.2 percent in the first half of 2025 rising to 15.7 percent in the second half, against a global average around 16.3 percent.
Both are true. A country of India's size can hold one of the world's largest absolute user bases and a population share near the global average simultaneously. Neither figure tells you what language those users type in, what they search for commercially, or whether any of them are in Pune buying what you sell.
That is the recurring shape of this whole topic. There is a great deal of data about India and almost none about the specific intersection that determines your content plan. Adoption is not language. Language support is not demand. Demand is not citation. Each step needs its own evidence, and where the evidence stops, so should the claim.
Budget, honestly
Language decisions get made emotionally more often than they get made with numbers, so it helps to look at where the money actually goes.
A native-language page is not a translation line item. Done properly it involves a native writer, a named native reviewer, keyword research in two scripts, its own measurement segment, and ongoing maintenance every time the underlying facts change. Multiply that by the number of pages you would need for the language to be credible rather than token, because a single Marathi page attached to an otherwise English site reads as a gesture and converts like one.
Against that, the validation sequence costs weeks and a modest paid-test budget. The asymmetry is the whole argument. Testing is cheap and building is not, so the order should be obvious, and yet content plans routinely commit to the language first and look for the evidence afterwards.
There is also a sequencing argument specific to exporters. If most of your revenue comes from outside India, the marginal English page aimed at a destination market almost certainly outperforms the marginal Marathi page aimed at a domestic segment you have not yet measured. That is not an argument against Indic languages. It is an argument about what to do in the first two quarters, with the rest deferred until there is data to defer to.
What we would actually recommend for a Pune exporter
English first, and by a wide margin, for anything aimed at buyers outside India. Not because Indian languages do not matter, but because the buyer in Frankfurt or Singapore is not searching in Marathi and the evidence for domestic Indic commercial volume in your category does not exist yet.
Then run the validation sequence for Marathi and Hindi as a small, time-boxed exercise rather than a strategic decision made in advance. Weeks, not quarters. The output is a decision with evidence behind it, which is a better position than either building on an assumption or dismissing the languages because nobody has proven they work.
If validation says yes, staff it properly. Native writers, native reviewers, both scripts covered, code-mixed variants accounted for. Machine-translated Marathi reads as machine-translated Marathi to the people it is meant to serve, and it damages the credibility of everything else on the site.
If validation says no, write that down with the date and the evidence, and revisit in a year. Markets change, platform support keeps expanding, and a documented no is a decision you can defend and reverse. An undocumented assumption is neither.
Frequently Asked Questions
How much of Pune's commercial search happens in Marathi?
No public source measuring it was located across four separate research passes conducted for this material. That is not a hedge, it is the finding. Anyone quoting a percentage for this should be asked to name the study, the sample and the method, because the figures in circulation could not be traced to any of those.
Google supports Marathi in AI Mode. Does that mean Marathi search matters commercially in Pune?
It means Marathi queries can receive AI answers, which is a supply fact rather than a demand fact. A platform adds a language on the basis of national population reach. Your question is far narrower: how many people, in one city, searching commercially, in your category, in that language, with budget. A national rollout tells you very little about that intersection.
Is Hindi a safe default for Pune because it is a national language?
No. Pune is in Maharashtra, where the state language is Marathi, and Maharashtra is not a Hindi-primary state. A national Hindi rollout is frequently misread as evidence of Hindi dominance locally. Hindi may still be worth testing, but it should be tested rather than assumed.
What is the fastest way to find out whether we need a Marathi page?
A five-step sequence that runs in weeks. Keyword research with Pune location targeting in both Devanagari and Latin script including code-mixed forms. Search Console data segmented by the language of the page receiving impressions. A paid test with language-matched ads and language-matched pages. Outcomes tracked by language of first touch in the CRM, measuring qualified enquiries rather than traffic. And a results-page review by a native speaker.
Why does a native speaker need to review the results page?
Because a romanised Marathi query can return English pages that happen to match the transliteration. That looks like a Marathi result and is not. Without someone who reads the language checking what is actually ranking, it is easy to conclude that native-language content is winning when English documents are matching a string.
Do Marathi pages actually get cited in Marathi AI answers?
Nobody has published a measurement of it. When a Marathi question receives a Marathi answer, the system may be citing a Marathi source or citing an English source and answering in Marathi. Those look identical to the user and are completely different for the business. Retrieval in Indic languages is documented as weaker than in English and Indic content is thin in training material, both of which make the second outcome plausible, but plausible is not measured. Until you have your own data, pair the native-script page with equally well-structured English covering the same facts.
Is machine translation acceptable for Indic-language pages?
No. Machine-translated Marathi reads as machine-translated Marathi to exactly the audience it is meant to serve, and it undermines the credibility of everything else on the site. If validation shows demand, staff it with native writers and named native reviewers, cover both scripts, and account for code-mixed forms. If validation does not show demand, the money is better spent elsewhere.
What should a Pune exporter default to?
English, by a wide margin, for anything aimed at buyers outside India, since a procurement lead in Frankfurt or Singapore is not searching in Marathi. Then run the Indic-language validation as a small time-boxed exercise rather than a strategic decision taken in advance. If the answer is no, record it with the date and the evidence and revisit in a year. A documented no is defensible and reversible. An undocumented assumption is neither.
Sources & References:
- Google India product announcements: AI Overviews launched in India in August 2024 in English and Hindi; AI Mode launched in English in June 2025, added Hindi in September 2025, and added Marathi together with six further Indian languages in October 2025; Search Live added Marathi in March 2026. Each dated first-party product announcement. These record product availability and are not measurements of commercial search demand.
- Marathi share of commercial search volume in Pune, Hindi share of business-to-business query volume in Pune, and English commercial query share at city level: no public source measuring any of these was located across four independent research passes conducted for this material. One research pass produced a numeric split for these quantities while simultaneously stating that no study measuring them was found. That numeric split was excluded.
- Retrieval quality in Indic languages: multiple published benchmarks and research reviews document weaker retrieval performance in Indian languages than in English, and note that Indic-language content is thinly represented in web-scale training material. Specific accuracy figures encountered in the reviewed research conflicted with one another and are deliberately not restated here.
- IndicRAGSuite: evaluates retrieval and response generation across Indian languages including Hindi and Marathi. A system-performance benchmark. It does not measure whether consumer assistants preferentially cite Indian websites for Indian users.
- Cultural bias in large language models, 2025 study reviewed for this material: weighted average recall gap of approximately 0.26 across models, with English-origin content retrieved more reliably than Indian-origin content, and a precision gap of similar magnitude. Measures content-origin asymmetry rather than language alone.
- Per-language accuracy figures: two research passes produced detailed per-language accuracy tables for Indic languages. The tables conflicted with one another and one was internally incoherent as transcribed. No per-language accuracy percentage is stated in this article for that reason.
- ChatGPT weekly active users in India: 100 million, disclosed by OpenAI, February 2026, first-party. Population-level generative AI diffusion in India: Microsoft AI Economy Institute, Global AI Adoption in 2025, 14.2 percent in H1 2025 rising to 15.7 percent in H2 2025 against a global average of approximately 16.3 percent. These measure different populations and must not be conflated.
- Whether well-structured Marathi pages earn citations in Marathi-language AI answers, or whether English sources carry the citation while the answer is delivered in Marathi: no published measurement located. Stated as an open question.
- Whether the same prompt issued from Pune retrieves different source nationalities than the same prompt issued from outside India: no public controlled study holding prompt intent, model version, session state and account history constant while varying only location and language was located.
- Code-mixing and script variation: Marathi commercial queries occur in Devanagari, in Latin transliteration, and in mixed Marathi and English forms. Keyword research conducted in a single script systematically underestimates the others. Practitioner observation supported across the reviewed research rather than a measured statistic.
- Validation sequence assembled from the cross-validated research behind Arfadia's Pune pages. It produces first-party data for a specific category and market rather than citing a general figure, because no general figure exists.
- Sinan Aral, Haiwen Li and Rui Zuo, The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale, Massachusetts Institute of Technology, arXiv 2602.13415v2. 24,000 queries across 243 countries producing 2.8 million AI and traditional results in 2024 and 2025, with 12,000 identical queries repeated across both years. India recorded a 54 percent rise in searches returning an AI result; Indonesia recorded 76 percent. Query-style split reported for a seven-country panel including India. This is a relative change between two years and is not a statement of absolute prevalence.
- This article makes no claim about the conversion performance of any language and offers no guarantee of ranking or citation outcomes in any language.