Ask five agencies what percentage of Philippine search queries run in English rather than Filipino or Taglish. You will get five confident percentages. All five are estimates.
No published study measures the split at the level of commercial search queries. The studies that do exist measure something adjacent and get quoted as though they measured this.
That matters commercially, because the answer decides whether you translate a content library or leave it in English. Translating a hundred pages against an assumption is an expensive way to discover you were wrong about your own audience. The good news, and it is genuinely good news: the number you actually need is not the national one. It is yours, and most of it is sitting in Search Console right now.
What the research actually establishes
Tagalog-English code-switching in Philippine online writing is well documented. TweetTaglish, published in the ACL Anthology in 2022, exists specifically as a dataset for studying it. Academic work on Filipino-English code-switching in online academic discourse found intensive within-turn switching, including Tagalog discourse markers appearing inside otherwise English exchanges.
None of that measures search queries.
A dataset of tweets tells you how Filipinos write to each other on a social platform. A search box is a different genre with different conventions, and people routinely type differently into it than they speak or post. Treating social-media code-switching prevalence as a proxy for query language is a category error that happens to sound reasonable, which is the most durable kind.
What is established on the model side is sharper, and it cuts the other way. FilBench, published at EMNLP 2025, evaluated 27 large language models on Filipino, Tagalog and Cebuano tasks. The best performer reached 72.23 percent aggregate. Its text-generation subscore was 46.48 percent. A separate 2025 study of retrieval-augmented models found accuracy highest in English, lower in Taglish, lowest in pure Tagalog.
So the honest position looks like this. Code-switching is real and common in Philippine online language. Its prevalence in commercial search is unmeasured. And machine handling of Filipino is demonstrably weaker than machine handling of English, which is a reason to be careful about generating Filipino content with tools rather than writing it.
Four Things That Get Confused With Each Other
Each measures a different population and a different behaviour. That confusion is what produces confident percentages nobody can source.
Code-switching in social writing
Well documented and academically studied, with dedicated datasets built from Philippine social media text. Tells you how people write posts. Says nothing about what they type into a search box.
Language mix in commercial queries
Unmeasured. No study analyses actual Philippine query logs by language at token level. Every percentage in circulation for this is an inference, however precisely it gets stated.
Model performance on Filipino
Measured, and lower than English. FilBench found the strongest of 27 models scored 72.23 percent aggregate with a generation subscore of 46.48 percent on Filipino tasks.
Business publishing language
English dominates corporate, professional and B2B publishing. The EF English Proficiency Index 2024 places the Philippines 22nd of 116 on a score of 570, second in Asia after Singapore.
The number you can actually get
Your own query-language mix, from your own Search Console export, segmented by landing page, device, region and intent. It takes an afternoon to produce. It is true for your category rather than for a national average that would flatten every category together. And it is the only version of this figure that should ever decide a translation budget.
Sources: TweetTaglish, ACL Anthology 2022 • FilBench, EMNLP 2025, arXiv 2508.03523 • Benchmarking Open-Source Large Language Models on Code-Switched Tagalog-English Queries, Journal of Advances in Information Technology, February 2025 • EF English Proficiency Index 2024
Created by Arfadia • arfadia.com/blog
Where the vocabulary actually lives
Before touching Search Console, it helps to know why Philippine consumer vocabulary is unusually hard to capture from search data alone. The reason is platform reach.
Facebook reached 94.9 percent of Philippine internet users aged 16 and over on a monthly basis, per DataReportal figures. Messenger 90.6 percent. YouTube 85 percent. TikTok 82.2 percent. Instagram 71.2 percent, at 29.0 million users aged 18 and above. Five platforms, and the top four all sit above eighty percent of the online population.
Referral traffic concentrates even harder. Facebook alone accounts for 84.17 percent of social-referral web traffic to third-party sites in the Philippines. Not a plurality. A supermajority, from one platform.
Worth flagging a measurement wrinkle here, because it is the kind of thing that turns into an argument later. One research source puts Facebook's ad reach at 97.7 percent of the Philippine internet user base as of October 2025, against the 94.9 percent monthly-reach figure above. Those are not the same measurement. Advertising reach and monthly active usage are different constructs with different denominators, and the sensible thing is to report both with their definitions rather than pick the more impressive one.
Why does any of this matter for a language question? Because consumer product vocabulary forms in those spaces first and reaches search second. TikTok Shop in the Philippines generated USD 2 billion in gross merchandise value in the fourth quarter of 2025, roughly 35 percent of the platform's combined quarterly sales, overtaking Lazada to become the second-largest platform by quarterly volume. Combined top-platform GMV across Shopee, Lazada and TikTok Shop reached USD 22 billion for the full year 2025, up 15 percent. Beauty and personal care led TikTok Shop at 28 percent of that value.
Live-selling, comment threads and creator captions are where a category gets named in everyday language. By the time that vocabulary shows up in a Search Console export, it has already been in circulation for months. So social comments are a genuinely useful vocabulary-discovery surface, and a genuinely poor volume-estimation surface. Two different jobs, and conflating them is how a keyword plan ends up chasing terms nobody searches.
How to measure your own mix in an afternoon
Seven steps. None of them requires anything you do not already have access to.
Export Search Console queries with dimensions attached. Not the top thousand by clicks. Everything available, with landing page, device, country and date range preserved. The tail is where mixed-language queries live, and a top-queries view is precisely the thing that removes the tail.
Pull Google Ads search terms if you run paid. Paid search terms surface phrasing that organic impressions never expose, particularly the conversational forms people abandon quickly.
Add internal site search and support transcripts. Your own site search is a query log you fully own, unsampled, with no privacy intermediary in the middle. Customer service transcripts are the closest thing to hearing a customer describe the problem in their own words before anyone has trained them into your vocabulary.
Label at token level, not query level. This is where most attempts quietly fail. A query like "requirements sa Pag-IBIG housing loan" is not English and it is not Filipino. Labelling it as either destroys the finding you were looking for. Use four labels, English, Filipino, other regional language, and mixed, and apply them per token before summarising per query.
Cluster intent separately from language. Learn, compare, validate, locate, buy, troubleshoot, contact. Language and intent are independent axes. The interesting result is usually a cell rather than a row: informational queries running mixed while commercial ones run English, or the reverse in a particular category.
Have a native speaker check the colloquial terms. Automated labelling misreads borrowed English nouns sitting inside Filipino syntax, and it misses which of several near-synonyms people actually use for a product category. Short review. Skipping it produces a clean-looking dataset that is wrong in one specific direction, which is worse than a messy one.
Then decide what deserves its own page. A stable mixed-language intent with real volume and a conversion path may justify a dedicated page. More often the right answer is natural synonyms in existing copy, an FAQ entry, a video transcript, or a social asset. Creating a page is the most expensive response available, and it should not be the default one.
What each signal is good for
| Source | What it tells you | Known limitation |
|---|---|---|
| Search Console queries | Real organic phrasing already reaching you, by page and device | Anonymised and filtered. Low-volume queries are withheld, so the tail is understated |
| Google Ads search terms | Conversational and long-tail phrasing organic never surfaces | Shaped by your keyword targeting, so it reflects what you bid on as much as what people type |
| Internal site search | Unsampled, fully owned, shows vocabulary after the person is already on your site | Biased toward people who failed to find something through navigation |
| Support and chat transcripts | How customers describe the problem before adopting your terminology | May contain personal data. Redact before analysis and keep it out of any external tool |
| Facebook comments and community threads | Natural code-switched vocabulary at scale, on a platform reaching 94.9 percent of online Filipinos monthly | Social register differs from search register. Vocabulary discovery, never volume estimation |
| TikTok comments and live-selling chat | Where consumer category language forms first, especially in beauty and personal care | Heavily trend-driven. Terms can appear and disappear inside a quarter |
| Keyword tool volume estimates | Rough relative demand between phrasings | Mixed-language variants are frequently underreported or bucketed, so absolute numbers mislead |
Two numbers for the same behaviour, and why both are correct
A short detour, because it illustrates the whole problem this article is about.
How many Filipinos use social media to discover products? One published compilation says 60 percent of Filipinos use social media to discover new products and services. DataReportal's Digital 2026 puts brand discovery via social platforms at 41.9 percent. Same country, overlapping windows, an 18-point gap.
Neither is wrong. They define discovery differently. One counts any use of social platforms in the course of finding out about products. The other counts social platforms as the specific channel where brand discovery happened. Averaging them would produce a number that describes nothing, and yet averaging them is exactly what a summary slide tends to do.
Hold that pattern in mind when someone hands you a query-language percentage. The question is never just what the number is. It is what was counted, across which population, over which window. A precise number with an undisclosed definition is less useful than an honest range with a stated one.
The translation decision, once you have the data
Three outcomes are common. The right response differs sharply between them.
If your mix comes back overwhelmingly English, which happens often in B2B, professional services and enterprise categories, stop debating translation and spend the budget on depth in English instead. Local entity density does the localisation work that translation was supposed to do. Name the Philippine regulator, the specific cities, the local publications, the actual peso context.
If a specific consumer intent comes back consistently mixed with real volume, build for it, and have a Filipino native speaker write it. This is where the FilBench result becomes practical rather than academic. Machine-generated Filipino content is handled less reliably by models and reads as machine-generated to humans. Worst of both.
If the data is thin because organic volume is low, do not treat that as evidence of an English-only audience. It may simply mean you have never ranked for the mixed-language variants and so never saw them. Test with paid before concluding anything from an absence.
Four Readings, Four Different Responses
The mistake is not choosing wrong. It is choosing before looking.
Overwhelmingly English
Common in B2B, professional services and enterprise. Stop debating translation. Spend the budget on depth and local entity density in English: named regulators, named cities, named publications, peso context.
One intent consistently mixed
Build for it, and have a Filipino native speaker write it. Do not machine-translate. Models score 46.48 percent on Filipino text generation, and that shows up in the output.
Thin data across the board
Absence of mixed-language queries is not evidence of an English-only audience. You may simply never have ranked for those variants. Test with paid before drawing a conclusion from silence.
Split by intent, not by language
The most common real pattern. Informational and troubleshooting queries lean mixed while commercial and comparison queries lean English, or the reverse in a given category. Serve each where it actually lives.
And in every case, resist the page reflex
A new page is the most expensive response available. Natural synonyms inside existing copy, an FAQ entry, a video transcript or a social asset will serve most mixed-language intents at a fraction of the cost, without creating a near-duplicate that competes with the page you already rank with. Given Facebook reaches 94.9 percent of online Filipinos monthly and accounts for 84.17 percent of social-referral traffic, a well-written social asset is frequently the higher-return option anyway.
Sources: FilBench, EMNLP 2025 • Journal of Advances in Information Technology, February 2025, code-switched Tagalog-English RAG benchmark • DataReportal platform reach and social-referral figures • Cross-validated Philippine SEO and GEO research, August 2026
Created by Arfadia • arfadia.com/blog
Why this is worth an afternoon
The national percentage does not exist. And it would not be the right input even if it did, because a national average across every category, region and intent would tell you almost nothing about whether your particular buyers type in English.
Your own data answers the question you are actually asking. It also produces something more durable than a decision. It produces a baseline you can re-measure next quarter, which means you find out when the mix shifts rather than assuming it never does.
Same discipline our SEO service for the Philippines applies to every figure it hands over. Measured on your property, definition attached, re-runnable by you. For the AI answer side of the same question, our GEO service for the Philippines works from the same principle.
Frequently Asked Questions
What percentage of Philippine searches are in English?
Nobody knows, and no credible published study measures it. Code-switching between Tagalog and English is well documented in Philippine online-language research, including academic datasets built from social media text, but no study analyses commercial search query logs by language. Any specific percentage offered for this is an estimate, not a measurement.
Is Taglish common in search queries?
Directionally, almost certainly yes, based on the linguistic literature on Philippine online communication. The magnitude is unknown. Social-media code-switching prevalence is not a valid proxy for search-query prevalence, because a search box is a different genre with different conventions from a post or a message.
Should we translate our website into Filipino?
Measure before deciding. Export your own Search Console queries with landing page, device and location preserved, label them at token level as English, Filipino, other regional language or mixed, then cluster by intent. If your mix is overwhelmingly English, which is common in B2B and professional categories, the budget is better spent on depth and local entity density in English. If a specific consumer intent comes back consistently mixed with real volume, build for it.
Can we use AI to generate Filipino content?
Not as a substitute for a native writer. FilBench, published at EMNLP 2025, evaluated 27 models on Filipino, Tagalog and Cebuano and found the strongest scored 72.23 percent aggregate with a text-generation subscore of just 46.48 percent. A separate 2025 study of retrieval-augmented models found accuracy highest in English, lower in Taglish and lowest in pure Tagalog. Machine-generated Filipino is handled less reliably by models and reads as machine-generated to Filipino readers.
Why label queries at token level rather than per query?
Because mixed queries are the point. A phrase combining an English noun with Filipino function words is neither English nor Filipino, and labelling the whole query as one or the other erases exactly the finding you are looking for. Label each token, then summarise per query into four buckets: English, Filipino, other regional language, and mixed.
Should we mine social comments for keyword research?
Yes for vocabulary, no for volume. Facebook reaches 94.9 percent of Philippine internet users aged 16 and over monthly and accounts for 84.17 percent of social-referral web traffic, while TikTok reaches 82.2 percent and is where consumer category language often forms first, particularly in beauty and personal care. That makes social comments an excellent discovery surface for how people actually name things. It makes them a poor basis for volume estimates, because social register and search register differ and trend-driven terms can appear and vanish inside a quarter.
Does a mixed-language query always deserve its own page?
No, and this is the most common overcorrection. A new page is the most expensive available response and creates a near-duplicate risk against pages you already rank with. Natural synonyms in existing copy, an FAQ entry, a video transcript or a social asset will serve most mixed-language intents adequately. Reserve a dedicated page for a stable intent with real volume and a distinct conversion path.
Why do published figures on Philippine social discovery disagree so much?
Definitions. One compilation reports that 60 percent of Filipinos use social media to discover new products and services, while DataReportal's Digital 2026 puts brand discovery via social platforms at 41.9 percent. Both describe real measurements of overlapping behaviour with different boundaries around what counts as discovery. The correct handling is to report both with their definitions rather than averaging them into a figure that describes neither.
Sources & References:
- Query-language split for Philippine commercial search is UNAVAILABLE. No study analysing Philippine search query logs by language was located across four independent research passes conducted for this article in August 2026.
- Code-switching evidence: TweetTaglish, a dataset for investigating Tagalog-English code-switching, ACL Anthology, 2022. Academic work on Filipino-English code-switching in online academic discourse documenting intensive within-turn switching. Both measure social and academic online writing, not search queries.
- FilBench, published at EMNLP 2025 Main, arXiv 2508.03523, evaluating 27 large language models on Filipino, Tagalog and Cebuano. Best model aggregate 72.23 percent; text-generation subscore 46.48 percent. Peer-reviewed.
- Benchmarking Open-Source Large Language Models on Code-Switched Tagalog-English Queries, Journal of Advances in Information Technology, February 2025. Reported performance highest in English, followed by Taglish, lowest in pure Tagalog, in retrieval-augmented generation tasks.
- EF English Proficiency Index 2024: the Philippines ranked 22nd of 116 countries on a score of 570, described as high proficiency, second in Asia after Singapore.
- Platform monthly reach among Philippine internet users aged 16 and over: Facebook 94.9 percent, Messenger 90.6 percent, YouTube 85 percent, TikTok 82.2 percent, Instagram 71.2 percent at 29.0 million users aged 18 and above. DataReportal. Facebook accounts for 84.17 percent of social-referral web traffic to third-party sites in the Philippines, per DataReportal. Separately, one research source puts Facebook advertising reach at 97.7 percent of the Philippine internet user base as of October 2025. Advertising reach and monthly active reach are different constructs with different denominators and are reported here separately rather than reconciled.
- Social commerce scale: TikTok Shop generated USD 2 billion GMV in the Philippines in Q4 2025, approximately 35 percent of the platform's combined quarterly sales, overtaking Lazada as second-largest platform by quarterly volume. Combined Shopee, Lazada and TikTok Shop GMV reached USD 22 billion for FY 2025, up 15 percent from USD 19 billion in FY 2024. Beauty and personal care led TikTok Shop at 28 percent of GMV. Cube Tradewinds Q4 2025 dataset, a commercial estimator, so figures are estimates rather than audited.
- Contradiction retained rather than resolved: one compilation reports 60 percent of Filipinos using social media to discover new products and services, while DataReportal Digital 2026 puts brand discovery via social platforms at 41.9 percent. Different definitions of discovery. Both reported, neither averaged.
- Search Console query data is anonymised and filtered, with low-volume queries withheld. Any query-language analysis built on it understates the long tail and should be read accordingly.
- Support and chat transcripts may contain personal data. Under Republic Act No. 10173, redaction before analysis and exclusion from external tools is the appropriate default for this workflow.