Translation exists. Complete localisation often does not.
The clearest result from the crawl is not just that directly verified multilingual content was a minority among the major websites we could classify. It is that publishing another language is only the first layer of implementation. URL structure, language annotations, metadata, crawlability and page-level consistency frequently lag behind the translated body content.
11 findings from 25,005 major websites
- Confirmed multilingual content was 26.3% of classifiable successes. The static HTML headline is 5,845 confirmed sites out of 22,185 non-uncertain successes.
- Multilingual implementation was more common among the highest-ranked sites. Confirmed multilingual content was 40.1% in the classifiable Tranco top 1,000 cohort versus 22.0% in the 25k+ cohort.
- Subdirectories were the most common observed URL architecture. They accounted for 51.6% of confirmed multilingual sites, but prevalence is not evidence that one architecture is universally “best.”
- Nearly one-third of confirmed multilingual sites had no detected hreflang. 1,769 of 5,845 confirmed sites published alternate-language main content without a detected hreflang annotation.
- When hreflang existed, most implementations passed the core checks we measured. 72.2% of confirmed hreflang users passed all applicable checks excluding x-default.
- Search-facing metadata was the clearest implementation gap. Only 11.6% of confirmed sites received full credit for title + meta localisation under the GlotEO scoring methodology.
- Mixed-language leftovers were common. 39.4% of confirmed multilingual sites triggered the study’s mixed-language deduction.
- A crawlable language selector was detected on one-third of confirmed sites. 33.0% exposed selector links; 65.8% had no selector detected by this crawl.
- Automatic locale redirects were present on 15.5% of confirmed multilingual sites. The research records the behaviour; it does not imply every redirect produces the same user or search outcome.
- The targeted JavaScript pass changed classifications without changing the one-decimal headline. Confirmed rose by 117 and uncertain fell by 523, while the JS-adjusted confirmed rate remained 26.3%.
- A stricter Yoruba/Tagalog alternate-page rule produced a 25.5% sensitivity estimate. Combining that overlay with the JS adjustment produced 25.4%. These are targeted classification scenarios, not replacements for the frozen 26.3% result, and not a statistical confidence interval.
How multilingual are major websites?
Among the classifiable major websites successfully analysed in GlotEO’s Tranco-based sample, 26.3% (5,845 / 22,185) had directly verified multilingual content in the static HTML freeze. “Confirmed” required more than a language dropdown, a lang attribute or an hreflang tag: the crawler had to fetch at least one alternate URL whose main-content language differed from the homepage, with at least 200 letters of main text on both sides.
That definition deliberately makes confirmation harder than simply detecting international SEO markup. It separates “the site advertises another locale” from “we actually fetched substantive content in another language.”
of classifiable successes were directly confirmed multilingual in the static HTML freeze.
5,845 / 22,185 · Uncertain rows excluded from this denominator · GlotEO Research 2026The denominator matters. Across all 25,005 successful analyses, 2,820 sites were uncertain. GlotEO kept them as their own class rather than forcing them into monolingual. On the all-success denominator, confirmed sites were 23.4% (5,845 / 25,005), probable sites 14.4%, monolingual 50.9%, and uncertain 11.3%.
The secondary sensitivity figure is 42.6% (9,448 / 22,185) when confirmed and probable are combined. It is useful for understanding sites with strong multilingual discovery signals but no qualifying alternate fetched; it is not the headline adoption rate.
Confirmed multilingual adoption falls through the rank cohorts
Static HTML v1.0 · classifiable successes within each Tranco cohort
The most prominent sites were much more likely to be multilingual
Confirmed multilingual content was 40.1% (192 / 479) among classifiable sites in the Tranco top 1,000 cohort. It fell to 30.4% in ranks 1,001–10,000, 27.6% in ranks 10,001–25,000, and 22.0% (1,845 / 8,394) in the 25k+ cohort.
The pattern survives an important sanity check: successful fetching actually rose as rank fell, from 53.6% in the top 1,000 attempted domains to 63.3% in the 25k+ attempted cohort. In other words, the adoption gradient is not explained by the crawler simply failing more often on lower-ranked domains.
| Tranco cohort | Classifiable sites | Confirmed multilingual | Confirmed + probable | Fetch success |
|---|---|---|---|---|
| Top 1,000 | 479 | 40.1% (192/479) | 62.4% | 53.6% |
| 1,001–10,000 | 4,958 | 30.4% (1,505/4,958) | 50.0% | 60.6% |
| 10,001–25,000 | 8,354 | 27.6% (2,303/8,354) | 44.5% | 61.6% |
| 25k+ | 8,394 | 22.0% (1,845/8,394) | 35.2% | 63.3% |
The 25k+ sample is cap-truncated: 2,351 candidate domains remained when the successful-analysis cap was reached. Treat this as a cohort result for the analysed sample, not a finished census of the long tail.
What sites advertise is not always what a crawler can verify
For confirmed multilingual sites, the study kept two language concepts separate. Observed languages came from the main content of pages that were actually fetched. Advertised languages came from hreflang annotations. That distinction is important because a tag can advertise a language that the limited crawl never reaches, while a crawler can also discover alternate-language pages that are not declared in hreflang.
Among the 5,845 confirmed multilingual sites, the median number of observed languages was 3 and the mean was 3.13. Almost half—46.8%—showed exactly two observed languages. Another 16.1% showed three, 14.1% showed four, and 23.0% reached the instrument's maximum of five observed languages.
The advertised-versus-observed comparison shows why a single signal is not enough. Advertised language count was greater than observed count on 38.7% (2,262 / 5,845) of confirmed sites. Observed count was greater than advertised on 34.1% (1,991 / 5,845). The two positive counts were equal on 27.2%.
In addition, 30.5% (1,780 / 5,845) of confirmed sites had zero advertised languages in the study's hreflang-derived field, despite the crawler having verified alternate-language main content. That is another way of seeing the gap between publishing a translation and explicitly describing the locale relationship in machine-readable markup.
Advertised language count vs observed language count
Confirmed multilingual sites · n=5,845 · mutually exclusive count comparison
One further measure needs especially careful wording: 40.4% of confirmed sites had at least one advertised language code that did not appear in the fetched-body language set. This is not evidence that those sites falsely advertise translations. A page outside the five-page crawl could easily contain the advertised language. The useful conclusion is narrower: hreflang declarations and a bounded crawl of substantive page content describe different parts of the system.
For website operators, that suggests two separate QA questions. First: can a crawler reach substantive pages in every language you intend to publish? Second: do your machine-readable locale annotations accurately map those pages to one another? Passing one test does not guarantee passing the other.
Subdirectories dominate, but multilingual URL architecture is still diverse
Among the 5,845 confirmed multilingual sites, 51.6% (3,015) used a subdirectory as the primary observed URL pattern. Hybrid implementations were next at 17.9%, followed by unknown at 13.0%, subdomains at 8.6% and ccTLDs at 8.3%. Parameter-based localisation was rare at 0.5%, and same-URL localisation was rarer still at 0.1%.
This is a prevalence result, not an SEO winner table. A subdirectory is common in the sample; the study did not test whether subdirectories outperform subdomains or ccTLDs.
URL architecture among confirmed multilingual sites
Primary mutually exclusive pattern · n=5,845
What “hybrid” means
A site is coded hybrid when the alternate URLs fetched use two or more schemes—such as ccTLD plus subdomain, or subdomain plus locale paths. Hybrid is not an error label.
Why different URLs matter
Google’s current guidance recommends separate URLs for different language versions and warns that locale-adaptive content on one URL may not all be crawled. That is external guidance, not a conclusion from GlotEO’s measurements.
External guidance: Google Search Central — Managing multi-regional and multilingual sites ↗.
The hreflang gap: translations can exist without the machine-readable map
Hreflang was present on 69.7% (4,076 / 5,845) of confirmed multilingual sites. That means 30.3% (1,769 / 5,845) had alternate-language main content that GlotEO actually fetched, but no hreflang was detected through HTML, HTTP Link headers or sitemap XHTML.
This distinction matters because hreflang is an annotation, not a translation detector. Its presence does not prove that substantive alternate-language content exists, and its absence does not prove that a translation is missing. The study measured both independently.
of confirmed multilingual sites had no detected hreflang.
1,769 / 5,845 confirmed multilingual sites · Static HTML v1.0When hreflang did exist on confirmed sites, most of the technical checks were relatively strong. Of the 4,076 confirmed sites with hreflang, 86.8% used valid language codes, 98.2% used absolute URLs, 91.1% were self-referential, 96.0% placed annotations in the HTML head where relevant, and 86.6% had reciprocal annotations among the 4,006 cases where the check applied.
Overall, 72.2% (2,944 / 4,076) passed all applicable checks measured by the study excluding x-default. x-default appeared on 65.2% of confirmed hreflang users, but its absence is treated here as a coverage choice—not an automatic specification failure.
Hreflang quality among confirmed sites that use it
Percent passing each measured check · denominators vary where noted
External guidance: Google Search Central — Tell Google about localized versions ↗.
The biggest weaknesses appeared after the body was translated
Once a site is confirmed multilingual, the remaining question is implementation completeness. GlotEO’s component analysis shows that some structural signals were common—while search-facing metadata, structured data and sitemap alternates were much less complete.
Title + meta localisation
11.6% (677 / 5,845) received full credit under the study’s title + meta localisation component. A further 2,385 were partial. This was the weakest high-weight component in the implementation index.
Mixed-language leftovers
39.4% (2,302 / 5,845) triggered the mixed-language deduction, meaning the crawler found meaningful strings in another language in navigation, footer or other surrounding page content under the study rules.
HTML language match
86.4% (5,048 / 5,845) received full credit for the HTML language declaration matching detected content. Google says it determines language from visible content rather than relying on the lang attribute.
Automatic locale redirects
15.5% (908 / 5,845) exhibited the automatic-locale redirect behaviour captured by the study. Google advises against automatic language redirection that can prevent access to alternate versions.
Where confirmed multilingual implementations are strongest—and weakest
Full-credit rate for selected GMWS components · n=5,845 confirmed sites
Language navigation is another example where “we found another language” and “the interface exposes a conventional selector” diverge. A crawlable language-selector link was detected on 33.0% (1,927 / 5,845) of confirmed sites, a JS-only selector on 1.3%, and no selector on 65.8%. The last figure should not be interpreted as a UX failure rate: the crawl could confirm alternate content through other routes.
Selectors are useful, but they are not proof of localisation
A language selector is a navigation mechanism. That sounds obvious, but it is an important research distinction. A dropdown can list ten languages even when the crawler cannot verify substantive alternate content, while a multilingual site may expose alternate pages through navigation, sitemaps or locale URL patterns without presenting a conventional selector in the sampled HTML. This is why the classifier never used a selector alone as confirmation.
The same principle applies to hreflang, ccTLDs and html lang. Each is a useful signal, but none is the content itself. The study's stronger confirmation standard—fetching substantial alternate-language main content—was designed to avoid turning markup prevalence into an inflated multilingual-adoption number.
Automatic redirects can hide the alternative path
Automatic locale redirects were detected on 15.5% of confirmed sites. GlotEO did not attempt to defeat those redirects by spoofing language preferences, and the crawler sent no Accept-Language header. This is important because a site that chooses a locale based on request context can present a different crawl surface from a site where every language has an independently reachable URL.
Google separately recommends letting users switch between language versions and warns that automatic redirection can prevent users and search engines from seeing every version. The research does not measure the ranking effect of redirects; it records how often the behaviour appeared in the confirmed multilingual sample.
Structured data and sitemaps are still lightly localised
Only 19.2% (1,125 / 5,845) of confirmed sites received full credit for the study's inLanguage structured-data component, and only 11.9% (693 / 5,845) received full credit for sitemap XHTML alternates. These are not universal requirements for every multilingual site, so low adoption should not be translated directly into an “error rate.” They are better understood as additional machine-readable layers that relatively few confirmed sites in the sample used completely.
Taken together, these results show why multilingual audits are easy to oversimplify. A team can correctly translate its main content, publish clean locale URLs and maintain strong canonicals while still leaving gaps in metadata, sitemaps, structured data or language navigation. Conversely, a site can have excellent markup while the translated page itself remains incomplete. A useful audit needs to inspect both the content and the surrounding implementation.
The GlotEO Multilingual Web Score shows a wide implementation spread
The GlotEO Multilingual Web Score (GMWS) is a 0–100 implementation-completeness index calculated only for the 5,845 confirmed multilingual sites. It is not a Google ranking factor, not an E-E-A-T score and not a score for the web as a whole.
The median confirmed site scored 75, with a mean of 67.1, a 25th percentile of 50 and a 75th percentile of 83. The distribution has a meaningful left tail: 15.0% of confirmed sites scored 20–39, even though another 36.6% scored 80–100.
GMWS distribution among confirmed multilingual sites
Implementation completeness index · n=5,845 · not a ranking factor
The score is useful because its components are published rather than hidden. Confirmed alternate content is 100% by definition. Crawlable distinct URLs received full credit on 86.4% of confirmed sites, canonical-vs-hreflang consistency on 82.3%, HTML language match on 86.4%, and slug translation on 69.0%. By contrast, title + meta localisation was 11.6%, sitemap XHTML alternates 11.9%, and inLanguage structured data 19.2%.
JavaScript changed hundreds of classifications, but not the one-decimal headline
The static HTML freeze deliberately did not render JavaScript. It identified 889 successful sites as js_required: 768 uncertain and 121 probable. None were confirmed or monolingual in the static freeze because their script-heavy shells prevented the normal classifier from obtaining enough substantive text to make those calls.
A targeted Chromium pass rendered those 889 sites only. It succeeded on 858 and kept the original class for 31 render failures. Among the 768 previously uncertain sites, 66 became confirmed, 48 probable, 412 monolingual and 242 remained uncertain. Among the 121 probable sites, 51 became confirmed, 67 stayed probable and 3 became uncertain.
After overlaying only those 889 results onto the unchanged 24,116 other successful rows, confirmed rose from 5,845 to 5,962 (+117) and uncertain fell from 2,820 to 2,297 (−523). Yet the confirmed-only rate among non-uncertain successes remained 26.3% to one decimal: 5,962 / 22,708.
Static freeze vs targeted JavaScript adjustment
Same 25,005 successful domains; only 889 js_required rows were re-evaluated
How robust is the 26.3% estimate?
The primary result remains the frozen static-HTML estimate: 26.3% (5,845 / 22,185) confirmed multilingual among classifiable successful analyses. We subsequently tested two parts of the methodology most likely to affect classification: JavaScript-dependent sites, and a set of classifications that depended on Yoruba (yo) or Tagalog (tl) language detection.
Neither targeted analysis changed the central conclusion. The JavaScript-adjusted result remained 26.3% (5,962 / 22,708). Applying a stricter alternate-page rule to Yoruba/Tagalog-dependent classifications produced a sensitivity estimate of 25.5% (5,647 / 22,165). Combining both targeted adjustments produced 25.4% (5,764 / 22,688).
These are sensitivity analyses, not replacements for the frozen primary result. They are also not four separate crawls: the static freeze is unchanged, and the two overlays re-evaluate only the rows they were designed to test.
Targeted robustness checks did not change the central conclusion
The confirmed multilingual rate remained close to one quarter of classifiable sites under each targeted alternative analysis.
Primary frozen result
Static HTML v1.0
26.3%
5,845 confirmed · 22,185 classifiable
| Analysis | Confirmed | Classifiable denominator | Confirmed rate | Status |
|---|---|---|---|---|
| Static HTML v1.0 | 5,845 | 22,185 | 26.3% | Primary frozen result |
| Targeted JS adjustment | 5,962 | 22,708 | 26.3% | Robustness analysis |
| Yoruba / Tagalog sensitivity | 5,647 | 22,165 | 25.5% | Sensitivity analysis |
| JS + Yoruba/Tagalog sensitivity | 5,764 | 22,688 | 25.4% | Combined sensitivity |
Inspect the underlying sensitivity-analysis data →
What changed under the stricter rule?
The language-ID overlay looked only at 242 rows whose multilingual class depended on lingua codes yo (Yoruba) or tl (Tagalog) as the only differing detected language: 228 freeze-confirmed, plus 14 probable. It used saved HTML. Zero live fetches. It was not a recrawl of the 25,005-site dataset. wikipedia.org was excluded from the 242-row dependency set because its confirmed multilingual classification did not depend on the tl homepage-language detection.
Under that overlay, confirmed moves from 5,845 to 5,647 (−198). The corresponding classifiable denominator is 22,165, so the sensitivity estimate is 25.5% (5,647 / 22,165). That figure is not the headline.
After stripping repeated site chrome, yo or tl still appeared in observed languages on 217 / 242 rows. Only 25 / 242 lost the language code entirely. Therefore most classification changes did not result from the language detector simply retracting Yoruba or Tagalog. Most came from a tighter alternate-page rule: a yo or tl alternate was counted as multilingual evidence only when its URL represented a genuine locale page, language edition or market edition—rather than simply being a different URL whose text was detected as another language. 30 still-confirmed rows include home-side yo/tl that survived chrome-strip.
This suggests that the main issue was not simply that the language detector incorrectly called pages Yoruba or Tagalog. In many cases the detected language remained after boilerplate was removed. The larger change came from asking a stricter question: was the alternate URL genuinely a localized edition of the site? Alternate-page equivalence proved to be an important source of classification uncertainty alongside language identification itself. That reading is limited to this targeted analysis; it is not a global error rate for the classifier.
An independent AI-assisted review, not a human gold-standard validation study, found that the arithmetic holds. The sensitivity figure does not replace the frozen headline.
Known classification examples
These examples show what the sensitivity analysis actually changed. They are methodology illustrations for specific domains under this research instrument—not judgements of those organisations’ global language programmes.
Homepage language-ID miss, classification unchanged
wikipedia.org
The homepage language detector classified the homepage as tl / Tagalog rather than en / English. That is a genuine homepage language-ID miss. Wikipedia remains confirmed multilingual: the crawl independently identified Japanese, German and Russian editions. Distinguish a language-detection error from a multilingual classification error. This case is the first, not the second.
Confirmed → monolingual
office.com
Under the stricter alternate-page rule, office.com moves from confirmed to monolingual because worldwide.aspx was not treated as a genuine localized or language-edition page. The example shows why page equivalence matters: a different URL is not automatically a locale edition.
Several classifications became more conservative
Amazon domains
Under the stricter rule: amazon.it → probable; amazon.com.au → probable; amazon.ae → monolingual; amazon.sa → monolingual. These labels refer only to what could be directly established for that specific domain under this methodology. They do not imply that Amazon is monolingual globally.
What this study can—and cannot—tell you
Large web studies become less useful when their caveats are treated as legal fine print. Several limits materially affect how these numbers should be used, so they belong in the interpretation rather than being hidden at the bottom of the page.
This is a study of major Tranco-ranked sites, not every website
The source population was Tranco list 46W9X. Tranco is designed to provide a research-oriented ranking of popular domains, so the sample is intentionally tilted toward prominent websites rather than the millions of small, inactive or rarely visited domains elsewhere on the internet. The correct claim is about the major websites successfully analysed in this sample, not “26.3% of the web.”
Exclusions are infrastructure outcomes, not monolingual sites
The crawler attempted 40,457 unique domains and successfully analysed 25,005. The remaining 15,452 were excluded for reasons including DNS failures, robots responses, HTTP errors, access restrictions and CAPTCHAs. Those domains were never converted into monolingual observations. Treating inaccessible sites as monolingual would materially bias the result.
“Confirmed” measures discoverable alternate content, not translation quality
A confirmed site passed a deliberately mechanical content test: the crawl fetched substantive main content in a language different from the homepage. The study did not grade linguistic fluency, cultural adaptation, terminology quality, conversion performance or human translation accuracy. A site can be technically confirmed multilingual and still contain poor translations; it can also provide excellent localisation in areas the crawl did not visit.
The crawl is bounded to five content pages per domain
The page cap keeps a 25,000-domain study computationally tractable and prevents large sites from dominating the crawl. The trade-off is incomplete corpus coverage. Observed languages, mixed-language findings, metadata checks and other page-level measures describe what appeared within the bounded sample of pages—not necessarily every URL the domain serves.
JavaScript was tested selectively, not universally
The static freeze used no browser rendering. The follow-up Chromium pass targeted the 889 successful rows explicitly flagged as requiring JavaScript. This was the right place to test the obvious classification hole, but it does not establish what a full browser-rendered recrawl of all 25,005 successes would produce.
Language identification and alternate-page equivalence can produce edge cases
Language identification and alternate-page equivalence are automated classifications and can produce edge cases, particularly on short, boilerplate-heavy or unconventional pages. A targeted Yoruba/Tagalog sensitivity analysis reduced the confirmed estimate from 26.3% to 25.5% under stricter alternate-page rules without changing the overall conclusion. The language-ID sensitivity review was AI-assisted rather than a human gold-standard validation study.
The Multilingual Web Score is a research framework, not a search-engine score
GMWS makes implementation differences easier to compare by publishing a transparent weighting of twelve components plus two deductions. The weights belong to GlotEO's research design. Google did not design or endorse them, and no relationship between the score and rankings, traffic or conversion was tested here.
Benchmark your own multilingual implementation
The research is most useful when it becomes an audit prompt. The checklist below lets you compare common implementation signals with the confirmed multilingual sample. Checking a box does not create a ranking score; it simply records which implementation layers your site has addressed.
Multilingual implementation checklist
Tick what your site has implemented. Benchmarks are static HTML v1.0 unless noted.
What teams should do differently
1. Treat translation and discoverability as separate workstreams. A page can exist in another language without clearly advertising its relationship to sibling pages. That is exactly what the 30.3% hreflang gap demonstrates. For implementation context on localized URLs and hreflang, see multilingual SEO.
2. QA the page shell, not just the article or product description. Mixed-language leftovers were detected on 39.4% of confirmed sites. Navigation, footer strings, forms and templates can undermine an otherwise complete translation.
3. Include metadata in the localisation workflow. The 11.6% full-credit rate for title + meta was the most obvious high-weight implementation gap in the score. Search-facing metadata should not be an afterthought.
4. Give each locale a stable destination. The research found distinct crawlable URLs to be common among confirmed sites, and Google separately recommends distinct URLs rather than relying only on locale-adaptive content.
5. Make uncertainty visible in audits. This study did not convert sites it could not classify into “monolingual.” Your own audits should be equally careful around JavaScript shells, geo-adaptation, authentication and crawler access.
How the study was run
This is an observational study of major websites in a Tranco-ranked sample, not a random sample of every registered domain and not a customer dataset.
Classification rules
Confirmed: the crawler fetched an alternate URL whose main-content language differed from the homepage, with at least 200 letters of main content on both sides.
Probable: the site exposed strong multilingual discovery signals—such as hreflang, locale-shaped URLs or a selector pointing to another path—but the crawl did not fetch a qualifying alternate.
Monolingual: the homepage was analysed, no qualifying alternate was fetched and no strong multilingual discovery cluster was detected.
Uncertain: evidence was contradictory or insufficient. Uncertain was kept visible and was never silently recoded as monolingual.
What is frozen, and what is an overlay
Static HTML v1.0
Primary frozen research dataset. The headline remains 26.3% (5,845 / 22,185) confirmed among classifiable successes.
Targeted JS overlay
Robustness analysis of 889 js_required successes only. The remaining 24,116 successful rows were unchanged. This is not a complete recrawl.
Yoruba/Tagalog sensitivity overlay
Alternative classification sensitivity on 242 saved-HTML rows whose class depended on yo or tl. Zero live fetches. Not a recrawl.
Combined targeted display
JavaScript plus language-ID sensitivity shown together as 25.4% (5,764 / 22,688). A display of two overlays, not a fourth crawl.
Methodology details and limitations
The unit of analysis was the Tranco pay-level domain. The crawler attempted HTTPS first, then HTTP on TLS failure. Robots rules were honoured, including the study’s handling of robots 5xx responses. It did not bypass authentication, CAPTCHAs or other access controls. The static v1.0 freeze did not execute JavaScript.
The study attempted 40,457 unique domains and successfully analysed 25,005, a 61.8% fetch/analysis success rate. The 15,452 exclusions are infrastructure/access outcomes, not monolingual websites. Major exclusion categories included DNS failures, robots 5xx responses, HTTP 4xx responses, robots disallow and CAPTCHAs.
The homepage + four-content-URL cap means observed language count is not a census of every locale a site may offer. Five observed languages is therefore the maximum visible under this instrument. Likewise, advertised hreflang languages that were not observed in fetched bodies are not proof that those pages do not exist elsewhere on the site.
The 25k+ cohort was truncated when the study reached its successful-analysis cap; 2,351 candidate domains remained. The sample is rank-biased and English-home heavy. Results should be described as properties of the analysed Tranco major-site sample, not “the web.”
The JavaScript-adjusted v1.1 result is a targeted robustness overlay on 889 js_required successes. It is not a complete Chromium recrawl of all 25,005 sites. The Yoruba/Tagalog overlay is a second targeted analysis of 242 saved-HTML rows whose class depended on yo or tl as the only differing language. Neither overlay replaces the frozen static-HTML dataset. Future validation work will use a blinded human-reviewed sample to estimate classifier precision and recall across multilingual status, page language and alternate-page equivalence.
Tranco source list: 46W9X ↗. Google links elsewhere on this page are labelled external guidance and are not mixed into the GlotEO result tables.
Use the data
Two workbooks are available. The frozen primary dataset is the source for the 26.3% headline and the downstream implementation statistics. A separate derived workbook supports the targeted JavaScript and Yoruba/Tagalog overlays. It does not replace the freeze.
Primary frozen dataset
The original frozen dataset underlying the primary State of the Multilingual Web 2026 findings, including the 26.3% confirmed multilingual estimate.
Open primary frozen dataset ↗File: State-of-the-Multilingual-Web-2026-freeze.xlsx
Robustness & sensitivity workbook
Derived analysis supporting the targeted JavaScript and Yoruba/Tagalog classification sensitivity checks. These overlays supplement the frozen dataset and do not replace the primary research results.
Open robustness & sensitivity data ↗File: State-of-the-Multilingual-Web-2026-derived-overlays.xlsx
Short citation
GlotEO. The State of the Multilingual Web 2026. GlotEO Research, 2026.
Recommended citation
GlotEO. “The State of the Multilingual Web 2026.” GlotEO Research, 2026. Static HTML freeze as of 28 August 2026; targeted JavaScript and language-ID sensitivity analyses as of 29 August 2026. Headline N=25,005 successfully analysed pay-level domains from Tranco list 46W9X. Frozen primary estimate: 26.3% (5,845 / 22,185) confirmed multilingual among classifiable successes.
Research integrity
Static HTML v1.0 headline: 26.3% (5,845 / 22,185) confirmed multilingual among non-uncertain successes. Targeted JS overlay: 26.3% (5,962 / 22,708) after re-evaluating the 889 js_required successes only. Yoruba/Tagalog sensitivity overlay: 25.5% (5,647 / 22,165). Combined targeted display: 25.4% (5,764 / 22,688). The 25.4%–26.3% spread is not a confidence interval.
When quoting a percentage, preserve its denominator and identify whether the figure is static v1.0, JS-adjusted, or a Yoruba/Tagalog sensitivity overlay. The 42.6% confirmed+probable figure is a sensitivity measure, not the headline adoption rate. Downstream implementation statistics on this page remain tied to the frozen 5,845 confirmed sites.
Research source: GlotEO original crawl and frozen derived dataset. External guidance: Google Search Central pages linked above. The external guidance is context only and is not incorporated into GlotEO’s original percentage calculations.
Published 29 August 2026. Research update: 29 August 2026 — added targeted Yoruba/Tagalog classification sensitivity analysis. The original static-HTML freeze and primary 26.3% result are unchanged.