Original GlotEO Research · 2026

The State of the Multilingual Web 2026

GlotEO analysed 25,005 major websites to measure how often they genuinely publish content in multiple languages—and how complete those multilingual implementations are once translation exists.

Freeze 28 Aug 2026 JS overlay 29 Aug 2026 Language-ID sensitivity 29 Aug 2026 Tranco 46W9X Homepage + up to 4 content URLs
25,005
successfully analysed major websites
Headline research N
26.3%
confirmed multilingual among classifiable successes
5,845 / 22,185
30.3%
of confirmed multilingual sites had no detected hreflang
1,769 / 5,845
39.4%
showed mixed-language leftovers under the study rules
2,302 / 5,845
11.6%
received full title + meta localisation credit
677 / 5,845
Executive summary

Translation exists. Complete localisation often does not.

The clearest result from the crawl is not just that directly verified multilingual content was a minority among the major websites we could classify. It is that publishing another language is only the first layer of implementation. URL structure, language annotations, metadata, crawlability and page-level consistency frequently lag behind the translated body content.

11 findings from 25,005 major websites

  1. Confirmed multilingual content was 26.3% of classifiable successes. The static HTML headline is 5,845 confirmed sites out of 22,185 non-uncertain successes.
  2. Multilingual implementation was more common among the highest-ranked sites. Confirmed multilingual content was 40.1% in the classifiable Tranco top 1,000 cohort versus 22.0% in the 25k+ cohort.
  3. Subdirectories were the most common observed URL architecture. They accounted for 51.6% of confirmed multilingual sites, but prevalence is not evidence that one architecture is universally “best.”
  4. Nearly one-third of confirmed multilingual sites had no detected hreflang. 1,769 of 5,845 confirmed sites published alternate-language main content without a detected hreflang annotation.
  5. When hreflang existed, most implementations passed the core checks we measured. 72.2% of confirmed hreflang users passed all applicable checks excluding x-default.
  6. Search-facing metadata was the clearest implementation gap. Only 11.6% of confirmed sites received full credit for title + meta localisation under the GlotEO scoring methodology.
  7. Mixed-language leftovers were common. 39.4% of confirmed multilingual sites triggered the study’s mixed-language deduction.
  8. A crawlable language selector was detected on one-third of confirmed sites. 33.0% exposed selector links; 65.8% had no selector detected by this crawl.
  9. Automatic locale redirects were present on 15.5% of confirmed multilingual sites. The research records the behaviour; it does not imply every redirect produces the same user or search outcome.
  10. The targeted JavaScript pass changed classifications without changing the one-decimal headline. Confirmed rose by 117 and uncertain fell by 523, while the JS-adjusted confirmed rate remained 26.3%.
  11. A stricter Yoruba/Tagalog alternate-page rule produced a 25.5% sensitivity estimate. Combining that overlay with the JS adjustment produced 25.4%. These are targeted classification scenarios, not replacements for the frozen 26.3% result, and not a statistical confidence interval.
The core thesis: a multilingual website is not simply a translated body of text. It is a stack of content, URLs, language relationships, metadata, crawlability and user-interface decisions. The research measures how often those layers appear together.
01 · Adoption

How multilingual are major websites?

Among the classifiable major websites successfully analysed in GlotEO’s Tranco-based sample, 26.3% (5,845 / 22,185) had directly verified multilingual content in the static HTML freeze. “Confirmed” required more than a language dropdown, a lang attribute or an hreflang tag: the crawler had to fetch at least one alternate URL whose main-content language differed from the homepage, with at least 200 letters of main text on both sides.

That definition deliberately makes confirmation harder than simply detecting international SEO markup. It separates “the site advertises another locale” from “we actually fetched substantive content in another language.”

26.3%

of classifiable successes were directly confirmed multilingual in the static HTML freeze.

5,845 / 22,185 · Uncertain rows excluded from this denominator · GlotEO Research 2026

The denominator matters. Across all 25,005 successful analyses, 2,820 sites were uncertain. GlotEO kept them as their own class rather than forcing them into monolingual. On the all-success denominator, confirmed sites were 23.4% (5,845 / 25,005), probable sites 14.4%, monolingual 50.9%, and uncertain 11.3%.

The secondary sensitivity figure is 42.6% (9,448 / 22,185) when confirmed and probable are combined. It is useful for understanding sites with strong multilingual discovery signals but no qualifying alternate fetched; it is not the headline adoption rate.

Confirmed multilingual adoption falls through the rank cohorts

Static HTML v1.0 · classifiable successes within each Tranco cohort

Source: GlotEO State of the Multilingual Web 2026 freezeObservational; not a causal ranking study
02 · Rank gradient

The most prominent sites were much more likely to be multilingual

Confirmed multilingual content was 40.1% (192 / 479) among classifiable sites in the Tranco top 1,000 cohort. It fell to 30.4% in ranks 1,001–10,000, 27.6% in ranks 10,001–25,000, and 22.0% (1,845 / 8,394) in the 25k+ cohort.

The pattern survives an important sanity check: successful fetching actually rose as rank fell, from 53.6% in the top 1,000 attempted domains to 63.3% in the 25k+ attempted cohort. In other words, the adoption gradient is not explained by the crawler simply failing more often on lower-ranked domains.

Do not read this as causation. The study does not show that multilingual websites rank higher, or that higher-ranked sites become multilingual because of rank. Large global brands are more likely to operate in multiple markets and may also expose locale URLs more clearly. Tranco rank is a sampling dimension, not a treatment.
Tranco cohortClassifiable sitesConfirmed multilingualConfirmed + probableFetch success
Top 1,00047940.1% (192/479)62.4%53.6%
1,001–10,0004,95830.4% (1,505/4,958)50.0%60.6%
10,001–25,0008,35427.6% (2,303/8,354)44.5%61.6%
25k+8,39422.0% (1,845/8,394)35.2%63.3%

The 25k+ sample is cap-truncated: 2,351 candidate domains remained when the successful-analysis cap was reached. Treat this as a cohort result for the analysed sample, not a finished census of the long tail.

03 · Language coverage

What sites advertise is not always what a crawler can verify

For confirmed multilingual sites, the study kept two language concepts separate. Observed languages came from the main content of pages that were actually fetched. Advertised languages came from hreflang annotations. That distinction is important because a tag can advertise a language that the limited crawl never reaches, while a crawler can also discover alternate-language pages that are not declared in hreflang.

Among the 5,845 confirmed multilingual sites, the median number of observed languages was 3 and the mean was 3.13. Almost half—46.8%—showed exactly two observed languages. Another 16.1% showed three, 14.1% showed four, and 23.0% reached the instrument's maximum of five observed languages.

Five is a crawl ceiling, not a language ceiling. The research fetched the homepage plus at most four content URLs. A site observed in five languages may support ten, twenty or one hundred languages elsewhere. The study does not claim to have counted each site's entire translation corpus.

The advertised-versus-observed comparison shows why a single signal is not enough. Advertised language count was greater than observed count on 38.7% (2,262 / 5,845) of confirmed sites. Observed count was greater than advertised on 34.1% (1,991 / 5,845). The two positive counts were equal on 27.2%.

In addition, 30.5% (1,780 / 5,845) of confirmed sites had zero advertised languages in the study's hreflang-derived field, despite the crawler having verified alternate-language main content. That is another way of seeing the gap between publishing a translation and explicitly describing the locale relationship in machine-readable markup.

Advertised language count vs observed language count

Confirmed multilingual sites · n=5,845 · mutually exclusive count comparison

Advertised = hreflang tags; observed = fetched page bodiesThe page cap limits what can be observed

One further measure needs especially careful wording: 40.4% of confirmed sites had at least one advertised language code that did not appear in the fetched-body language set. This is not evidence that those sites falsely advertise translations. A page outside the five-page crawl could easily contain the advertised language. The useful conclusion is narrower: hreflang declarations and a bounded crawl of substantive page content describe different parts of the system.

For website operators, that suggests two separate QA questions. First: can a crawler reach substantive pages in every language you intend to publish? Second: do your machine-readable locale annotations accurately map those pages to one another? Passing one test does not guarantee passing the other.

04 · Architecture

Subdirectories dominate, but multilingual URL architecture is still diverse

Among the 5,845 confirmed multilingual sites, 51.6% (3,015) used a subdirectory as the primary observed URL pattern. Hybrid implementations were next at 17.9%, followed by unknown at 13.0%, subdomains at 8.6% and ccTLDs at 8.3%. Parameter-based localisation was rare at 0.5%, and same-URL localisation was rarer still at 0.1%.

This is a prevalence result, not an SEO winner table. A subdirectory is common in the sample; the study did not test whether subdirectories outperform subdomains or ccTLDs.

URL architecture among confirmed multilingual sites

Primary mutually exclusive pattern · n=5,845

Subdirectory = plurality, not “best practice” proofHybrid is a GlotEO coding category for multiple observed schemes

What “hybrid” means

A site is coded hybrid when the alternate URLs fetched use two or more schemes—such as ccTLD plus subdomain, or subdomain plus locale paths. Hybrid is not an error label.

Why different URLs matter

Google’s current guidance recommends separate URLs for different language versions and warns that locale-adaptive content on one URL may not all be crawled. That is external guidance, not a conclusion from GlotEO’s measurements.

External guidance: Google Search Central — Managing multi-regional and multilingual sites ↗.

05 · International SEO

The hreflang gap: translations can exist without the machine-readable map

Hreflang was present on 69.7% (4,076 / 5,845) of confirmed multilingual sites. That means 30.3% (1,769 / 5,845) had alternate-language main content that GlotEO actually fetched, but no hreflang was detected through HTML, HTTP Link headers or sitemap XHTML.

This distinction matters because hreflang is an annotation, not a translation detector. Its presence does not prove that substantive alternate-language content exists, and its absence does not prove that a translation is missing. The study measured both independently.

30.3%

of confirmed multilingual sites had no detected hreflang.

1,769 / 5,845 confirmed multilingual sites · Static HTML v1.0

When hreflang did exist on confirmed sites, most of the technical checks were relatively strong. Of the 4,076 confirmed sites with hreflang, 86.8% used valid language codes, 98.2% used absolute URLs, 91.1% were self-referential, 96.0% placed annotations in the HTML head where relevant, and 86.6% had reciprocal annotations among the 4,006 cases where the check applied.

Overall, 72.2% (2,944 / 4,076) passed all applicable checks measured by the study excluding x-default. x-default appeared on 65.2% of confirmed hreflang users, but its absence is treated here as a coverage choice—not an automatic specification failure.

Hreflang quality among confirmed sites that use it

Percent passing each measured check · denominators vary where noted

Core all-pass metric excludes x-defaultReciprocity denominator: 4,006 applicable cases
Practical implication: publishing translated content and explicitly mapping its language/region relationships are separate implementation jobs. A translation workflow can solve the first without fully solving the second.

External guidance: Google Search Central — Tell Google about localized versions ↗.

06 · The localisation gap

The biggest weaknesses appeared after the body was translated

Once a site is confirmed multilingual, the remaining question is implementation completeness. GlotEO’s component analysis shows that some structural signals were common—while search-facing metadata, structured data and sitemap alternates were much less complete.

Title + meta localisation

11.6% (677 / 5,845) received full credit under the study’s title + meta localisation component. A further 2,385 were partial. This was the weakest high-weight component in the implementation index.

Mixed-language leftovers

39.4% (2,302 / 5,845) triggered the mixed-language deduction, meaning the crawler found meaningful strings in another language in navigation, footer or other surrounding page content under the study rules.

HTML language match

86.4% (5,048 / 5,845) received full credit for the HTML language declaration matching detected content. Google says it determines language from visible content rather than relying on the lang attribute.

Automatic locale redirects

15.5% (908 / 5,845) exhibited the automatic-locale redirect behaviour captured by the study. Google advises against automatic language redirection that can prevent access to alternate versions.

Where confirmed multilingual implementations are strongest—and weakest

Full-credit rate for selected GMWS components · n=5,845 confirmed sites

Observed language count is capped by homepage + 4 content URLsFive observed languages is an observational ceiling, not a site maximum

Language navigation is another example where “we found another language” and “the interface exposes a conventional selector” diverge. A crawlable language-selector link was detected on 33.0% (1,927 / 5,845) of confirmed sites, a JS-only selector on 1.3%, and no selector on 65.8%. The last figure should not be interpreted as a UX failure rate: the crawl could confirm alternate content through other routes.

Selectors are useful, but they are not proof of localisation

A language selector is a navigation mechanism. That sounds obvious, but it is an important research distinction. A dropdown can list ten languages even when the crawler cannot verify substantive alternate content, while a multilingual site may expose alternate pages through navigation, sitemaps or locale URL patterns without presenting a conventional selector in the sampled HTML. This is why the classifier never used a selector alone as confirmation.

The same principle applies to hreflang, ccTLDs and html lang. Each is a useful signal, but none is the content itself. The study's stronger confirmation standard—fetching substantial alternate-language main content—was designed to avoid turning markup prevalence into an inflated multilingual-adoption number.

Automatic redirects can hide the alternative path

Automatic locale redirects were detected on 15.5% of confirmed sites. GlotEO did not attempt to defeat those redirects by spoofing language preferences, and the crawler sent no Accept-Language header. This is important because a site that chooses a locale based on request context can present a different crawl surface from a site where every language has an independently reachable URL.

Google separately recommends letting users switch between language versions and warns that automatic redirection can prevent users and search engines from seeing every version. The research does not measure the ranking effect of redirects; it records how often the behaviour appeared in the confirmed multilingual sample.

Structured data and sitemaps are still lightly localised

Only 19.2% (1,125 / 5,845) of confirmed sites received full credit for the study's inLanguage structured-data component, and only 11.9% (693 / 5,845) received full credit for sitemap XHTML alternates. These are not universal requirements for every multilingual site, so low adoption should not be translated directly into an “error rate.” They are better understood as additional machine-readable layers that relatively few confirmed sites in the sample used completely.

Taken together, these results show why multilingual audits are easy to oversimplify. A team can correctly translate its main content, publish clean locale URLs and maintain strong canonicals while still leaving gaps in metadata, sitemaps, structured data or language navigation. Conversely, a site can have excellent markup while the translated page itself remains incomplete. A useful audit needs to inspect both the content and the surrounding implementation.

07 · Implementation completeness

The GlotEO Multilingual Web Score shows a wide implementation spread

The GlotEO Multilingual Web Score (GMWS) is a 0–100 implementation-completeness index calculated only for the 5,845 confirmed multilingual sites. It is not a Google ranking factor, not an E-E-A-T score and not a score for the web as a whole.

The median confirmed site scored 75, with a mean of 67.1, a 25th percentile of 50 and a 75th percentile of 83. The distribution has a meaningful left tail: 15.0% of confirmed sites scored 20–39, even though another 36.6% scored 80–100.

GMWS distribution among confirmed multilingual sites

Implementation completeness index · n=5,845 · not a ranking factor

Median 75 · Mean 67.1 · p25 50 · p75 83Weights are GlotEO’s published research framework

The score is useful because its components are published rather than hidden. Confirmed alternate content is 100% by definition. Crawlable distinct URLs received full credit on 86.4% of confirmed sites, canonical-vs-hreflang consistency on 82.3%, HTML language match on 86.4%, and slug translation on 69.0%. By contrast, title + meta localisation was 11.6%, sitemap XHTML alternates 11.9%, and inLanguage structured data 19.2%.

Interpretation: the median confirmed site is not “badly multilingual.” The more useful finding is unevenness: a site can have strong alternate content and clean URLs while still leaving metadata, sitemap or structured-data localisation incomplete.
08 · Robustness check

JavaScript changed hundreds of classifications, but not the one-decimal headline

The static HTML freeze deliberately did not render JavaScript. It identified 889 successful sites as js_required: 768 uncertain and 121 probable. None were confirmed or monolingual in the static freeze because their script-heavy shells prevented the normal classifier from obtaining enough substantive text to make those calls.

A targeted Chromium pass rendered those 889 sites only. It succeeded on 858 and kept the original class for 31 render failures. Among the 768 previously uncertain sites, 66 became confirmed, 48 probable, 412 monolingual and 242 remained uncertain. Among the 121 probable sites, 51 became confirmed, 67 stayed probable and 3 became uncertain.

After overlaying only those 889 results onto the unchanged 24,116 other successful rows, confirmed rose from 5,845 to 5,962 (+117) and uncertain fell from 2,820 to 2,297 (−523). Yet the confirmed-only rate among non-uncertain successes remained 26.3% to one decimal: 5,962 / 22,708.

Static freeze vs targeted JavaScript adjustment

Same 25,005 successful domains; only 889 js_required rows were re-evaluated

Bar length is each class as a share of 25,005 successful analysesv1.1 is not a full recrawl · 31 render failures kept their v1.0 class
Why this matters: the JavaScript pass supports the stability of the headline rate, but it does not turn the study into a fully rendered-web census. The other 24,116 successful rows were not re-rendered in Chromium. A second targeted overlay—on Yoruba and Tagalog language-ID dependence—is reported next. It also does not replace the frozen 26.3% headline.
09 · Robustness & sensitivity

How robust is the 26.3% estimate?

The primary result remains the frozen static-HTML estimate: 26.3% (5,845 / 22,185) confirmed multilingual among classifiable successful analyses. We subsequently tested two parts of the methodology most likely to affect classification: JavaScript-dependent sites, and a set of classifications that depended on Yoruba (yo) or Tagalog (tl) language detection.

Neither targeted analysis changed the central conclusion. The JavaScript-adjusted result remained 26.3% (5,962 / 22,708). Applying a stricter alternate-page rule to Yoruba/Tagalog-dependent classifications produced a sensitivity estimate of 25.5% (5,647 / 22,165). Combining both targeted adjustments produced 25.4% (5,764 / 22,688).

These are sensitivity analyses, not replacements for the frozen primary result. They are also not four separate crawls: the static freeze is unchanged, and the two overlays re-evaluate only the rows they were designed to test.

Important: 25.4%–26.3% is not a statistical confidence interval. These values come from different targeted classification scenarios. The frozen 26.3% static-HTML result remains the primary estimate.

Targeted robustness checks did not change the central conclusion

The confirmed multilingual rate remained close to one quarter of classifiable sites under each targeted alternative analysis.

Static HTML v1.0
26.3%
Targeted JS adjustment
26.3%
Yoruba / Tagalog sensitivity
25.5%
JS + yo/tl sensitivity
25.4%

Primary frozen result

Static HTML v1.0

26.3%

5,845 confirmed · 22,185 classifiable

Select a scenario to inspect its numerator, denominator and statusNot a statistical confidence interval
Confirmed multilingual rates under the frozen study and three targeted overlays
Analysis Confirmed Classifiable denominator Confirmed rate Status
Static HTML v1.0 5,845 22,185 26.3% Primary frozen result
Targeted JS adjustment 5,962 22,708 26.3% Robustness analysis
Yoruba / Tagalog sensitivity 5,647 22,165 25.5% Sensitivity analysis
JS + Yoruba/Tagalog sensitivity 5,764 22,688 25.4% Combined sensitivity

Inspect the underlying sensitivity-analysis data →

What changed under the stricter rule?

The language-ID overlay looked only at 242 rows whose multilingual class depended on lingua codes yo (Yoruba) or tl (Tagalog) as the only differing detected language: 228 freeze-confirmed, plus 14 probable. It used saved HTML. Zero live fetches. It was not a recrawl of the 25,005-site dataset. wikipedia.org was excluded from the 242-row dependency set because its confirmed multilingual classification did not depend on the tl homepage-language detection.

Under that overlay, confirmed moves from 5,845 to 5,647 (−198). The corresponding classifiable denominator is 22,165, so the sensitivity estimate is 25.5% (5,647 / 22,165). That figure is not the headline.

After stripping repeated site chrome, yo or tl still appeared in observed languages on 217 / 242 rows. Only 25 / 242 lost the language code entirely. Therefore most classification changes did not result from the language detector simply retracting Yoruba or Tagalog. Most came from a tighter alternate-page rule: a yo or tl alternate was counted as multilingual evidence only when its URL represented a genuine locale page, language edition or market edition—rather than simply being a different URL whose text was detected as another language. 30 still-confirmed rows include home-side yo/tl that survived chrome-strip.

This suggests that the main issue was not simply that the language detector incorrectly called pages Yoruba or Tagalog. In many cases the detected language remained after boilerplate was removed. The larger change came from asking a stricter question: was the alternate URL genuinely a localized edition of the site? Alternate-page equivalence proved to be an important source of classification uncertainty alongside language identification itself. That reading is limited to this targeted analysis; it is not a global error rate for the classifier.

An independent AI-assisted review, not a human gold-standard validation study, found that the arithmetic holds. The sensitivity figure does not replace the frozen headline.

Known classification examples

These examples show what the sensitivity analysis actually changed. They are methodology illustrations for specific domains under this research instrument—not judgements of those organisations’ global language programmes.

Homepage language-ID miss, classification unchanged

wikipedia.org

The homepage language detector classified the homepage as tl / Tagalog rather than en / English. That is a genuine homepage language-ID miss. Wikipedia remains confirmed multilingual: the crawl independently identified Japanese, German and Russian editions. Distinguish a language-detection error from a multilingual classification error. This case is the first, not the second.

Confirmed → monolingual

office.com

Under the stricter alternate-page rule, office.com moves from confirmed to monolingual because worldwide.aspx was not treated as a genuine localized or language-edition page. The example shows why page equivalence matters: a different URL is not automatically a locale edition.

Several classifications became more conservative

Amazon domains

Under the stricter rule: amazon.it → probable; amazon.com.au → probable; amazon.ae → monolingual; amazon.sa → monolingual. These labels refer only to what could be directly established for that specific domain under this methodology. They do not imply that Amazon is monolingual globally.

10 · Interpretation

What this study can—and cannot—tell you

Large web studies become less useful when their caveats are treated as legal fine print. Several limits materially affect how these numbers should be used, so they belong in the interpretation rather than being hidden at the bottom of the page.

This is a study of major Tranco-ranked sites, not every website

The source population was Tranco list 46W9X. Tranco is designed to provide a research-oriented ranking of popular domains, so the sample is intentionally tilted toward prominent websites rather than the millions of small, inactive or rarely visited domains elsewhere on the internet. The correct claim is about the major websites successfully analysed in this sample, not “26.3% of the web.”

Exclusions are infrastructure outcomes, not monolingual sites

The crawler attempted 40,457 unique domains and successfully analysed 25,005. The remaining 15,452 were excluded for reasons including DNS failures, robots responses, HTTP errors, access restrictions and CAPTCHAs. Those domains were never converted into monolingual observations. Treating inaccessible sites as monolingual would materially bias the result.

“Confirmed” measures discoverable alternate content, not translation quality

A confirmed site passed a deliberately mechanical content test: the crawl fetched substantive main content in a language different from the homepage. The study did not grade linguistic fluency, cultural adaptation, terminology quality, conversion performance or human translation accuracy. A site can be technically confirmed multilingual and still contain poor translations; it can also provide excellent localisation in areas the crawl did not visit.

The crawl is bounded to five content pages per domain

The page cap keeps a 25,000-domain study computationally tractable and prevents large sites from dominating the crawl. The trade-off is incomplete corpus coverage. Observed languages, mixed-language findings, metadata checks and other page-level measures describe what appeared within the bounded sample of pages—not necessarily every URL the domain serves.

JavaScript was tested selectively, not universally

The static freeze used no browser rendering. The follow-up Chromium pass targeted the 889 successful rows explicitly flagged as requiring JavaScript. This was the right place to test the obvious classification hole, but it does not establish what a full browser-rendered recrawl of all 25,005 successes would produce.

Language identification and alternate-page equivalence can produce edge cases

Language identification and alternate-page equivalence are automated classifications and can produce edge cases, particularly on short, boilerplate-heavy or unconventional pages. A targeted Yoruba/Tagalog sensitivity analysis reduced the confirmed estimate from 26.3% to 25.5% under stricter alternate-page rules without changing the overall conclusion. The language-ID sensitivity review was AI-assisted rather than a human gold-standard validation study.

The Multilingual Web Score is a research framework, not a search-engine score

GMWS makes implementation differences easier to compare by publishing a transparent weighting of twelve components plus two deductions. The weights belong to GlotEO's research design. Google did not design or endorse them, and no relationship between the score and rankings, traffic or conversion was tested here.

A useful way to read the report: treat the percentages as benchmarks for implementation patterns among a large sample of major websites. Use them to ask better audit questions—not as universal laws of the internet or causal SEO claims.
11 · Use the research

Benchmark your own multilingual implementation

The research is most useful when it becomes an audit prompt. The checklist below lets you compare common implementation signals with the confirmed multilingual sample. Checking a box does not create a ranking score; it simply records which implementation layers your site has addressed.

Multilingual implementation checklist

Tick what your site has implemented. Benchmarks are static HTML v1.0 unless noted.

0 / 8
implementation checks selected
86.4% full credit
69.7% present
72.2% all-pass*
11.6% full credit
86.4% full credit
69.0% full credit
33.0% crawlable
39.4% had leftovers
*Among confirmed sites with hreflang, passing all applicable study checks excluding x-default.

What teams should do differently

1. Treat translation and discoverability as separate workstreams. A page can exist in another language without clearly advertising its relationship to sibling pages. That is exactly what the 30.3% hreflang gap demonstrates. For implementation context on localized URLs and hreflang, see multilingual SEO.

2. QA the page shell, not just the article or product description. Mixed-language leftovers were detected on 39.4% of confirmed sites. Navigation, footer strings, forms and templates can undermine an otherwise complete translation.

3. Include metadata in the localisation workflow. The 11.6% full-credit rate for title + meta was the most obvious high-weight implementation gap in the score. Search-facing metadata should not be an afterthought.

4. Give each locale a stable destination. The research found distinct crawlable URLs to be common among confirmed sites, and Google separately recommends distinct URLs rather than relying only on locale-adaptive content.

5. Make uncertainty visible in audits. This study did not convert sites it could not classify into “monolingual.” Your own audits should be equally careful around JavaScript shells, geo-adaptation, authentication and crawler access.

12 · Methodology

How the study was run

This is an observational study of major websites in a Tranco-ranked sample, not a random sample of every registered domain and not a customer dataset.

40,457unique domains attempted
25,005successfully analysed
Homepage + 4maximum content-page fetch scope
Tranco 46W9Xsingle source list
lingua 2.2.0language identification library
No Accept-Languagecrawler request behaviour
889js_required successes in the targeted Chromium pass
242yo/tl-dependent rows in the language-ID sensitivity

Classification rules

Confirmed: the crawler fetched an alternate URL whose main-content language differed from the homepage, with at least 200 letters of main content on both sides.

Probable: the site exposed strong multilingual discovery signals—such as hreflang, locale-shaped URLs or a selector pointing to another path—but the crawl did not fetch a qualifying alternate.

Monolingual: the homepage was analysed, no qualifying alternate was fetched and no strong multilingual discovery cluster was detected.

Uncertain: evidence was contradictory or insufficient. Uncertain was kept visible and was never silently recoded as monolingual.

What is frozen, and what is an overlay

Static HTML v1.0

Primary frozen research dataset. The headline remains 26.3% (5,845 / 22,185) confirmed among classifiable successes.

Targeted JS overlay

Robustness analysis of 889 js_required successes only. The remaining 24,116 successful rows were unchanged. This is not a complete recrawl.

Yoruba/Tagalog sensitivity overlay

Alternative classification sensitivity on 242 saved-HTML rows whose class depended on yo or tl. Zero live fetches. Not a recrawl.

Combined targeted display

JavaScript plus language-ID sensitivity shown together as 25.4% (5,764 / 22,688). A display of two overlays, not a fourth crawl.

Methodology details and limitations

The unit of analysis was the Tranco pay-level domain. The crawler attempted HTTPS first, then HTTP on TLS failure. Robots rules were honoured, including the study’s handling of robots 5xx responses. It did not bypass authentication, CAPTCHAs or other access controls. The static v1.0 freeze did not execute JavaScript.

The study attempted 40,457 unique domains and successfully analysed 25,005, a 61.8% fetch/analysis success rate. The 15,452 exclusions are infrastructure/access outcomes, not monolingual websites. Major exclusion categories included DNS failures, robots 5xx responses, HTTP 4xx responses, robots disallow and CAPTCHAs.

The homepage + four-content-URL cap means observed language count is not a census of every locale a site may offer. Five observed languages is therefore the maximum visible under this instrument. Likewise, advertised hreflang languages that were not observed in fetched bodies are not proof that those pages do not exist elsewhere on the site.

The 25k+ cohort was truncated when the study reached its successful-analysis cap; 2,351 candidate domains remained. The sample is rank-biased and English-home heavy. Results should be described as properties of the analysed Tranco major-site sample, not “the web.”

The JavaScript-adjusted v1.1 result is a targeted robustness overlay on 889 js_required successes. It is not a complete Chromium recrawl of all 25,005 sites. The Yoruba/Tagalog overlay is a second targeted analysis of 242 saved-HTML rows whose class depended on yo or tl as the only differing language. Neither overlay replaces the frozen static-HTML dataset. Future validation work will use a blinded human-reviewed sample to estimate classifier precision and recall across multilingual status, page language and alternate-page equivalence.

Tranco source list: 46W9X ↗. Google links elsewhere on this page are labelled external guidance and are not mixed into the GlotEO result tables.

13 · Reproduce and cite

Use the data

Two workbooks are available. The frozen primary dataset is the source for the 26.3% headline and the downstream implementation statistics. A separate derived workbook supports the targeted JavaScript and Yoruba/Tagalog overlays. It does not replace the freeze.

Primary frozen dataset

The original frozen dataset underlying the primary State of the Multilingual Web 2026 findings, including the 26.3% confirmed multilingual estimate.

Open primary frozen dataset ↗

File: State-of-the-Multilingual-Web-2026-freeze.xlsx

Robustness & sensitivity workbook

Derived analysis supporting the targeted JavaScript and Yoruba/Tagalog classification sensitivity checks. These overlays supplement the frozen dataset and do not replace the primary research results.

Open robustness & sensitivity data ↗

File: State-of-the-Multilingual-Web-2026-derived-overlays.xlsx

Short citation

GlotEO. The State of the Multilingual Web 2026. GlotEO Research, 2026.

Recommended citation

GlotEO. “The State of the Multilingual Web 2026.” GlotEO Research, 2026. Static HTML freeze as of 28 August 2026; targeted JavaScript and language-ID sensitivity analyses as of 29 August 2026. Headline N=25,005 successfully analysed pay-level domains from Tranco list 46W9X. Frozen primary estimate: 26.3% (5,845 / 22,185) confirmed multilingual among classifiable successes.

Research integrity

Static HTML v1.0 headline: 26.3% (5,845 / 22,185) confirmed multilingual among non-uncertain successes. Targeted JS overlay: 26.3% (5,962 / 22,708) after re-evaluating the 889 js_required successes only. Yoruba/Tagalog sensitivity overlay: 25.5% (5,647 / 22,165). Combined targeted display: 25.4% (5,764 / 22,688). The 25.4%–26.3% spread is not a confidence interval.

Freeze SHA-256: fbd917e414a41b7c7eb9d907f1f2347730bf2a778060aaae4e565de0bf380394

When quoting a percentage, preserve its denominator and identify whether the figure is static v1.0, JS-adjusted, or a Yoruba/Tagalog sensitivity overlay. The 42.6% confirmed+probable figure is a sensitivity measure, not the headline adoption rate. Downstream implementation statistics on this page remain tied to the frozen 5,845 confirmed sites.

Research source: GlotEO original crawl and frozen derived dataset. External guidance: Google Search Central pages linked above. The external guidance is context only and is not incorporated into GlotEO’s original percentage calculations.

Published 29 August 2026. Research update: 29 August 2026 — added targeted Yoruba/Tagalog classification sensitivity analysis. The original static-HTML freeze and primary 26.3% result are unchanged.

Put the findings into practiceTry GlotEO for multilingual website and app translation, hreflang, localized URLs, and metadata.

Try For Free All research