How Chinese compares to other languages: a linguistic analysis
Mandarin next to English, Spanish, French, German, Japanese, Korean, Vietnamese and Arabic, measured on open linguistic datasets. Pick your language.

Japanese, Korean and Vietnamese are full of words borrowed from Chinese. None of them is related to Chinese.
If you speak English, Spanish, French, German, Japanese, Korean, Vietnamese or Arabic and wonder what Mandarin will feel like, this guide measures it on open linguistic data. You can pick your own language and see it feature by feature.
Key takeaways
- Vietnamese is closest in grammar. It differs from Mandarin on 42% of the learner features we checked, with Korean and Japanese close behind.
- None of the eight shares Mandarin's basic words. On core words like "eye" and "water", all eight are no closer to Mandarin than chance.
- The shared layer is borrowed. Japanese took more than a quarter of its core word list from Chinese, while Mandarin borrowed almost nothing.
- English hides its similarities. It shares a lot of quiet structure with Mandarin but differs on almost everything you hear in a first lesson.
- Only Japanese still writes with Chinese characters. Korea teaches some hanja in school, and Vietnam replaced characters with the Latin alphabet.
- Grammar isn't the whole story. Vietnamese is closest in grammar, but writing and vocabulary weigh at least as much in official learning times.
Pick the language you speak best and see what Mandarin will feel like:
What will Mandarin feel like for you?
Pick the language you speak best to see, feature by feature, what Mandarin will feel like. Vietnamese, Korean and Japanese share the most grammar with Mandarin; none of the eight shares basic words with it beyond chance.
Grammar that works like Mandarin
8 of 36
Learner features, WALS. All shared WALS features: 75 of 140.
Basic-word distance
101.6
≈100 means no more alike than chance (ASJP, 40 core words). Cantonese scores 80.5.
Chinese-derived words
0
of 1,516 words · Chinese loans in WOLD's English core list (even 'tea' came via Dutch)
US diplomat training time
88 weeks
FSI Category IV, 2,200 class hours, against 30 weeks for Spanish or French.
What Mandarin will feel like if you speak English
- Feels familiar (2): Subject, verb, object, Adjectives go first
- Easier (7): Verbs never change, No tense endings, No I/me, he/him, Plurals are optional, Questions keep statement order, No -er or 'more', Short, simple syllables
- New for you (8): Tones change the word, A few new sounds, Measure words, Aspect, not tense, Adjectives act like verbs, A polite 'you', Characters, not letters, Almost no shared words
- Rewire the order (2): Descriptions go before the noun, Place and time before the verb
Every language at a glance
Language · Grammar like Mandarin (of 36 learner features) · Basic-word distance (≈100 = chance) · Chinese-derived words
EnglishIndo-European (Germanic) · Latin alphabet
8 of 36
101.6
0 of 1,516 words: Chinese loans in WOLD's English core list (even 'tea' came via Dutch)
SpanishIndo-European (Romance) · Latin alphabet
12 of 35
97.7
≈0: No Chinese-derived word layer (not covered by WOLD)
FrenchIndo-European (Romance) · Latin alphabet
10 of 36
101.6
≈0: No Chinese-derived word layer (not covered by WOLD)
GermanIndo-European (Germanic) · Latin alphabet
7 of 33
101.1
≈0: No Chinese-derived word layer (not covered by WOLD)
JapaneseJaponic · Kanji + hiragana + katakana
19 of 35
101.3
27.3% of WOLD's Japanese core list came from Chinese (581 of 2,131 words)
KoreanKoreanic · Hangul (hanja taught in school)
18 of 33
95.5
57% of Standard Korean Language Dictionary headwords are Sino-Korean (NIKL 2000) (251,478 of 440,262 headwords)
VietnameseAustroasiatic (Vietic) · Latin alphabet (chữ Quốc ngữ)
21 of 36
99.4
25.5% of WOLD's Vietnamese core list came from Chinese (391 of 1,534 words)
ArabicAfroasiatic (Semitic) · Arabic abjad, right to left
8 of 35
101.6
≈0: No Chinese-derived word layer; شاي shāy 'tea' is one famous exception
US diplomat training (FSI) takes about 88 weeks, or 2,200 class hours, for English speakers learning Mandarin, against 30 weeks for Spanish or French and 36 for German. FSI publishes no estimates for speakers of other languages.
"Feels familiar" means a feature works the same way in your language, "easier" that Mandarin drops something your language makes you track, "new" that Mandarin adds something, and "rewire" the same pieces in a different order. These labels are our reading of the WALS values. Data: WALS Online v2020.6, ASJP v21.1 and WOLD v4.2 (CC BY 4.0), and PHOIBLE (CC BY-SA 3.0).
Is Chinese related to Japanese, Korean or Vietnamese?
No. Mandarin is a Sino-Tibetan language. Japanese and Korean each belong to small families of their own, Vietnamese is Austroasiatic, Arabic is a Semitic language in the Afroasiatic family, and English, Spanish, French and German are Indo-European, according to WALS's language data.
What links Mandarin with Japanese, Korean and Vietnamese is two thousand years of borrowing from written Chinese, not a common ancestor.
You can see the difference between borrowing and ancestry in basic words. ASJP keeps a 40-word list for thousands of languages, with words like "I", "two", "eye", "water", "fire" and "name" that languages rarely borrow.
We compared Mandarin's list with each language's using LDND, the normalised edit distance of the Wichmann method. On its scale, about 100 means the words are no more alike than random pairs.
Only Cantonese, Mandarin's sister language, scores clearly below chance, at 80.5. Korean, Vietnamese and Japanese all score near 100, about as far from Mandarin as English, French and Arabic. The chart in the next section has every score.
For scale, English and German score 75.4, and English and Spanish 94.1, by our computation on ASJP's word lists.
Japanese, Korean and Vietnamese sit at chance level because their core words are native. Here's "eye" in all four:
| Language | "Eye" |
|---|---|
| Mandarin | 眼睛 |
| Japanese | 目 (me) |
| Korean | 눈 nun |
| Vietnamese | mắt |
The Chinese loans come in a later, more educated layer of the vocabulary. That's where the next sections find them.
Grammar puts Vietnamese, Korean and Japanese closest
On grammar, the order is the same whichever way we count. Vietnamese, Korean and Japanese are closer to Mandarin than any of the European languages or Arabic.
Thai, an unrelated language we added as a reference, lands right next to them. That's the first warning about what "close" means here:
How far is each language from Mandarin?
Grammar distance is the share of WALS features whose values differ; basic-word distance is ASJP's LDND on 40 core words, where about 100 means no more alike than chance. Neither measures how long a language takes to learn.
Distance from Mandarin (95% ranges from 2,000 resamples of the features)
Cantonese (reference)
- 36 learner features that differ
- 22% (5 of 23) range 6%–39%
- All shared features that differ
- 19% (16 of 83) range 11%–28%
- Basic words (LDND)
- 80.5
Vietnamese
- 36 learner features that differ
- 42% (15 of 36) range 25%–58%
- All shared features that differ
- 33% (45 of 136) range 25%–41%
- Basic words (LDND)
- 99.4
Thai (reference)
- 36 learner features that differ
- 44% (15 of 34) range 28%–61%
- All shared features that differ
- 34% (44 of 128) range 27%–43%
- Basic words (LDND)
- 100.8
Korean
- 36 learner features that differ
- 45% (15 of 33) range 28%–63%
- All shared features that differ
- 36% (46 of 129) range 27%–44%
- Basic words (LDND)
- 95.5
Japanese
- 36 learner features that differ
- 46% (16 of 35) range 31%–62%
- All shared features that differ
- 44% (58 of 131) range 36%–53%
- Basic words (LDND)
- 101.3
Spanish
- 36 learner features that differ
- 66% (23 of 35) range 50%–81%
- All shared features that differ
- 53% (73 of 138) range 44%–61%
- Basic words (LDND)
- 97.7
French
- 36 learner features that differ
- 72% (26 of 36) range 58%–86%
- All shared features that differ
- 63% (87 of 139) range 55%–71%
- Basic words (LDND)
- 101.6
Arabic
- 36 learner features that differ
- 77% (27 of 35) range 63%–91%
- All shared features that differ
- 59% (78 of 133) range 50%–67%
- Basic words (LDND)
- 101.6
English
- 36 learner features that differ
- 78% (28 of 36) range 64%–89%
- All shared features that differ
- 46% (65 of 140) range 38%–55%
- Basic words (LDND)
- 101.6
German
- 36 learner features that differ
- 79% (26 of 33) range 64%–91%
- All shared features that differ
- 58% (77 of 132) range 50%–67%
- Basic words (LDND)
- 101.1
Distance from English
German
- 36 learner features that differ
- 18% (6 of 33) range 6%–32%
- All shared features that differ
- 33% (45 of 135) range 25%–41%
- Basic words (LDND)
- 75.4
French
- 36 learner features that differ
- 31% (11 of 36) range 17%–47%
- All shared features that differ
- 36% (51 of 142) range 28%–44%
- Basic words (LDND)
- 92.8
Spanish
- 36 learner features that differ
- 37% (13 of 35) range 22%–54%
- All shared features that differ
- 35% (50 of 141) range 28%–44%
- Basic words (LDND)
- 94.1
Arabic
- 36 learner features that differ
- 60% (21 of 35) range 43%–75%
- All shared features that differ
- 50% (68 of 136) range 41%–59%
- Basic words (LDND)
- 96.4
Korean
- 36 learner features that differ
- 61% (20 of 33) range 44%–77%
- All shared features that differ
- 53% (70 of 132) range 44%–62%
- Basic words (LDND)
- 99.8
Japanese
- 36 learner features that differ
- 69% (24 of 35) range 53%–83%
- All shared features that differ
- 61% (82 of 134) range 53%–69%
- Basic words (LDND)
- 99.3
Cantonese (reference)
- 36 learner features that differ
- 70% (16 of 23) range 50%–88%
- All shared features that differ
- 49% (42 of 86) range 38%–60%
- Basic words (LDND)
- 96.9
Thai (reference)
- 36 learner features that differ
- 71% (24 of 34) range 54%–85%
- All shared features that differ
- 48% (63 of 131) range 40%–57%
- Basic words (LDND)
- 99.5
Vietnamese
- 36 learner features that differ
- 72% (26 of 36) range 58%–86%
- All shared features that differ
- 51% (70 of 138) range 42%–59%
- Basic words (LDND)
- 103.9
Mandarin
- 36 learner features that differ
- 78% (28 of 36) range 64%–89%
- All shared features that differ
- 46% (65 of 140) range 38%–55%
- Basic words (LDND)
- 101.6
Cantonese and Thai are references: a sister language of Mandarin and an unrelated tone language. Data: WALS Online v2020.6, ASJP v21.1 and WOLD v4.2 (CC BY 4.0), and PHOIBLE (CC BY-SA 3.0); distances computed by SayMei.
We measured grammar distance as the share of WALS features whose values differ between two languages. WALS, the World Atlas of Language Structures, codes up to 192 features per language, from "does the language have tones?" to "does the relative clause come before the noun?". We used two feature sets:
- 36 learner features: the ones a learner notices, chosen before we looked at results. They cover tones, syllables, verb endings, gender, plurals, cases, tense and aspect, measure words, word order and questions.
- All shared features: every WALS feature coded for at least seven of the nine languages, minus a handful of word-meaning features, up to 143 per pair.
English is the interesting case. On the 36 features a learner feels, it differs from Mandarin on 78%, near the bottom.
On everything WALS codes, it differs on just 46%, statistically tied with Japanese: when we resampled the features, their ranges overlapped almost entirely. English and Mandarin share a lot of quiet structure, like subject-verb-object order, no case endings on nouns and few verb forms next to Spanish or Korean.
What they don't share is everything you hear in the first lesson.
Measured from English, German is closest, differing on 18% of learner features, and French and Spanish come next. Mandarin differs on 78%, the most of the nine.
Tones, short syllables and a few new sounds
Mandarin's sound system isn't unusually big. PHOIBLE, a database of sound inventories, lists 24 to 27 consonants for Mandarin across its four descriptions.
That's the same range as English across its nine, and WALS puts both in its "average" class. The differences are which sounds you use, and what pitch does.
Pitch is the big one. Of the nine languages, only Mandarin and Vietnamese use it to tell words apart, WALS finds; so do Cantonese and Thai, our two reference languages. Here's how each language uses pitch:
Mandarin
- Does pitch change the word?
- Yes: four tones
- Example
- "mom", 妈; "hemp", 麻; "horse", 马; "to scold", 骂
Vietnamese
- Does pitch change the word?
- Yes: six tones
- Example
- ma, má, mà, mả, mã, mạ are six different words
Japanese
- Does pitch change the word?
- Partly: pitch accent, a "simple tone system"
- Example
- In Tokyo Japanese, 箸 (hashi) "chopsticks" and 橋 (hashi) "bridge" differ by pitch
English, Spanish, French, German, Korean, Arabic
- Does pitch change the word?
- No
- Example
- Pitch shows emphasis or a question, never a different word
Vietnamese speakers already work this way. Everyone else builds a new habit, and our tone pair trainer plays all 20 tone combinations to help. Our guide to learning Mandarin tones explains how adults get them right.
Mandarin syllables are short, too. They end in a vowel, -n or -ng: 中, 南, 好.
By our count, CC-CEDICT's single-character readings use just 419 distinct syllables, or 1,291 once tones are counted. English, French, German and Arabic allow far heavier consonant clusters.
Most speakers meet the same short list of new sounds. Here's how each language starts, based on PHOIBLE's sound inventories:
English
- Already have
- –
- Close
- p/b as a puff of air (pin vs spin), z/c, h, zh/ch/sh/r
- New
- j/q/x, ü, tones
Spanish
- Already have
- h (your jota)
- Close
- –
- New
- puffs of air on p/t/k, zh/ch/sh/r, j/q/x, z/c, ü, tones
French
- Already have
- ü (your u in tu)
- Close
- zh/ch/sh/r
- New
- puffs of air on p/t/k, j/q/x, z/c, h, tones
German
- Already have
- ü, z (as in Zeit), h (as in ach)
- Close
- p/b, zh/ch/sh/r, x (as in ich)
- New
- j/q, tones
Japanese
- Already have
- –
- Close
- j/q/x (chi, shi), z/c (tsu), h, pitch
- New
- puffs of air on p/t/k, zh/ch/sh/r, ü
Korean
- Already have
- puffs of air (ㅂ/ㅍ, ㅈ/ㅊ), j/q/x
- Close
- h
- New
- zh/ch/sh/r, z/c, ü, f, tones
Vietnamese
- Already have
- h (your kh), tones
- Close
- t/d (your th/t), j (your ch)
- New
- zh/ch/sh/r, z/c, ü
Arabic
- Already have
- h (your خ kh)
- Close
- zh/ch/sh/r
- New
- p (Arabic has no p), j/q/x, z/c, ü, tones
Source: SayMei reading of PHOIBLE inventories 2175, 2210, 2182, 2184, 2196, 2197, 2233 and 2157 (CC BY-SA 3.0) plus WALS 11A and 13A. Not yet reviewed by a phonetician.
The pinyin chart has audio for every syllable, if you want to hear these.
The grammar drops endings and adds three habits
Mandarin removes most of what makes European grammar heavy. Then it adds three things of its own: measure words, aspect particles and a strict "modifiers first" order.
Here are all 36 learner features, side by side. Filled cells work the same way as in Mandarin:
36 learner features, side by side
Each row is one WALS feature and each cell that language's value. ● marks a value that matches Mandarin's; – means WALS hasn't coded it.
WALS values for the 36 learner features (● = same as Mandarin)
How many consonants Sounds · WALS 1A
- Mandarin
- Average
- English
- ● Average
- Spanish
- ● Average
- French
- ● Average
- German
- ● Average
- Japanese
- Moderately small
- Korean
- ● Average
- Vietnamese
- ● Average
- Arabic
- Moderately large
How many vowel qualities Sounds · WALS 2A
- Mandarin
- Average (5-6)
- English
- Large (7-14)
- Spanish
- ● Average (5-6)
- French
- Large (7-14)
- German
- Large (7-14)
- Japanese
- ● Average (5-6)
- Korean
- Large (7-14)
- Vietnamese
- Large (7-14)
- Arabic
- ● Average (5-6)
Voiced vs voiceless consonants (b/p, d/t) Sounds · WALS 4A
- Mandarin
- In fricatives alone
- English
- In both plosives and fricatives
- Spanish
- ● In fricatives alone
- French
- In both plosives and fricatives
- German
- In both plosives and fricatives
- Japanese
- In both plosives and fricatives
- Korean
- No voicing contrast
- Vietnamese
- ● In fricatives alone
- Arabic
- In both plosives and fricatives
Front rounded vowels (like ü) Sounds · WALS 11A
- Mandarin
- High only
- English
- None
- Spanish
- None
- French
- High and mid
- German
- High and mid
- Japanese
- None
- Korean
- None
- Vietnamese
- None
- Arabic
- None
Syllable structure Sounds · WALS 12A
- Mandarin
- Moderately complex
- English
- Complex
- Spanish
- ● Moderately complex
- French
- Complex
- German
- Complex
- Japanese
- ● Moderately complex
- Korean
- ● Moderately complex
- Vietnamese
- ● Moderately complex
- Arabic
- Complex
Tone Sounds · WALS 13A
- Mandarin
- Complex tone system
- English
- No tones
- Spanish
- No tones
- French
- No tones
- German
- No tones
- Japanese
- Simple tone system
- Korean
- No tones
- Vietnamese
- ● Complex tone system
- Arabic
- No tones
How grammar attaches to words Word forms · WALS 20A
- Mandarin
- Isolating/concatenative
- English
- Exclusively concatenative
- Spanish
- Exclusively concatenative
- French
- Exclusively concatenative
- German
- Exclusively concatenative
- Japanese
- Exclusively concatenative
- Korean
- Exclusively concatenative
- Vietnamese
- Exclusively isolating
- Arabic
- Ablaut/concatenative
Grammatical categories packed into a verb Word forms · WALS 22A
- Mandarin
- 0-1 category per word
- English
- 2-3 categories per word
- Spanish
- 4-5 categories per word
- French
- 4-5 categories per word
- German
- 2-3 categories per word
- Japanese
- 4-5 categories per word
- Korean
- 6-7 categories per word
- Vietnamese
- ● 0-1 category per word
- Arabic
- 6-7 categories per word
Reduplication (看看 kànkan) Word forms · WALS 27A
- Mandarin
- Productive full and partial reduplication
- English
- No productive reduplication
- Spanish
- No productive reduplication
- French
- No productive reduplication
- German
- No productive reduplication
- Japanese
- Full reduplication only
- Korean
- ● Productive full and partial reduplication
- Vietnamese
- ● Productive full and partial reduplication
- Arabic
- ● Productive full and partial reduplication
Verb endings that agree with the subject Word forms · WALS 29A
- Mandarin
- No subject person/number marking
- English
- Syncretic
- Spanish
- Syncretic
- French
- Syncretic
- German
- Syncretic
- Japanese
- ● No subject person/number marking
- Korean
- ● No subject person/number marking
- Vietnamese
- ● No subject person/number marking
- Arabic
- Syncretic
Grammatical gender Word forms · WALS 30A
- Mandarin
- None
- English
- Three
- Spanish
- Two
- French
- Two
- German
- Three
- Japanese
- –
- Korean
- –
- Vietnamese
- ● None
- Arabic
- Two
How plurals are marked Word forms · WALS 33A
- Mandarin
- Plural suffix
- English
- ● Plural suffix
- Spanish
- ● Plural suffix
- French
- ● Plural suffix
- German
- ● Plural suffix
- Japanese
- ● Plural suffix
- Korean
- ● Plural suffix
- Vietnamese
- Plural word
- Arabic
- Mixed morphological plural
When plurals must be marked Word forms · WALS 34A
- Mandarin
- Only human nouns, optional
- English
- All nouns, always obligatory
- Spanish
- All nouns, always obligatory
- French
- All nouns, always obligatory
- German
- All nouns, always obligatory
- Japanese
- ● Only human nouns, optional
- Korean
- –
- Vietnamese
- All nouns, always optional
- Arabic
- All nouns, always obligatory
Grammatical cases Word forms · WALS 49A
- Mandarin
- No morphological case-marking
- English
- 2 cases
- Spanish
- ● No morphological case-marking
- French
- ● No morphological case-marking
- German
- 4 cases
- Japanese
- 8-9 cases
- Korean
- 6-7 cases
- Vietnamese
- ● No morphological case-marking
- Arabic
- ● No morphological case-marking
Can you drop 'I/you/he'? Word forms · WALS 101A
- Mandarin
- Optional pronouns in subject position
- English
- Obligatory pronouns in subject position
- Spanish
- Subject affixes on verb
- French
- Obligatory pronouns in subject position
- German
- Obligatory pronouns in subject position
- Japanese
- ● Optional pronouns in subject position
- Korean
- ● Optional pronouns in subject position
- Vietnamese
- ● Optional pronouns in subject position
- Arabic
- Subject affixes on verb
Person marking on the verb Word forms · WALS 102A
- Mandarin
- No person marking
- English
- Only the A argument
- Spanish
- Both the A and P arguments
- French
- Only the A argument
- German
- Only the A argument
- Japanese
- ● No person marking
- Korean
- ● No person marking
- Vietnamese
- ● No person marking
- Arabic
- Both the A and P arguments
Perfective vs imperfective aspect Time · WALS 65A
- Mandarin
- Grammatical marking
- English
- No grammatical marking
- Spanish
- ● Grammatical marking
- French
- ● Grammatical marking
- German
- No grammatical marking
- Japanese
- No grammatical marking
- Korean
- ● Grammatical marking
- Vietnamese
- No grammatical marking
- Arabic
- ● Grammatical marking
Past tense Time · WALS 66A
- Mandarin
- No past tense
- English
- Present, no remoteness distinctions
- Spanish
- Present, no remoteness distinctions
- French
- Present, no remoteness distinctions
- German
- Present, no remoteness distinctions
- Japanese
- Present, no remoteness distinctions
- Korean
- Present, no remoteness distinctions
- Vietnamese
- ● No past tense
- Arabic
- Present, no remoteness distinctions
Future tense ending Time · WALS 67A
- Mandarin
- No inflectional future
- English
- ● No inflectional future
- Spanish
- Inflectional future exists
- French
- Inflectional future exists
- German
- ● No inflectional future
- Japanese
- ● No inflectional future
- Korean
- ● No inflectional future
- Vietnamese
- ● No inflectional future
- Arabic
- Inflectional future exists
Perfect ('have done') Time · WALS 68A
- Mandarin
- No perfect
- English
- From possessive
- Spanish
- From possessive
- French
- From possessive
- German
- From possessive
- Japanese
- ● No perfect
- Korean
- ● No perfect
- Vietnamese
- Other perfect
- Arabic
- ● No perfect
Measure words (numeral classifiers) Noun phrase · WALS 55A
- Mandarin
- Obligatory
- English
- Absent
- Spanish
- –
- French
- Absent
- German
- Absent
- Japanese
- ● Obligatory
- Korean
- ● Obligatory
- Vietnamese
- ● Obligatory
- Arabic
- Absent
Possessor before or after the noun Noun phrase · WALS 86A
- Mandarin
- Genitive-Noun
- English
- No dominant order
- Spanish
- Noun-Genitive
- French
- Noun-Genitive
- German
- Noun-Genitive
- Japanese
- ● Genitive-Noun
- Korean
- ● Genitive-Noun
- Vietnamese
- Noun-Genitive
- Arabic
- Noun-Genitive
Adjective before or after the noun Noun phrase · WALS 87A
- Mandarin
- Adjective-Noun
- English
- ● Adjective-Noun
- Spanish
- Noun-Adjective
- French
- Noun-Adjective
- German
- ● Adjective-Noun
- Japanese
- ● Adjective-Noun
- Korean
- ● Adjective-Noun
- Vietnamese
- Noun-Adjective
- Arabic
- Noun-Adjective
'This/that' before or after the noun Noun phrase · WALS 88A
- Mandarin
- Demonstrative-Noun
- English
- ● Demonstrative-Noun
- Spanish
- ● Demonstrative-Noun
- French
- ● Demonstrative-Noun
- German
- ● Demonstrative-Noun
- Japanese
- ● Demonstrative-Noun
- Korean
- ● Demonstrative-Noun
- Vietnamese
- Noun-Demonstrative
- Arabic
- Noun-Demonstrative
Number before or after the noun Noun phrase · WALS 89A
- Mandarin
- Numeral-Noun
- English
- ● Numeral-Noun
- Spanish
- ● Numeral-Noun
- French
- ● Numeral-Noun
- German
- ● Numeral-Noun
- Japanese
- ● Numeral-Noun
- Korean
- ● Numeral-Noun
- Vietnamese
- ● Numeral-Noun
- Arabic
- No dominant order
Relative clause before or after the noun Noun phrase · WALS 90A
- Mandarin
- Relative clause-Noun
- English
- Noun-Relative clause
- Spanish
- Noun-Relative clause
- French
- Noun-Relative clause
- German
- Noun-Relative clause
- Japanese
- ● Relative clause-Noun
- Korean
- ● Relative clause-Noun
- Vietnamese
- Noun-Relative clause
- Arabic
- Noun-Relative clause
Basic word order Sentence · WALS 81A
- Mandarin
- SVO
- English
- ● SVO
- Spanish
- ● SVO
- French
- ● SVO
- German
- No dominant order
- Japanese
- SOV
- Korean
- SOV
- Vietnamese
- ● SVO
- Arabic
- ● SVO
Where 'at home', 'with a pen' go Sentence · WALS 84A
- Mandarin
- XVO
- English
- VOX
- Spanish
- VOX
- French
- VOX
- German
- No dominant order
- Japanese
- XOV
- Korean
- –
- Vietnamese
- VOX
- Arabic
- VOX
Prepositions or postpositions Sentence · WALS 85A
- Mandarin
- No dominant order
- English
- Prepositions
- Spanish
- Prepositions
- French
- Prepositions
- German
- Prepositions
- Japanese
- Postpositions
- Korean
- Postpositions
- Vietnamese
- Prepositions
- Arabic
- Prepositions
Question particle position Sentence · WALS 92A
- Mandarin
- Final
- English
- No question particle
- Spanish
- No question particle
- French
- Initial
- German
- No question particle
- Japanese
- ● Final
- Korean
- No question particle
- Vietnamese
- ● Final
- Arabic
- Initial
Where 'what/who/where' go in questions Sentence · WALS 93A
- Mandarin
- Not initial interrogative phrase
- English
- Initial interrogative phrase
- Spanish
- Initial interrogative phrase
- French
- Initial interrogative phrase
- German
- Initial interrogative phrase
- Japanese
- ● Not initial interrogative phrase
- Korean
- ● Not initial interrogative phrase
- Vietnamese
- ● Not initial interrogative phrase
- Arabic
- ● Not initial interrogative phrase
How yes/no questions are made Sentence · WALS 116A
- Mandarin
- Question particle
- English
- Interrogative word order
- Spanish
- Interrogative word order
- French
- ● Question particle
- German
- Interrogative word order
- Japanese
- ● Question particle
- Korean
- Interrogative verb morphology
- Vietnamese
- ● Question particle
- Arabic
- ● Question particle
Adjectives used like verbs ('I busy') Sentence · WALS 118A
- Mandarin
- Verbal encoding
- English
- Nonverbal encoding
- Spanish
- Nonverbal encoding
- French
- Nonverbal encoding
- German
- –
- Japanese
- Mixed
- Korean
- Mixed
- Vietnamese
- ● Verbal encoding
- Arabic
- Nonverbal encoding
Leaving out 'to be' with nouns Sentence · WALS 120A
- Mandarin
- Impossible
- English
- ● Impossible
- Spanish
- ● Impossible
- French
- ● Impossible
- German
- –
- Japanese
- ● Impossible
- Korean
- ● Impossible
- Vietnamese
- Possible
- Arabic
- Possible
How comparisons are built Sentence · WALS 121A
- Mandarin
- Exceed
- English
- Particle
- Spanish
- Particle
- French
- Particle
- German
- –
- Japanese
- Locational
- Korean
- Locational
- Vietnamese
- ● Exceed
- Arabic
- –
Polite 'you' Sentence · WALS 45A
- Mandarin
- Binary politeness distinction
- English
- No politeness distinction
- Spanish
- ● Binary politeness distinction
- French
- ● Binary politeness distinction
- German
- ● Binary politeness distinction
- Japanese
- Pronouns avoided for politeness
- Korean
- Pronouns avoided for politeness
- Vietnamese
- Pronouns avoided for politeness
- Arabic
- No politeness distinction
Values from WALS Online v2020.6 (CC BY 4.0). Arabic is Egyptian Arabic, the best-covered Arabic in WALS. Same or different compares the WALS value labels exactly, so partial overlaps count as different.
What Mandarin drops
Verbs never change. "I eat" is 我吃, "he eats" is 他吃, and "we eat" is 我们吃.
WALS counts zero to one inflectional category per Mandarin verb. Spanish, French and Japanese have four to five, and Korean and Egyptian Arabic six to seven.
There's no grammatical gender. "He", "she" and "it", 他, 她 and 它, differ only in writing, and all three are pronounced tā.
Nouns don't take plural endings: "one book" is 一本书, and "three books" is 三本书. The suffix 们 is used only for people, and optionally, exactly the WALS value for Japanese.
There's no past tense, either. Time words carry the time: "I ate yesterday" is 我昨天吃了, and "I'll eat tomorrow" is 我明天吃.
What Mandarin adds
That 了 in "I ate yesterday" is an aspect marker. It says the action is complete, not that it happened in the past.
WALS codes Mandarin, Spanish, French, Korean and Arabic as marking aspect grammatically, and English, German, Japanese and Vietnamese as not. So Spanish and French speakers already have the concept, as in comí versus comía, or j'ai mangé versus je mangeais, even though 了 isn't a past tense.
Every counted noun needs a measure word:
- "a book": 一本书
- "a sheet of paper": 一张纸
- "a cat": 一只猫
Japanese, Korean and Vietnamese speakers have counters already; English, French, German and Arabic speakers don't. The measure word finder lists the 36 you'll use most.
Adjectives also work like verbs, so "I am busy" has no "am": 我很忙.
What Mandarin rearranges
The basic order is subject-verb-object, as in English, Spanish, French and Vietnamese. But everything that describes a noun goes before it, however long, and only Japanese and Korean share that habit. Place, time and questions follow their own rules:
- Descriptions first: "the book I bought yesterday" is 我昨天买的书, literally "I yesterday buy DE book".
- Place before the verb: "I eat at home" is 我在家吃饭. Time goes there too.
- Yes-or-no questions keep statement order: "Are you a student?" is 你是学生吗? nǐ shì xuésheng ma?, which adds 吗 at the end, the way Japanese adds か ka.
- Question words stay put: "Where are you going?" is 你去哪儿? nǐ qù nǎr?, with "where" left where the answer goes.
Borrowed words give some learners a head start
An English speaker learning Spanish recognises información and problema on day one. An English speaker learning Mandarin gets almost nothing, and the data shows why.
Mandarin borrows less than any language in the World Loanword Database. By our count, just 26 of the 2,130 words in its list are clearly or probably borrowed.
That's 1.2%, the lowest of the database's 41 languages. The next lowest, Old High German, borrowed 5.6%.
English contributed two of Mandarin's loans: "lemon", 柠檬, and "bus", 巴士. For new things, Mandarin usually builds words from its own parts: a computer is 电脑, literally "electric brain". So friendly loanwords like "coffee", 咖啡, and "sofa", 沙发, stay a short list.
Japanese, Korean and Vietnamese went the other way. They borrowed so much from written Chinese that linguists call the shared layer Sino-Xenic vocabulary:
学生
- Japanese
- 学生 (gakusei)
- Korean
- 학생 haksaeng
- Vietnamese
- học sinh
- English
- student
大学
- Japanese
- 大学 (daigaku)
- Korean
- 대학 daehak
- Vietnamese
- đại học
- English
- university
电话
- Japanese
- 電話 (denwa)
- Korean
- 전화 jeonhwa
- Vietnamese
- điện thoại
- English
- telephone
文化
- Japanese
- 文化 (bunka)
- Korean
- 문화 munhwa
- Vietnamese
- văn hóa
- English
- culture
Sources: Mandarin forms checked against CC-CEDICT; Japanese, Korean and Vietnamese forms against Wiktionary. Not yet reviewed by native speakers.
How big is that layer? It depends on what you count:
Japanese
- Chinese-derived
- 27.3%
- What was counted
- 581 of the 2,131 words in its WOLD list
Vietnamese
- Chinese-derived
- 25.5%
- What was counted
- 391 of 1,534 WOLD words, Old Chinese included
Korean
- Chinese-derived
- 57.12%
- What was counted
- 251,478 of 440,262 dictionary headwords; 25.28% are native
Sources: our counts on WOLD v4.2; for Korean, the Standard Korean Language Dictionary as counted by Jeong 2000, NIKL.
The WOLD lists cover about 1,460 everyday meanings, while the Korean figure counts a full dictionary, technical terms included.
That's why figures of 50–70% circulate online: a dictionary count weighs in thousands of rare, formal nouns.
The Korean dictionary also shows where the layer sits:
| Part of the Korean dictionary | Sino-Korean share |
|---|---|
| Nouns | 62% |
| Particles, all 356 | None |
| Verb endings, all 2,523 | None |
Source: Jeong 2000, NIKL.
The words transfer; the grammar doesn't.
Even shared words need care. The sounds have drifted for centuries, and some meanings moved: in Japanese, 手紙, read tegami, is a letter, while Mandarin 手纸 is toilet paper.
One word did travel everywhere. In all nine languages, the word for tea comes from Chinese, WALS shows:
- Mandarin, Japanese, Korean, Vietnamese and Arabic took cha, as in Mandarin 茶.
- English, Spanish, French and German took Min Nan te. English got it through Dutch thee, according to WOLD.
Only Japanese still writes with Chinese characters
Mandarin is written in characters, each usually a syllable with a meaning, and pinyin is the Latin-letter bridge for learners. Of the eight languages, only Japanese still writes everyday text with Chinese characters.
Here's how each one writes today:
| Language | Script | Chinese characters today |
|---|---|---|
| Japanese | Kanji plus two syllabaries | 2,136 official everyday kanji |
| Korean | Hangul | 1,800 basic hanja taught for classical Chinese |
| Vietnamese | Latin alphabet, quốc ngữ | Replaced Classical Chinese and chữ Nôm |
| English, Spanish, French, German | Latin alphabet | None |
| Arabic | A right-to-left abjad | None |
Sources: Agency for Cultural Affairs, 2010; Kim 2015; Vietnamese Nôm Preservation Foundation.
For Japanese readers, shared characters often keep their meaning, but the readings differ completely, and many characters were simplified differently in China and Japan. "Library" is 图书馆 in Mandarin and 図書館, read toshokan, in Japanese.
Korea's list of basic hanja was last adjusted in 2000. Knowing hanja helps with Mandarin, but daily reading in Korean doesn't use them.
Vietnamese was written for centuries in Classical Chinese characters and in chữ Nôm, a character script built for Vietnamese, before the Latin-based quốc ngữ replaced them. So Vietnamese learners get the vocabulary but not the writing.
How many characters does a Mandarin reader need? In film and TV subtitles, 1,111 characters cover 95% of running text.
To cover 99%, you need 2,076, as our characters guide found using SUBTLEX-CH. For the two script standards in China and Taiwan, see Simplified vs Traditional Chinese.
Distance only partly predicts learning time
These distances line up only partly with official learning times, and those times exist only for English speakers. Here's how long the US Foreign Service Institute budgets for professional speaking and listening:
Source: US Department of State, FSI language categories, checked October 2026. Cantonese and Arabic share Mandarin's 88 weeks, which come to 2,200 class hours.
Set that next to the grammar distances. Vietnamese is the closest of the eight to Mandarin in grammar, yet for an English speaker it's in Category III, while Japanese and Korean are in Category IV like Mandarin. Writing systems and vocabulary weigh at least as much as grammar.
FSI publishes no estimates for speakers of other languages, so there's no official number for a Korean or Spanish speaker learning Mandarin. Our guides to whether Chinese is hard to learn and how long it takes to learn Mandarin cover difficulty from the learner's side and the hours.
Two caveats on the FSI page itself:
- Its hours don't multiply out. 88 weeks at the stated 23 class hours a week is 2,024 hours, not 2,200.
- It changed recently. Between late 2024 and 2026, its goal changed from speaking and reading to speaking and listening, and so did its hours for Categories I to III.
What these numbers don't mean
- Distance isn't difficulty. A feature you lack can be easy, like having no gender to learn, or hard, like tones. The cards label each one, and the labels are our reading of the data.
- The feature list changes the answer. English moves from near the bottom to the middle depending on which WALS features you count. We publish both sets and the 95% ranges.
- WALS codes the dominant pattern of a standard variety. Arabic here is Egyptian Arabic, the best-covered Arabic in WALS, with 145 features against 30 for Modern Standard Arabic. Modern Standard Arabic puts the verb first and question words at the front.
- LDND is coarse. Scores in the 90s don't separate distant relatives from unrelated languages: English and Spanish, both Indo-European, score 94.1.
- Loanword shares depend on the word list. WOLD's everyday list and a full dictionary give very different percentages for the same language.
- No learner data. None of this measures how fast real learners progress.
What to do next
- Pick your language in the tool above. See which features will feel familiar and which will need rewiring.
- Train tones from the start, unless you speak Vietnamese. Of the eight, only Vietnamese has lexical tone, as Mandarin does.
- Use your borrowed words if you speak Japanese, Korean or Vietnamese. Expect familiar meanings in new sounds, and watch for false friends.
- Expect the first lessons to feel new, even from English. The shared structure is quiet; tones, measure words and word order aren't.
- Hear it for yourself. The quickest way to feel the difference is to hear the tones and say a few sentences out loud.
FAQ
Is Chinese similar to Japanese?
In vocabulary and writing, yes; in sound and sentence structure, much less. More than a quarter of the words in WOLD's Japanese list came from Chinese, and Japanese still writes them with kanji. But Japanese puts the verb last and inflects it, and Mandarin doesn't: on 36 learner features from WALS, the two differ on 46%. Basic words like "eye" and "water" aren't related at all.
Can Japanese people read Chinese?
Partly. Japanese uses 2,136 everyday kanji, so a Japanese reader can often guess the meaning of Mandarin compounds like "culture", 文化. But mainland Mandarin uses simplified forms that often differ from Japanese ones, as in "library", 图书馆, which Japanese writes 図書館. Grammar words are different, and some shared words are false friends: Mandarin 手纸 means toilet paper, while Japanese 手紙, read tegami, means a letter.
Does Korean still use Chinese characters?
Mostly not for everyday writing: modern Korean is written in hangul. South Korean schools still teach hanja, from a list of 1,800 basic characters for classical Chinese. The vocabulary is a different story. 57% of the headwords in the Standard Korean Language Dictionary are Sino-Korean, so many Mandarin words come with a familiar meaning.
Is Vietnamese related to Chinese?
Not by descent: Vietnamese is Austroasiatic, and Mandarin is Sino-Tibetan. But Vietnamese borrowed heavily, with 25.5% of the WOLD Vietnamese list coming from Chinese. Its grammar is the closest to Mandarin of the eight languages we compared, with tones, classifiers and no verb endings. It was written in character scripts for centuries before switching to the Latin alphabet.
Is Chinese grammar similar to English grammar?
On the surface, partly: both put subject, verb and object in that order, and neither has case endings on nouns. Across all shared WALS features, they differ on 46%, about as much as Mandarin and Japanese. But the parts a learner meets first differ on 78% of the learner features we checked: tones, measure words, aspect particles and modifiers before the noun.
What language is closest to Mandarin?
Other Chinese languages, such as Cantonese, which differs on 19% of shared WALS features and shares far more basic vocabulary. Among non-Chinese languages in this comparison, Vietnamese is closest in grammar, with 33% of shared features differing, followed by Korean and Japanese. None of them shares basic vocabulary with Mandarin beyond chance. See Mandarin vs Cantonese.
Is Japanese a tonal language like Chinese?
Not in the same way. WALS codes Japanese as a "simple tone system": pitch accent marks where the pitch falls in a word, and it can tell words apart: "chopsticks", 箸, and "bridge", 橋, are both read hashi but differ in pitch. Mandarin gives every syllable one of four tones or a neutral tone, so 妈, 麻, 马 and 骂 are four different words.
Does Chinese have tenses?
No. WALS codes Mandarin with no past tense and no inflectional future. Time comes from words like "yesterday", 昨天, and "tomorrow", 明天, while particles describe the shape of an action: 了 marks it as completed, and 过 as experienced. The verb itself never changes form.
How we did this
Grammar distance is the share of WALS features whose values differ between two languages, over the features coded for both. We used two sets: the 36 learner features, all listed in the matrix above, and every feature coded for at least seven of the nine languages, leaving out six word-meaning features. The 95% ranges come from resampling the features 2,000 times.
Basic-word distance is LDND on ASJP's 40-item lists, using the first listed form per meaning. LDND is the Levenshtein distance, normalised by word length and divided by the average distance between words of different meanings, times 100.
The ASJP doculects were MANDARIN, ENGLISH, SPANISH, FRENCH, STANDARD_GERMAN_2, TOKYO_JAPANESE, KOREAN, VIETNAMESE, STANDARD_ARABIC_2, CANTONESE and THAI.
Loanwords are the WOLD words coded "clearly borrowed" or "probably borrowed", with the source language recorded for Chinese loans.
Sounds come from PHOIBLE inventories, several per language. We report ranges, and use one inventory per language for the sound table.
Cards: the "feels familiar", "easier", "new" and "rewire" labels are our editorial reading of the WALS values, which every card shows.
Data versions:
| Dataset | Version |
|---|---|
| WALS | v2020.6 |
| ASJP | v21.1 |
| WOLD | v4.2 |
| PHOIBLE | development snapshot 5f82b9c |
| CC-CEDICT | 2025-11-29 |
Data licence: the comparison data behind the tools on this page is shared under CC BY 4.0, derived from WALS, ASJP and WOLD, and CC BY-SA 3.0 for the sound groups, derived from PHOIBLE, with the credits below.
Sources
- Dryer, Matthew S. & Martin Haspelmath (eds.). 2013. WALS Online (v2020.6). Zenodo. CC BY 4.0. https://wals.info. Features cited: 1A, 11A, 12A, 13A, 22A, 30A, 34A, 55A, 65A, 66A, 81A, 84A, 90A, 92A, 93A, 118A and 138A.
- Wichmann, Søren, Eric W. Holman & Cecil H. Brown (eds.). The ASJP Database (v21.1). CC BY 4.0. https://asjp.clld.org
- Haspelmath, Martin & Uri Tadmor (eds.). 2009. World Loanword Database (v4.2). CC BY 4.0. https://wold.clld.org
- Moran, Steven & Daniel McCloy (eds.). PHOIBLE 2.0, development snapshot 5f82b9c. CC BY-SA 3.0. https://phoible.org
- CC-CEDICT, MDBG. CC BY-SA 4.0. https://www.mdbg.net/chinese/dictionary?page=cc-cedict
- Jeong Ho-seong. 2000. 표준국어대사전 수록 정보의 통계적 분석 [A statistical analysis of the Standard Korean Language Dictionary]. Saegugeosaenghwal 10(1). National Institute of the Korean Language. https://www.korean.go.kr/nkview/nklife/2000_1/2000_0104.pdf
- Agency for Cultural Affairs, Japan. 2010. 常用漢字表 (Cabinet Notice No. 2).
- Kim Wang-kyu. 2015. Status and trends of education in Chinese characters in Korea. Han-ja Han-mun Education 36.
- Vietnamese Nôm Preservation Foundation. What is Nôm?
- US Department of State. Foreign Language Training (FSI language categories). Checked 2026-10-06.
- Cai, Qing & Marc Brysbaert. 2010. SUBTLEX-CH. PLoS ONE 5(6): e10729.
Methods
This analysis uses four open datasets: WALS v2020.6, ASJP v21.1, WOLD v4.2 (all CC BY 4.0) and PHOIBLE (CC BY-SA 3.0). Mandarin forms were checked against CC-CEDICT and other-language forms against Wiktionary. It has not yet been reviewed by native speakers of the other eight languages. If you spot a mistake, tell us: we fix errors.
