SayMei

How Chinese compares to other languages: a linguistic analysis

Mandarin next to English, Spanish, French, German, Japanese, Korean, Vietnamese and Arabic, measured on open linguistic datasets. Pick your language.

Japanese, Korean and Vietnamese are full of words borrowed from Chinese. None of them is related to Chinese.

If you speak English, Spanish, French, German, Japanese, Korean, Vietnamese or Arabic and wonder what Mandarin will feel like, this guide measures it on open linguistic data. You can pick your own language and see it feature by feature.

Key takeaways

  • Vietnamese is closest in grammar. It differs from Mandarin on 42% of the learner features we checked, with Korean and Japanese close behind.
  • None of the eight shares Mandarin's basic words. On core words like "eye" and "water", all eight are no closer to Mandarin than chance.
  • The shared layer is borrowed. Japanese took more than a quarter of its core word list from Chinese, while Mandarin borrowed almost nothing.
  • English hides its similarities. It shares a lot of quiet structure with Mandarin but differs on almost everything you hear in a first lesson.
  • Only Japanese still writes with Chinese characters. Korea teaches some hanja in school, and Vietnam replaced characters with the Latin alphabet.
  • Grammar isn't the whole story. Vietnamese is closest in grammar, but writing and vocabulary weigh at least as much in official learning times.

Pick the language you speak best and see what Mandarin will feel like:

What will Mandarin feel like for you?

Pick the language you speak best to see, feature by feature, what Mandarin will feel like. Vietnamese, Korean and Japanese share the most grammar with Mandarin; none of the eight shares basic words with it beyond chance.

Grammar that works like Mandarin

8 of 36

Learner features, WALS. All shared WALS features: 75 of 140.

Basic-word distance

101.6

≈100 means no more alike than chance (ASJP, 40 core words). Cantonese scores 80.5.

Chinese-derived words

0

of 1,516 words · Chinese loans in WOLD's English core list (even 'tea' came via Dutch)

US diplomat training time

88 weeks

FSI Category IV, 2,200 class hours, against 30 weeks for Spanish or French.

What Mandarin will feel like if you speak English

  • Feels familiar (2): Subject, verb, object, Adjectives go first
  • Easier (7): Verbs never change, No tense endings, No I/me, he/him, Plurals are optional, Questions keep statement order, No -er or 'more', Short, simple syllables
  • New for you (8): Tones change the word, A few new sounds, Measure words, Aspect, not tense, Adjectives act like verbs, A polite 'you', Characters, not letters, Almost no shared words
  • Rewire the order (2): Descriptions go before the noun, Place and time before the verb

Every language at a glance

Language · Grammar like Mandarin (of 36 learner features) · Basic-word distance (≈100 = chance) · Chinese-derived words

  • EnglishIndo-European (Germanic) · Latin alphabet

    8 of 36

    101.6

    0 of 1,516 words: Chinese loans in WOLD's English core list (even 'tea' came via Dutch)

  • SpanishIndo-European (Romance) · Latin alphabet

    12 of 35

    97.7

    ≈0: No Chinese-derived word layer (not covered by WOLD)

  • FrenchIndo-European (Romance) · Latin alphabet

    10 of 36

    101.6

    ≈0: No Chinese-derived word layer (not covered by WOLD)

  • GermanIndo-European (Germanic) · Latin alphabet

    7 of 33

    101.1

    ≈0: No Chinese-derived word layer (not covered by WOLD)

  • JapaneseJaponic · Kanji + hiragana + katakana

    19 of 35

    101.3

    27.3% of WOLD's Japanese core list came from Chinese (581 of 2,131 words)

  • KoreanKoreanic · Hangul (hanja taught in school)

    18 of 33

    95.5

    57% of Standard Korean Language Dictionary headwords are Sino-Korean (NIKL 2000) (251,478 of 440,262 headwords)

  • VietnameseAustroasiatic (Vietic) · Latin alphabet (chữ Quốc ngữ)

    21 of 36

    99.4

    25.5% of WOLD's Vietnamese core list came from Chinese (391 of 1,534 words)

  • ArabicAfroasiatic (Semitic) · Arabic abjad, right to left

    8 of 35

    101.6

    ≈0: No Chinese-derived word layer; شاي shāy 'tea' is one famous exception

"Feels familiar" means a feature works the same way in your language, "easier" that Mandarin drops something your language makes you track, "new" that Mandarin adds something, and "rewire" the same pieces in a different order. These labels are our reading of the WALS values. Data: WALS Online v2020.6, ASJP v21.1 and WOLD v4.2 (CC BY 4.0), and PHOIBLE (CC BY-SA 3.0).

No. Mandarin is a Sino-Tibetan language. Japanese and Korean each belong to small families of their own, Vietnamese is Austroasiatic, Arabic is a Semitic language in the Afroasiatic family, and English, Spanish, French and German are Indo-European, according to WALS's language data.

What links Mandarin with Japanese, Korean and Vietnamese is two thousand years of borrowing from written Chinese, not a common ancestor.

You can see the difference between borrowing and ancestry in basic words. ASJP keeps a 40-word list for thousands of languages, with words like "I", "two", "eye", "water", "fire" and "name" that languages rarely borrow.

We compared Mandarin's list with each language's using LDND, the normalised edit distance of the Wichmann method. On its scale, about 100 means the words are no more alike than random pairs.

Only Cantonese, Mandarin's sister language, scores clearly below chance, at 80.5. Korean, Vietnamese and Japanese all score near 100, about as far from Mandarin as English, French and Arabic. The chart in the next section has every score.

For scale, English and German score 75.4, and English and Spanish 94.1, by our computation on ASJP's word lists.

Japanese, Korean and Vietnamese sit at chance level because their core words are native. Here's "eye" in all four:

Language"Eye"
Mandarin眼睛(yǎnjing)
Japanese目 (me)
Korean눈 nun
Vietnamesemắt

The Chinese loans come in a later, more educated layer of the vocabulary. That's where the next sections find them.

Grammar puts Vietnamese, Korean and Japanese closest

On grammar, the order is the same whichever way we count. Vietnamese, Korean and Japanese are closer to Mandarin than any of the European languages or Arabic.

Thai, an unrelated language we added as a reference, lands right next to them. That's the first warning about what "close" means here:

How far is each language from Mandarin?

Grammar distance is the share of WALS features whose values differ; basic-word distance is ASJP's LDND on 40 core words, where about 100 means no more alike than chance. Neither measures how long a language takes to learn.

Distance from Mandarin (95% ranges from 2,000 resamples of the features)

  • Cantonese (reference)

    36 learner features that differ
    22% (5 of 23) range 6%–39%
    All shared features that differ
    19% (16 of 83) range 11%–28%
    Basic words (LDND)
    80.5
  • Vietnamese

    36 learner features that differ
    42% (15 of 36) range 25%–58%
    All shared features that differ
    33% (45 of 136) range 25%–41%
    Basic words (LDND)
    99.4
  • Thai (reference)

    36 learner features that differ
    44% (15 of 34) range 28%–61%
    All shared features that differ
    34% (44 of 128) range 27%–43%
    Basic words (LDND)
    100.8
  • Korean

    36 learner features that differ
    45% (15 of 33) range 28%–63%
    All shared features that differ
    36% (46 of 129) range 27%–44%
    Basic words (LDND)
    95.5
  • Japanese

    36 learner features that differ
    46% (16 of 35) range 31%–62%
    All shared features that differ
    44% (58 of 131) range 36%–53%
    Basic words (LDND)
    101.3
  • Spanish

    36 learner features that differ
    66% (23 of 35) range 50%–81%
    All shared features that differ
    53% (73 of 138) range 44%–61%
    Basic words (LDND)
    97.7
  • French

    36 learner features that differ
    72% (26 of 36) range 58%–86%
    All shared features that differ
    63% (87 of 139) range 55%–71%
    Basic words (LDND)
    101.6
  • Arabic

    36 learner features that differ
    77% (27 of 35) range 63%–91%
    All shared features that differ
    59% (78 of 133) range 50%–67%
    Basic words (LDND)
    101.6
  • English

    36 learner features that differ
    78% (28 of 36) range 64%–89%
    All shared features that differ
    46% (65 of 140) range 38%–55%
    Basic words (LDND)
    101.6
  • German

    36 learner features that differ
    79% (26 of 33) range 64%–91%
    All shared features that differ
    58% (77 of 132) range 50%–67%
    Basic words (LDND)
    101.1

Distance from English

  • German

    36 learner features that differ
    18% (6 of 33) range 6%–32%
    All shared features that differ
    33% (45 of 135) range 25%–41%
    Basic words (LDND)
    75.4
  • French

    36 learner features that differ
    31% (11 of 36) range 17%–47%
    All shared features that differ
    36% (51 of 142) range 28%–44%
    Basic words (LDND)
    92.8
  • Spanish

    36 learner features that differ
    37% (13 of 35) range 22%–54%
    All shared features that differ
    35% (50 of 141) range 28%–44%
    Basic words (LDND)
    94.1
  • Arabic

    36 learner features that differ
    60% (21 of 35) range 43%–75%
    All shared features that differ
    50% (68 of 136) range 41%–59%
    Basic words (LDND)
    96.4
  • Korean

    36 learner features that differ
    61% (20 of 33) range 44%–77%
    All shared features that differ
    53% (70 of 132) range 44%–62%
    Basic words (LDND)
    99.8
  • Japanese

    36 learner features that differ
    69% (24 of 35) range 53%–83%
    All shared features that differ
    61% (82 of 134) range 53%–69%
    Basic words (LDND)
    99.3
  • Cantonese (reference)

    36 learner features that differ
    70% (16 of 23) range 50%–88%
    All shared features that differ
    49% (42 of 86) range 38%–60%
    Basic words (LDND)
    96.9
  • Thai (reference)

    36 learner features that differ
    71% (24 of 34) range 54%–85%
    All shared features that differ
    48% (63 of 131) range 40%–57%
    Basic words (LDND)
    99.5
  • Vietnamese

    36 learner features that differ
    72% (26 of 36) range 58%–86%
    All shared features that differ
    51% (70 of 138) range 42%–59%
    Basic words (LDND)
    103.9
  • Mandarin

    36 learner features that differ
    78% (28 of 36) range 64%–89%
    All shared features that differ
    46% (65 of 140) range 38%–55%
    Basic words (LDND)
    101.6

Cantonese and Thai are references: a sister language of Mandarin and an unrelated tone language. Data: WALS Online v2020.6, ASJP v21.1 and WOLD v4.2 (CC BY 4.0), and PHOIBLE (CC BY-SA 3.0); distances computed by SayMei.

We measured grammar distance as the share of WALS features whose values differ between two languages. WALS, the World Atlas of Language Structures, codes up to 192 features per language, from "does the language have tones?" to "does the relative clause come before the noun?". We used two feature sets:

  • 36 learner features: the ones a learner notices, chosen before we looked at results. They cover tones, syllables, verb endings, gender, plurals, cases, tense and aspect, measure words, word order and questions.
  • All shared features: every WALS feature coded for at least seven of the nine languages, minus a handful of word-meaning features, up to 143 per pair.

English is the interesting case. On the 36 features a learner feels, it differs from Mandarin on 78%, near the bottom.

On everything WALS codes, it differs on just 46%, statistically tied with Japanese: when we resampled the features, their ranges overlapped almost entirely. English and Mandarin share a lot of quiet structure, like subject-verb-object order, no case endings on nouns and few verb forms next to Spanish or Korean.

What they don't share is everything you hear in the first lesson.

Measured from English, German is closest, differing on 18% of learner features, and French and Spanish come next. Mandarin differs on 78%, the most of the nine.

Tones, short syllables and a few new sounds

Mandarin's sound system isn't unusually big. PHOIBLE, a database of sound inventories, lists 24 to 27 consonants for Mandarin across its four descriptions.

That's the same range as English across its nine, and WALS puts both in its "average" class. The differences are which sounds you use, and what pitch does.

Pitch is the big one. Of the nine languages, only Mandarin and Vietnamese use it to tell words apart, WALS finds; so do Cantonese and Thai, our two reference languages. Here's how each language uses pitch:

  • Mandarin

    Does pitch change the word?
    Yes: four tones
    Example
    "mom", 妈(mā); "hemp", 麻(má); "horse", 马(mǎ); "to scold", 骂(mà)
  • Vietnamese

    Does pitch change the word?
    Yes: six tones
    Example
    ma, má, mà, mả, mã, mạ are six different words
  • Japanese

    Does pitch change the word?
    Partly: pitch accent, a "simple tone system"
    Example
    In Tokyo Japanese, 箸 (hashi) "chopsticks" and 橋 (hashi) "bridge" differ by pitch
  • English, Spanish, French, German, Korean, Arabic

    Does pitch change the word?
    No
    Example
    Pitch shows emphasis or a question, never a different word

Vietnamese speakers already work this way. Everyone else builds a new habit, and our tone pair trainer plays all 20 tone combinations to help. Our guide to learning Mandarin tones explains how adults get them right.

Mandarin syllables are short, too. They end in a vowel, -n or -ng: 中(zhōng), 南(nán), 好(hǎo).

By our count, CC-CEDICT's single-character readings use just 419 distinct syllables, or 1,291 once tones are counted. English, French, German and Arabic allow far heavier consonant clusters.

Most speakers meet the same short list of new sounds. Here's how each language starts, based on PHOIBLE's sound inventories:

  • English

    Already have
    –
    Close
    p/b as a puff of air (pin vs spin), z/c, h, zh/ch/sh/r
    New
    j/q/x, ü, tones
  • Spanish

    Already have
    h (your jota)
    Close
    –
    New
    puffs of air on p/t/k, zh/ch/sh/r, j/q/x, z/c, ü, tones
  • French

    Already have
    ü (your u in tu)
    Close
    zh/ch/sh/r
    New
    puffs of air on p/t/k, j/q/x, z/c, h, tones
  • German

    Already have
    ü, z (as in Zeit), h (as in ach)
    Close
    p/b, zh/ch/sh/r, x (as in ich)
    New
    j/q, tones
  • Japanese

    Already have
    –
    Close
    j/q/x (chi, shi), z/c (tsu), h, pitch
    New
    puffs of air on p/t/k, zh/ch/sh/r, ü
  • Korean

    Already have
    puffs of air (ㅂ/ㅍ, ㅈ/ㅊ), j/q/x
    Close
    h
    New
    zh/ch/sh/r, z/c, ü, f, tones
  • Vietnamese

    Already have
    h (your kh), tones
    Close
    t/d (your th/t), j (your ch)
    New
    zh/ch/sh/r, z/c, ü
  • Arabic

    Already have
    h (your خ kh)
    Close
    zh/ch/sh/r
    New
    p (Arabic has no p), j/q/x, z/c, ü, tones

Source: SayMei reading of PHOIBLE inventories 2175, 2210, 2182, 2184, 2196, 2197, 2233 and 2157 (CC BY-SA 3.0) plus WALS 11A and 13A. Not yet reviewed by a phonetician.

The pinyin chart has audio for every syllable, if you want to hear these.

The grammar drops endings and adds three habits

Mandarin removes most of what makes European grammar heavy. Then it adds three things of its own: measure words, aspect particles and a strict "modifiers first" order.

Here are all 36 learner features, side by side. Filled cells work the same way as in Mandarin:

36 learner features, side by side

Each row is one WALS feature and each cell that language's value. ● marks a value that matches Mandarin's; – means WALS hasn't coded it.

WALS values for the 36 learner features (● = same as Mandarin)

  • How many consonants Sounds · WALS 1A

    Mandarin
    Average
    English
    ● Average
    Spanish
    ● Average
    French
    ● Average
    German
    ● Average
    Japanese
    Moderately small
    Korean
    ● Average
    Vietnamese
    ● Average
    Arabic
    Moderately large
  • How many vowel qualities Sounds · WALS 2A

    Mandarin
    Average (5-6)
    English
    Large (7-14)
    Spanish
    ● Average (5-6)
    French
    Large (7-14)
    German
    Large (7-14)
    Japanese
    ● Average (5-6)
    Korean
    Large (7-14)
    Vietnamese
    Large (7-14)
    Arabic
    ● Average (5-6)
  • Voiced vs voiceless consonants (b/p, d/t) Sounds · WALS 4A

    Mandarin
    In fricatives alone
    English
    In both plosives and fricatives
    Spanish
    ● In fricatives alone
    French
    In both plosives and fricatives
    German
    In both plosives and fricatives
    Japanese
    In both plosives and fricatives
    Korean
    No voicing contrast
    Vietnamese
    ● In fricatives alone
    Arabic
    In both plosives and fricatives
  • Front rounded vowels (like ü) Sounds · WALS 11A

    Mandarin
    High only
    English
    None
    Spanish
    None
    French
    High and mid
    German
    High and mid
    Japanese
    None
    Korean
    None
    Vietnamese
    None
    Arabic
    None
  • Syllable structure Sounds · WALS 12A

    Mandarin
    Moderately complex
    English
    Complex
    Spanish
    ● Moderately complex
    French
    Complex
    German
    Complex
    Japanese
    ● Moderately complex
    Korean
    ● Moderately complex
    Vietnamese
    ● Moderately complex
    Arabic
    Complex
  • Tone Sounds · WALS 13A

    Mandarin
    Complex tone system
    English
    No tones
    Spanish
    No tones
    French
    No tones
    German
    No tones
    Japanese
    Simple tone system
    Korean
    No tones
    Vietnamese
    ● Complex tone system
    Arabic
    No tones
  • How grammar attaches to words Word forms · WALS 20A

    Mandarin
    Isolating/concatenative
    English
    Exclusively concatenative
    Spanish
    Exclusively concatenative
    French
    Exclusively concatenative
    German
    Exclusively concatenative
    Japanese
    Exclusively concatenative
    Korean
    Exclusively concatenative
    Vietnamese
    Exclusively isolating
    Arabic
    Ablaut/concatenative
  • Grammatical categories packed into a verb Word forms · WALS 22A

    Mandarin
    0-1 category per word
    English
    2-3 categories per word
    Spanish
    4-5 categories per word
    French
    4-5 categories per word
    German
    2-3 categories per word
    Japanese
    4-5 categories per word
    Korean
    6-7 categories per word
    Vietnamese
    ● 0-1 category per word
    Arabic
    6-7 categories per word
  • Reduplication (看看 kànkan) Word forms · WALS 27A

    Mandarin
    Productive full and partial reduplication
    English
    No productive reduplication
    Spanish
    No productive reduplication
    French
    No productive reduplication
    German
    No productive reduplication
    Japanese
    Full reduplication only
    Korean
    ● Productive full and partial reduplication
    Vietnamese
    ● Productive full and partial reduplication
    Arabic
    ● Productive full and partial reduplication
  • Verb endings that agree with the subject Word forms · WALS 29A

    Mandarin
    No subject person/number marking
    English
    Syncretic
    Spanish
    Syncretic
    French
    Syncretic
    German
    Syncretic
    Japanese
    ● No subject person/number marking
    Korean
    ● No subject person/number marking
    Vietnamese
    ● No subject person/number marking
    Arabic
    Syncretic
  • Grammatical gender Word forms · WALS 30A

    Mandarin
    None
    English
    Three
    Spanish
    Two
    French
    Two
    German
    Three
    Japanese
    –
    Korean
    –
    Vietnamese
    ● None
    Arabic
    Two
  • How plurals are marked Word forms · WALS 33A

    Mandarin
    Plural suffix
    English
    ● Plural suffix
    Spanish
    ● Plural suffix
    French
    ● Plural suffix
    German
    ● Plural suffix
    Japanese
    ● Plural suffix
    Korean
    ● Plural suffix
    Vietnamese
    Plural word
    Arabic
    Mixed morphological plural
  • When plurals must be marked Word forms · WALS 34A

    Mandarin
    Only human nouns, optional
    English
    All nouns, always obligatory
    Spanish
    All nouns, always obligatory
    French
    All nouns, always obligatory
    German
    All nouns, always obligatory
    Japanese
    ● Only human nouns, optional
    Korean
    –
    Vietnamese
    All nouns, always optional
    Arabic
    All nouns, always obligatory
  • Grammatical cases Word forms · WALS 49A

    Mandarin
    No morphological case-marking
    English
    2 cases
    Spanish
    ● No morphological case-marking
    French
    ● No morphological case-marking
    German
    4 cases
    Japanese
    8-9 cases
    Korean
    6-7 cases
    Vietnamese
    ● No morphological case-marking
    Arabic
    ● No morphological case-marking
  • Can you drop 'I/you/he'? Word forms · WALS 101A

    Mandarin
    Optional pronouns in subject position
    English
    Obligatory pronouns in subject position
    Spanish
    Subject affixes on verb
    French
    Obligatory pronouns in subject position
    German
    Obligatory pronouns in subject position
    Japanese
    ● Optional pronouns in subject position
    Korean
    ● Optional pronouns in subject position
    Vietnamese
    ● Optional pronouns in subject position
    Arabic
    Subject affixes on verb
  • Person marking on the verb Word forms · WALS 102A

    Mandarin
    No person marking
    English
    Only the A argument
    Spanish
    Both the A and P arguments
    French
    Only the A argument
    German
    Only the A argument
    Japanese
    ● No person marking
    Korean
    ● No person marking
    Vietnamese
    ● No person marking
    Arabic
    Both the A and P arguments
  • Perfective vs imperfective aspect Time · WALS 65A

    Mandarin
    Grammatical marking
    English
    No grammatical marking
    Spanish
    ● Grammatical marking
    French
    ● Grammatical marking
    German
    No grammatical marking
    Japanese
    No grammatical marking
    Korean
    ● Grammatical marking
    Vietnamese
    No grammatical marking
    Arabic
    ● Grammatical marking
  • Past tense Time · WALS 66A

    Mandarin
    No past tense
    English
    Present, no remoteness distinctions
    Spanish
    Present, no remoteness distinctions
    French
    Present, no remoteness distinctions
    German
    Present, no remoteness distinctions
    Japanese
    Present, no remoteness distinctions
    Korean
    Present, no remoteness distinctions
    Vietnamese
    ● No past tense
    Arabic
    Present, no remoteness distinctions
  • Future tense ending Time · WALS 67A

    Mandarin
    No inflectional future
    English
    ● No inflectional future
    Spanish
    Inflectional future exists
    French
    Inflectional future exists
    German
    ● No inflectional future
    Japanese
    ● No inflectional future
    Korean
    ● No inflectional future
    Vietnamese
    ● No inflectional future
    Arabic
    Inflectional future exists
  • Perfect ('have done') Time · WALS 68A

    Mandarin
    No perfect
    English
    From possessive
    Spanish
    From possessive
    French
    From possessive
    German
    From possessive
    Japanese
    ● No perfect
    Korean
    ● No perfect
    Vietnamese
    Other perfect
    Arabic
    ● No perfect
  • Measure words (numeral classifiers) Noun phrase · WALS 55A

    Mandarin
    Obligatory
    English
    Absent
    Spanish
    –
    French
    Absent
    German
    Absent
    Japanese
    ● Obligatory
    Korean
    ● Obligatory
    Vietnamese
    ● Obligatory
    Arabic
    Absent
  • Possessor before or after the noun Noun phrase · WALS 86A

    Mandarin
    Genitive-Noun
    English
    No dominant order
    Spanish
    Noun-Genitive
    French
    Noun-Genitive
    German
    Noun-Genitive
    Japanese
    ● Genitive-Noun
    Korean
    ● Genitive-Noun
    Vietnamese
    Noun-Genitive
    Arabic
    Noun-Genitive
  • Adjective before or after the noun Noun phrase · WALS 87A

    Mandarin
    Adjective-Noun
    English
    ● Adjective-Noun
    Spanish
    Noun-Adjective
    French
    Noun-Adjective
    German
    ● Adjective-Noun
    Japanese
    ● Adjective-Noun
    Korean
    ● Adjective-Noun
    Vietnamese
    Noun-Adjective
    Arabic
    Noun-Adjective
  • 'This/that' before or after the noun Noun phrase · WALS 88A

    Mandarin
    Demonstrative-Noun
    English
    ● Demonstrative-Noun
    Spanish
    ● Demonstrative-Noun
    French
    ● Demonstrative-Noun
    German
    ● Demonstrative-Noun
    Japanese
    ● Demonstrative-Noun
    Korean
    ● Demonstrative-Noun
    Vietnamese
    Noun-Demonstrative
    Arabic
    Noun-Demonstrative
  • Number before or after the noun Noun phrase · WALS 89A

    Mandarin
    Numeral-Noun
    English
    ● Numeral-Noun
    Spanish
    ● Numeral-Noun
    French
    ● Numeral-Noun
    German
    ● Numeral-Noun
    Japanese
    ● Numeral-Noun
    Korean
    ● Numeral-Noun
    Vietnamese
    ● Numeral-Noun
    Arabic
    No dominant order
  • Relative clause before or after the noun Noun phrase · WALS 90A

    Mandarin
    Relative clause-Noun
    English
    Noun-Relative clause
    Spanish
    Noun-Relative clause
    French
    Noun-Relative clause
    German
    Noun-Relative clause
    Japanese
    ● Relative clause-Noun
    Korean
    ● Relative clause-Noun
    Vietnamese
    Noun-Relative clause
    Arabic
    Noun-Relative clause
  • Basic word order Sentence · WALS 81A

    Mandarin
    SVO
    English
    ● SVO
    Spanish
    ● SVO
    French
    ● SVO
    German
    No dominant order
    Japanese
    SOV
    Korean
    SOV
    Vietnamese
    ● SVO
    Arabic
    ● SVO
  • Where 'at home', 'with a pen' go Sentence · WALS 84A

    Mandarin
    XVO
    English
    VOX
    Spanish
    VOX
    French
    VOX
    German
    No dominant order
    Japanese
    XOV
    Korean
    –
    Vietnamese
    VOX
    Arabic
    VOX
  • Prepositions or postpositions Sentence · WALS 85A

    Mandarin
    No dominant order
    English
    Prepositions
    Spanish
    Prepositions
    French
    Prepositions
    German
    Prepositions
    Japanese
    Postpositions
    Korean
    Postpositions
    Vietnamese
    Prepositions
    Arabic
    Prepositions
  • Question particle position Sentence · WALS 92A

    Mandarin
    Final
    English
    No question particle
    Spanish
    No question particle
    French
    Initial
    German
    No question particle
    Japanese
    ● Final
    Korean
    No question particle
    Vietnamese
    ● Final
    Arabic
    Initial
  • Where 'what/who/where' go in questions Sentence · WALS 93A

    Mandarin
    Not initial interrogative phrase
    English
    Initial interrogative phrase
    Spanish
    Initial interrogative phrase
    French
    Initial interrogative phrase
    German
    Initial interrogative phrase
    Japanese
    ● Not initial interrogative phrase
    Korean
    ● Not initial interrogative phrase
    Vietnamese
    ● Not initial interrogative phrase
    Arabic
    ● Not initial interrogative phrase
  • How yes/no questions are made Sentence · WALS 116A

    Mandarin
    Question particle
    English
    Interrogative word order
    Spanish
    Interrogative word order
    French
    ● Question particle
    German
    Interrogative word order
    Japanese
    ● Question particle
    Korean
    Interrogative verb morphology
    Vietnamese
    ● Question particle
    Arabic
    ● Question particle
  • Adjectives used like verbs ('I busy') Sentence · WALS 118A

    Mandarin
    Verbal encoding
    English
    Nonverbal encoding
    Spanish
    Nonverbal encoding
    French
    Nonverbal encoding
    German
    –
    Japanese
    Mixed
    Korean
    Mixed
    Vietnamese
    ● Verbal encoding
    Arabic
    Nonverbal encoding
  • Leaving out 'to be' with nouns Sentence · WALS 120A

    Mandarin
    Impossible
    English
    ● Impossible
    Spanish
    ● Impossible
    French
    ● Impossible
    German
    –
    Japanese
    ● Impossible
    Korean
    ● Impossible
    Vietnamese
    Possible
    Arabic
    Possible
  • How comparisons are built Sentence · WALS 121A

    Mandarin
    Exceed
    English
    Particle
    Spanish
    Particle
    French
    Particle
    German
    –
    Japanese
    Locational
    Korean
    Locational
    Vietnamese
    ● Exceed
    Arabic
    –
  • Polite 'you' Sentence · WALS 45A

    Mandarin
    Binary politeness distinction
    English
    No politeness distinction
    Spanish
    ● Binary politeness distinction
    French
    ● Binary politeness distinction
    German
    ● Binary politeness distinction
    Japanese
    Pronouns avoided for politeness
    Korean
    Pronouns avoided for politeness
    Vietnamese
    Pronouns avoided for politeness
    Arabic
    No politeness distinction

Values from WALS Online v2020.6 (CC BY 4.0). Arabic is Egyptian Arabic, the best-covered Arabic in WALS. Same or different compares the WALS value labels exactly, so partial overlaps count as different.

What Mandarin drops

Verbs never change. "I eat" is 我吃(wǒ chī), "he eats" is 他吃(tā chī), and "we eat" is 我们吃(wǒmen chī).

WALS counts zero to one inflectional category per Mandarin verb. Spanish, French and Japanese have four to five, and Korean and Egyptian Arabic six to seven.

There's no grammatical gender. "He", "she" and "it", 他, 她 and 它, differ only in writing, and all three are pronounced tā.

Nouns don't take plural endings: "one book" is 一本书(yì běn shū), and "three books" is 三本书(sān běn shū). The suffix 们(men) is used only for people, and optionally, exactly the WALS value for Japanese.

There's no past tense, either. Time words carry the time: "I ate yesterday" is 我昨天吃了(wǒ zuótiān chī le), and "I'll eat tomorrow" is 我明天吃(wǒ míngtiān chī).

What Mandarin adds

That 了(le) in "I ate yesterday" is an aspect marker. It says the action is complete, not that it happened in the past.

WALS codes Mandarin, Spanish, French, Korean and Arabic as marking aspect grammatically, and English, German, Japanese and Vietnamese as not. So Spanish and French speakers already have the concept, as in comí versus comía, or j'ai mangé versus je mangeais, even though 了(le) isn't a past tense.

Every counted noun needs a measure word:

  • "a book": 一本书(yì běn shū)
  • "a sheet of paper": 一张纸(yì zhāng zhǐ)
  • "a cat": 一只猫(yì zhī māo)

Japanese, Korean and Vietnamese speakers have counters already; English, French, German and Arabic speakers don't. The measure word finder lists the 36 you'll use most.

Adjectives also work like verbs, so "I am busy" has no "am": 我很忙(wǒ hěn máng).

What Mandarin rearranges

The basic order is subject-verb-object, as in English, Spanish, French and Vietnamese. But everything that describes a noun goes before it, however long, and only Japanese and Korean share that habit. Place, time and questions follow their own rules:

  • Descriptions first: "the book I bought yesterday" is 我昨天买的书(wǒ zuótiān mǎi de shū), literally "I yesterday buy DE book".
  • Place before the verb: "I eat at home" is 我在家吃饭(wǒ zài jiā chīfàn). Time goes there too.
  • Yes-or-no questions keep statement order: "Are you a student?" is 你是学生吗? nǐ shì xuésheng ma?, which adds 吗(ma) at the end, the way Japanese adds か ka.
  • Question words stay put: "Where are you going?" is 你去哪儿? nǐ qù nǎr?, with "where" left where the answer goes.

Borrowed words give some learners a head start

An English speaker learning Spanish recognises información and problema on day one. An English speaker learning Mandarin gets almost nothing, and the data shows why.

Mandarin borrows less than any language in the World Loanword Database. By our count, just 26 of the 2,130 words in its list are clearly or probably borrowed.

That's 1.2%, the lowest of the database's 41 languages. The next lowest, Old High German, borrowed 5.6%.

English contributed two of Mandarin's loans: "lemon", 柠檬(níngméng), and "bus", 巴士(bāshì). For new things, Mandarin usually builds words from its own parts: a computer is 电脑(diànnǎo), literally "electric brain". So friendly loanwords like "coffee", 咖啡(kāfēi), and "sofa", 沙发(shāfā), stay a short list.

Japanese, Korean and Vietnamese went the other way. They borrowed so much from written Chinese that linguists call the shared layer Sino-Xenic vocabulary:

  • 学生(xuésheng)

    Japanese
    学生 (gakusei)
    Korean
    학생 haksaeng
    Vietnamese
    học sinh
    English
    student
  • 大学(dàxué)

    Japanese
    大学 (daigaku)
    Korean
    대학 daehak
    Vietnamese
    đại học
    English
    university
  • 电话(diànhuà)

    Japanese
    電話 (denwa)
    Korean
    전화 jeonhwa
    Vietnamese
    điện thoại
    English
    telephone
  • 文化(wénhuà)

    Japanese
    文化 (bunka)
    Korean
    문화 munhwa
    Vietnamese
    văn hóa
    English
    culture

Sources: Mandarin forms checked against CC-CEDICT; Japanese, Korean and Vietnamese forms against Wiktionary. Not yet reviewed by native speakers.

How big is that layer? It depends on what you count:

  • Japanese

    Chinese-derived
    27.3%
    What was counted
    581 of the 2,131 words in its WOLD list
  • Vietnamese

    Chinese-derived
    25.5%
    What was counted
    391 of 1,534 WOLD words, Old Chinese included
  • Korean

    Chinese-derived
    57.12%
    What was counted
    251,478 of 440,262 dictionary headwords; 25.28% are native

Sources: our counts on WOLD v4.2; for Korean, the Standard Korean Language Dictionary as counted by Jeong 2000, NIKL.

The WOLD lists cover about 1,460 everyday meanings, while the Korean figure counts a full dictionary, technical terms included.

That's why figures of 50–70% circulate online: a dictionary count weighs in thousands of rare, formal nouns.

The Korean dictionary also shows where the layer sits:

Part of the Korean dictionarySino-Korean share
Nouns62%
Particles, all 356None
Verb endings, all 2,523None

Source: Jeong 2000, NIKL.

The words transfer; the grammar doesn't.

Even shared words need care. The sounds have drifted for centuries, and some meanings moved: in Japanese, 手紙, read tegami, is a letter, while Mandarin 手纸(shǒuzhǐ) is toilet paper.

One word did travel everywhere. In all nine languages, the word for tea comes from Chinese, WALS shows:

  • Mandarin, Japanese, Korean, Vietnamese and Arabic took cha, as in Mandarin 茶(chá).
  • English, Spanish, French and German took Min Nan te. English got it through Dutch thee, according to WOLD.

Only Japanese still writes with Chinese characters

Mandarin is written in characters, each usually a syllable with a meaning, and pinyin is the Latin-letter bridge for learners. Of the eight languages, only Japanese still writes everyday text with Chinese characters.

Here's how each one writes today:

LanguageScriptChinese characters today
JapaneseKanji plus two syllabaries2,136 official everyday kanji
KoreanHangul1,800 basic hanja taught for classical Chinese
VietnameseLatin alphabet, quốc ngữReplaced Classical Chinese and chữ Nôm
English, Spanish, French, GermanLatin alphabetNone
ArabicA right-to-left abjadNone

Sources: Agency for Cultural Affairs, 2010; Kim 2015; Vietnamese Nôm Preservation Foundation.

For Japanese readers, shared characters often keep their meaning, but the readings differ completely, and many characters were simplified differently in China and Japan. "Library" is 图书馆(túshūguǎn) in Mandarin and 図書館, read toshokan, in Japanese.

Korea's list of basic hanja was last adjusted in 2000. Knowing hanja helps with Mandarin, but daily reading in Korean doesn't use them.

Vietnamese was written for centuries in Classical Chinese characters and in chữ Nôm, a character script built for Vietnamese, before the Latin-based quốc ngữ replaced them. So Vietnamese learners get the vocabulary but not the writing.

How many characters does a Mandarin reader need? In film and TV subtitles, 1,111 characters cover 95% of running text.

To cover 99%, you need 2,076, as our characters guide found using SUBTLEX-CH. For the two script standards in China and Taiwan, see Simplified vs Traditional Chinese.

Distance only partly predicts learning time

These distances line up only partly with official learning times, and those times exist only for English speakers. Here's how long the US Foreign Service Institute budgets for professional speaking and listening:

Closest to Mandarin in grammar, Vietnamese takes half the time

Weeks of full-time training for English speakers to reach professional speaking and listening, ILR level 3

LanguageWeeks
Spanish30
French30
German~36
Vietnamese~44
Mandarin88
Japanese88
Korean88

Source: US Department of State, FSI language categories, checked October 2026. Cantonese and Arabic share Mandarin's 88 weeks, which come to 2,200 class hours.

Set that next to the grammar distances. Vietnamese is the closest of the eight to Mandarin in grammar, yet for an English speaker it's in Category III, while Japanese and Korean are in Category IV like Mandarin. Writing systems and vocabulary weigh at least as much as grammar.

FSI publishes no estimates for speakers of other languages, so there's no official number for a Korean or Spanish speaker learning Mandarin. Our guides to whether Chinese is hard to learn and how long it takes to learn Mandarin cover difficulty from the learner's side and the hours.

Two caveats on the FSI page itself:

  • Its hours don't multiply out. 88 weeks at the stated 23 class hours a week is 2,024 hours, not 2,200.
  • It changed recently. Between late 2024 and 2026, its goal changed from speaking and reading to speaking and listening, and so did its hours for Categories I to III.

What these numbers don't mean

  • Distance isn't difficulty. A feature you lack can be easy, like having no gender to learn, or hard, like tones. The cards label each one, and the labels are our reading of the data.
  • The feature list changes the answer. English moves from near the bottom to the middle depending on which WALS features you count. We publish both sets and the 95% ranges.
  • WALS codes the dominant pattern of a standard variety. Arabic here is Egyptian Arabic, the best-covered Arabic in WALS, with 145 features against 30 for Modern Standard Arabic. Modern Standard Arabic puts the verb first and question words at the front.
  • LDND is coarse. Scores in the 90s don't separate distant relatives from unrelated languages: English and Spanish, both Indo-European, score 94.1.
  • Loanword shares depend on the word list. WOLD's everyday list and a full dictionary give very different percentages for the same language.
  • No learner data. None of this measures how fast real learners progress.

What to do next

  • Pick your language in the tool above. See which features will feel familiar and which will need rewiring.
  • Train tones from the start, unless you speak Vietnamese. Of the eight, only Vietnamese has lexical tone, as Mandarin does.
  • Use your borrowed words if you speak Japanese, Korean or Vietnamese. Expect familiar meanings in new sounds, and watch for false friends.
  • Expect the first lessons to feel new, even from English. The shared structure is quiet; tones, measure words and word order aren't.
  • Hear it for yourself. The quickest way to feel the difference is to hear the tones and say a few sentences out loud.

FAQ

Is Chinese similar to Japanese?

In vocabulary and writing, yes; in sound and sentence structure, much less. More than a quarter of the words in WOLD's Japanese list came from Chinese, and Japanese still writes them with kanji. But Japanese puts the verb last and inflects it, and Mandarin doesn't: on 36 learner features from WALS, the two differ on 46%. Basic words like "eye" and "water" aren't related at all.

Can Japanese people read Chinese?

Partly. Japanese uses 2,136 everyday kanji, so a Japanese reader can often guess the meaning of Mandarin compounds like "culture", 文化(wénhuà). But mainland Mandarin uses simplified forms that often differ from Japanese ones, as in "library", 图书馆(túshūguǎn), which Japanese writes 図書館. Grammar words are different, and some shared words are false friends: Mandarin 手纸(shǒuzhǐ) means toilet paper, while Japanese 手紙, read tegami, means a letter.

Does Korean still use Chinese characters?

Mostly not for everyday writing: modern Korean is written in hangul. South Korean schools still teach hanja, from a list of 1,800 basic characters for classical Chinese. The vocabulary is a different story. 57% of the headwords in the Standard Korean Language Dictionary are Sino-Korean, so many Mandarin words come with a familiar meaning.

Not by descent: Vietnamese is Austroasiatic, and Mandarin is Sino-Tibetan. But Vietnamese borrowed heavily, with 25.5% of the WOLD Vietnamese list coming from Chinese. Its grammar is the closest to Mandarin of the eight languages we compared, with tones, classifiers and no verb endings. It was written in character scripts for centuries before switching to the Latin alphabet.

Is Chinese grammar similar to English grammar?

On the surface, partly: both put subject, verb and object in that order, and neither has case endings on nouns. Across all shared WALS features, they differ on 46%, about as much as Mandarin and Japanese. But the parts a learner meets first differ on 78% of the learner features we checked: tones, measure words, aspect particles and modifiers before the noun.

What language is closest to Mandarin?

Other Chinese languages, such as Cantonese, which differs on 19% of shared WALS features and shares far more basic vocabulary. Among non-Chinese languages in this comparison, Vietnamese is closest in grammar, with 33% of shared features differing, followed by Korean and Japanese. None of them shares basic vocabulary with Mandarin beyond chance. See Mandarin vs Cantonese.

Is Japanese a tonal language like Chinese?

Not in the same way. WALS codes Japanese as a "simple tone system": pitch accent marks where the pitch falls in a word, and it can tell words apart: "chopsticks", 箸, and "bridge", 橋, are both read hashi but differ in pitch. Mandarin gives every syllable one of four tones or a neutral tone, so 妈(mā), 麻(má), 马(mǎ) and 骂(mà) are four different words.

Does Chinese have tenses?

No. WALS codes Mandarin with no past tense and no inflectional future. Time comes from words like "yesterday", 昨天(zuótiān), and "tomorrow", 明天(míngtiān), while particles describe the shape of an action: 了(le) marks it as completed, and 过(guo) as experienced. The verb itself never changes form.

How we did this

Grammar distance is the share of WALS features whose values differ between two languages, over the features coded for both. We used two sets: the 36 learner features, all listed in the matrix above, and every feature coded for at least seven of the nine languages, leaving out six word-meaning features. The 95% ranges come from resampling the features 2,000 times.

Basic-word distance is LDND on ASJP's 40-item lists, using the first listed form per meaning. LDND is the Levenshtein distance, normalised by word length and divided by the average distance between words of different meanings, times 100.

The ASJP doculects were MANDARIN, ENGLISH, SPANISH, FRENCH, STANDARD_GERMAN_2, TOKYO_JAPANESE, KOREAN, VIETNAMESE, STANDARD_ARABIC_2, CANTONESE and THAI.

Loanwords are the WOLD words coded "clearly borrowed" or "probably borrowed", with the source language recorded for Chinese loans.

Sounds come from PHOIBLE inventories, several per language. We report ranges, and use one inventory per language for the sound table.

Cards: the "feels familiar", "easier", "new" and "rewire" labels are our editorial reading of the WALS values, which every card shows.

Data versions:

DatasetVersion
WALSv2020.6
ASJPv21.1
WOLDv4.2
PHOIBLEdevelopment snapshot 5f82b9c
CC-CEDICT2025-11-29

Data licence: the comparison data behind the tools on this page is shared under CC BY 4.0, derived from WALS, ASJP and WOLD, and CC BY-SA 3.0 for the sound groups, derived from PHOIBLE, with the credits below.

Sources

  • Dryer, Matthew S. & Martin Haspelmath (eds.). 2013. WALS Online (v2020.6). Zenodo. CC BY 4.0. https://wals.info. Features cited: 1A, 11A, 12A, 13A, 22A, 30A, 34A, 55A, 65A, 66A, 81A, 84A, 90A, 92A, 93A, 118A and 138A.
  • Wichmann, Søren, Eric W. Holman & Cecil H. Brown (eds.). The ASJP Database (v21.1). CC BY 4.0. https://asjp.clld.org
  • Haspelmath, Martin & Uri Tadmor (eds.). 2009. World Loanword Database (v4.2). CC BY 4.0. https://wold.clld.org
  • Moran, Steven & Daniel McCloy (eds.). PHOIBLE 2.0, development snapshot 5f82b9c. CC BY-SA 3.0. https://phoible.org
  • CC-CEDICT, MDBG. CC BY-SA 4.0. https://www.mdbg.net/chinese/dictionary?page=cc-cedict
  • Jeong Ho-seong. 2000. 표준국어대사전 수록 정보의 통계적 분석 [A statistical analysis of the Standard Korean Language Dictionary]. Saegugeosaenghwal 10(1). National Institute of the Korean Language. https://www.korean.go.kr/nkview/nklife/2000_1/2000_0104.pdf
  • Agency for Cultural Affairs, Japan. 2010. 常用漢字表 (Cabinet Notice No. 2).
  • Kim Wang-kyu. 2015. Status and trends of education in Chinese characters in Korea. Han-ja Han-mun Education 36.
  • Vietnamese Nôm Preservation Foundation. What is Nôm?
  • US Department of State. Foreign Language Training (FSI language categories). Checked 2026-10-06.
  • Cai, Qing & Marc Brysbaert. 2010. SUBTLEX-CH. PLoS ONE 5(6): e10729.

Methods

This analysis uses four open datasets: WALS v2020.6, ASJP v21.1, WOLD v4.2 (all CC BY 4.0) and PHOIBLE (CC BY-SA 3.0). Mandarin forms were checked against CC-CEDICT and other-language forms against Wiktionary. It has not yet been reviewed by native speakers of the other eight languages. If you spot a mistake, tell us: we fix errors.

Mei Lin, SayMei's AI teacher, smiling and waving

Ready to get speaking?

Talk with Mei Lin, your AI teacher, free for 10 minutes.

Start speaking

Mei Lin is an AI tutor. No account or card needed.