Simplified vs Traditional Chinese: the history, the differences, and which to learn
Simplified or Traditional? 35% to 41% of the 1,000 most common characters differ. Where each script is used, how they split, and which to learn first.

Write the same Chinese sentence in Simplified and in Traditional characters, and about a quarter of the characters change. The language underneath doesn't.
If you're about to start, moving to Taiwan or Hong Kong, or confused by "Mandarin vs Simplified", this guide shows how the scripts differ, where each is used and which to learn first.
Key takeaways
- Same language, two scripts. The words, grammar and pronunciation stay the same; the shapes of some characters change.
- Over a third of common characters look different. The exact share depends on whose standard you use, China's conversion table or Taiwan's forms.
- About a quarter of a typical page changes. The most frequent characters are the ones least likely to differ.
- Switching is cheaper than it looks. Many characters change in shared parts, so they convert in batches.
- One direction is harder. Going from Simplified to Traditional, 96 merged characters each split into two or more.
- Start with the script of your place. With no place in mind, start with Simplified, the script of the HSK and mainland China.
Read a story in Traditional. SayMei's graded stories are free to read: open any book, tap the gear, and set Characters to Traditional. Open the stories
What's the difference between Simplified and Traditional Chinese?
They're two standard ways to write one language, with the same words, grammar and pronunciation. What changes is the shape of some characters. "To study", 学, is 學 in Traditional, and "to speak", 说, is 說. "Person", 人, is the same in both.
Simplified characters, 简体字, are the forms the People's Republic of China standardised from 1956 on. Traditional characters, 繁體字, are the older forms, still the standard in Taiwan, Hong Kong and Macau.
Taiwan calls them "standard characters", 正體字. That's the label on the registration form for its national Chinese test, the TOCFL.
So "Mandarin or Simplified?" isn't a real choice. Mandarin is a spoken language, and you can write it in either script. Cantonese is a different spoken language, and in Hong Kong it's written in Traditional characters too. Our Mandarin vs Cantonese guide covers that choice.
One thing the labels hide: switching characters doesn't switch vocabulary. "Taxi" is 出租车 in Beijing, but usually 計程車 in Taiwan.
A converter that only swaps characters turns 出租车 into 出租車: correct Traditional, but not the usual Taiwan word. OpenCC, the open-source converter most apps use, keeps a separate Taiwan phrase list for exactly this reason.
Over a third of common characters differ
More than a third of the 1,000 most frequent characters have a different Traditional form. The exact share depends on which standard you compare against:
| Of the 1,000 most frequent characters | Characters that differ |
|---|---|
| By China's 2013 conversion table | 352 (35.2%) |
| By the standard forms used in Taiwan | 408 (40.8%) |
Source: our count, with characters ranked by frequency in SUBTLEX-CH, Cai and Brysbaert 2010. How we measured.
Those 1,000 characters make up 94% of everyday text in the film-and-TV subtitle corpus we used to rank them. Here's how the share holds as you go further down the list:
How many common characters look different?
Of the 1,000 most frequent characters, 35.2% have a different Traditional form by China's 2013 table and 40.8% by Taiwan's standard forms. Those 1,000 characters make up 94% of everyday text.
Share of the most frequent characters whose Traditional form differs, from the top 10 to the top 5,000 (ratio scale along the bottom)
- China's 2013 table
- Taiwan's standard forms
| Top | Share of everyday text | Differ by China's table | Differ by Taiwan's forms |
|---|---|---|---|
| 10 | 23.14% | 30% | 30% |
| 50 | 46.52% | 32% | 34% |
| 100 | 58.17% | 38% | 41% |
| 250 | 74.28% | 34.4% | 36.8% |
| 500 | 85.02% | 33.4% | 37.4% |
| 1,000 | 93.97% | 35.2% | 40.8% |
| 1,500 | 97.36% | 34.1% | 39.7% |
| 2,000 | 98.86% | 34.2% | 40.2% |
| 2,500 | 99.53% | 33.5% | 39.4% |
| 3,000 | 99.81% | 33% | 38.9% |
| 4,000 | 99.97% | 31.9% | 37.3% |
| 5,000 | 100% | 31.3% | 36.2% |
Frequency ranks from SUBTLEX-CH (film and TV subtitles). Characters that always differ, never keeping their Simplified shape: 31.7% of the top 1,000 by China's table and 33.7% by Taiwan's forms; 29.8% and 32.4% of the top 5,000. Our count, 6 October 2026; the method is under How we measured.
Why two numbers? The Taiwan figure also counts two kinds of difference that China doesn't call simplification:
- Variant choices. In a separate 1955 decision, China kept one form among several variants. It kept "at", 于, where Taiwan writes 於.
- Printed shapes. Some characters differ only slightly in print, like "not have", 没, which Taiwan prints as 沒.
A learner switching scripts meets both kinds, so the Taiwan figure is the better guide to the work involved.
A typical page changes by about a quarter
On a real page, the share is lower. The most frequent characters are the ones least likely to change: the possessive 的, "is", 是, "I", 我, and "you", 你, are the same in both scripts.
Characters that always change make up 26% to 27% of the subtitle corpus, depending on the standard.
When we converted all of SayMei's graded stories, 25.7% of the characters changed.
Just five characters made up 23% of all the changes:
| Meaning | Simplified | Traditional |
|---|---|---|
| come | 来 | 來 |
| the plural for people | 们 | 們 |
| the general measure word | 个 | 個 |
| this | 这 | 這 |
| the second syllable of "what", 什么 | 么 | 麼 |
The share barely moves as you climb the HSK levels:
| HSK level (2025 syllabus) | Characters to read | Differ by China's table | Differ by Taiwan's forms |
|---|---|---|---|
| Level 1 | 246 | 32.1% | 35.8% |
| Levels 1–3 | 655 | 35.6% | 39.4% |
| Levels 1–6 | 1,940 | 34.5% | 40.4% |
| Levels 1–9 | 3,088 | 33.9% | 39.7% |
Source: character lists from the new HSK syllabus; differences are our count.
Simplified does save ink: about 19% fewer strokes in everyday text, by our count on Unihan's stroke data. The savings pile up on the most frequent characters:
Simplification made handwriting official
Mostly, simplification made official the shortcuts people already wrote by hand. Then it applied each simplified part to every character that contains it.
The first scheme simplified hundreds of characters and dozens of components. In a joint notice of 7 March 1964, China ruled that most of them also apply when they appear inside other characters.
Here's the usual classification of the methods, as Wikipedia's article on simplified characters summarises it:
| Method | Examples |
|---|---|
| Handwriting shapes made official (草书, cursive) | 书/書 shū "book", 东/東 dōng "east", 长/長 cháng "long" |
| Keep one part, drop the rest | 广/廣 guǎng "wide", 飞/飛 fēi "to fly", 习/習 xí "to practise" |
| A simple mark replaces a busy part | 对/對 duì "correct", 汉/漢 hàn "Chinese", 鸡/雞 jī "chicken" |
| A simpler sound part | 远/遠 yuǎn "far", 运/運 yùn "to move", 进/進 jìn "to enter" |
| A new sound-and-meaning character | 护/護 hù "to protect", 惊/驚 jīng "to startle", 响/響 xiǎng "loud" |
| An old or folk form brought back | 从/從 cóng "from", 众/眾 zhòng "crowd", 电/電 diàn "electricity" |
| Several characters become one | 发 for 發 "to send out" and 髮 "hair"; 后 for 後 "after" and 后 "empress" |
| One part simplified everywhere | 言 → 讠 in 说/說, 话/話 huà "speech"; 門 → 门 in 问/問 wèn "to ask" |
The last row is why switching is less work than the percentages suggest. Most of China's 1964 General List was built that way, from a short list of simplified parts:
| China's 1964 General List | Count |
|---|---|
| Characters in the list | 2,238 |
| Built by applying simplified parts to other characters | 1,754 |
| Simplified characters used as parts | 132 |
| Simplified components used as parts | 14 |
Source: General List, 1964.
Once you know that the "speech" part, 讠, is 言, and the "metal" part, 钅, is 金, dozens of characters convert themselves. Explore 80 common pairs, grouped by how they were simplified:
How characters were simplified
80 common pairs, grouped by how they were simplified, with stroke counts and frequency ranks, plus China's 96 one-to-many groups.
Shapes people already wrote fast with a brush (草书 cǎoshū) were turned into print forms.
Readings and examples checked against CC-CEDICT. Methods overlap; each pair is filed under its main method.
One Simplified character can be two Traditional ones
Simplification merged characters that sounded alike, so going back needs the whole word. 发 stands for both "to send out", 發, and "hair", 髮:
- "to discover", 发现, becomes 發現
- "hair", 头发, becomes 頭髮
China's 2013 conversion table splits 96 groups like this. Its official explainer warns that converting "dipper", 斗, blindly gives 北鬥星 for the Big Dipper, instead of 北斗星.
Half of them, 48 groups, involve characters among the 1,000 most frequent, by our count, so learners meet them early:
| Simplified | Traditional forms |
|---|---|
| 里 | 裡 "inside", or 里, a unit of distance |
| 只 | 只 "only", or 隻, a measure word for animals |
| 干 | 乾 "dry", or 幹 "to do" |
| 面 | 面 "face", or 麵 "noodles" |
| 台 | 臺 "stage", 颱 "typhoon" or 檯 "table" |
Good converters read the whole word, and most of the time they get it right. We tested the one SayMei uses for its Traditional view, opencc-js, on all 166 of our graded stories.
It got about 1 character in 1,260 wrong, by our audit. The mistakes cluster in a few patterns:
- A longer wrong word matches first. In "will it come clean after washing?", 洗了还能干净吗, the converter reads 能干, "capable". So it writes 能幹淨 instead of 能乾淨.
- No helpful neighbour. In "the noodles are tasty", 面很好吃, nothing next to 面 says "noodles", so it stays 面, "face". In our noodle-shop story, 57 of the 74 面 that mean "noodles" were left as 面.
- A wrong but real word. "In the sea", 海里边, contains "nautical mile", 海里, so 里 isn't changed to 裡.
- Names. 曹冲, a boy in a well-known story, became 曹衝 instead of 曹沖. In our own tests, the surname 余 became 餘, "surplus".
Try it on five short paragraphs:
Flip a paragraph between the scripts
In five short paragraphs, 49 of 147 characters change in Traditional, and the converter SayMei's Traditional view uses wrote 8 of them wrong.
- A haircut and a bowl of noodles. 我昨天去理发了。理发师说我的头发太干了。后来我去吃了一碗面条,面很好吃。 → 我昨天去理髮了。理髮師說我的頭髮太乾了。後來我去吃了一碗麵條,麵很好吃。 I got a haircut yesterday. The barber said my hair was too dry. Afterwards I had a bowl of noodles, and the noodles were delicious. 12 of 32 characters change. The converter wrote 面 for 麵.
- Grandma's clock. 墙上的钟停了,我只好看手表。那个钟是我奶奶的,她只有这一个钟。 → 牆上的鐘停了,我只好看手錶。那個鐘是我奶奶的,她只有這一個鐘。 The clock on the wall stopped, so I had to look at my watch. That clock was my grandmother's; she only had this one clock. 8 of 27 characters change. The converter wrote 鍾 for 鐘 (2 times).
- Laundry on a rainy day. 下雨了。我的衣服还没干,这件白衬衫洗了还能干净吗? → 下雨了。我的衣服還沒乾,這件白襯衫洗了還能乾淨嗎? It is raining. My clothes are not dry yet. Can this white shirt still come clean after washing? 9 of 22 characters change. The converter wrote 幹 for 乾 (2 times).
- From SayMei Stories: The Little Noodle Shop. 我想给她做一碗长长的面。可是我不会。那是长寿面。面很长,人也有很多很多岁。可是面不能断。断了,就不好。 → 我想給她做一碗長長的麵。可是我不會。那是長壽麵。麵很長,人也有很多很多歲。可是麵不能斷。斷了,就不好。 I want to make her a bowl of long, long noodles, but I don't know how. Those are longevity noodles. The noodles are long, and people live many, many years. But the noodles must not break. If they break, it's bad luck. 14 of 43 characters change. The converter wrote 面 for 麵 (3 times).
- Same characters, different words. 我坐出租车去地铁站,然后用手机上的软件打印了车票。 → 我坐出租車去地鐵站,然後用手機上的軟件打印了車票。 I took a taxi to the subway station, then printed the ticket with an app on my phone. 6 of 23 characters change. The converter got every character right. Every character here is converted correctly, but a Taiwanese writer would more likely say 計程車 (taxi), 捷運 (MRT), 軟體 (software) and 列印 (print). Switching characters does not switch vocabulary.
What a converter can't do:
- It converts characters only, not vocabulary: 出租车 becomes 出租車, not the Taiwan word 計程車.
- Phrase matching is greedy, so a longer wrong phrase can win: 能干净 becomes 能幹淨, 海里边 becomes 海里邊, 那个钟 becomes 那個鍾.
- A character with no helpful neighbours falls back to a default: a bare 面 stays 面 ("face", not "noodles"), and a bare 云 stays 云.
- Names are guesses: 曹冲 becomes 曹衝 (it should be 曹沖), and the surname 余 becomes 餘.
- Taiwan and Hong Kong standards differ (說/説, 裡/裏, 麵/麪); these examples follow Taiwan.
- Traditional to Simplified is mostly safe, but a few characters keep their Traditional form in some words (乾坤, 著作, 瞭望).
Traditional means Taiwan's standard forms, checked word by word against CC-CEDICT. Converter: opencc-js 1.4.2, Simplified to Taiwan standard.
The lesson for learners isn't to avoid converters. It's to know which characters to double-check: the 96 merged ones, and a dozen more that Taiwan keeps apart, such as 游, which Taiwan writes 游 for "swim" and 遊 for "travel". Our flipper underlines them.
Where is each script used?
Simplified is the official standard in mainland China and Singapore, and Malaysia's Chinese-medium schools have taught it since 1983. Traditional is the standard in Taiwan, Hong Kong and Macau. Everywhere else, it depends on the community, and most large communities use both.
Pick a place to see its script, its Chinese population and the sources:
Where each script is used
Simplified is official in mainland China and Singapore; Traditional is the standard in Taiwan, Hong Kong and Macau. Malaysia teaches Simplified but lives with both, and large communities abroad mix the two.
- Mainland ChinaSimplified
- TaiwanTraditional
- Hong KongTraditional
- MacauTraditional
- SingaporeSimplified3.0 million Chinese residents, 74.3% of residents (2020)
- MalaysiaSimplified + bothabout 6.9 million (22.2% of 30.9 million citizens) (2025)
- United StatesBoth5,205,461 Chinese (except Taiwanese), alone or in combination (2020)
- CanadaBoth1.7 million (4.7% of Canadians) (2021)
- AustraliaBoth5.5% report Chinese ancestry (2021)
- United KingdomBoth445,619 in England and Wales (0.7%) (2021)
- IndonesiaSimplified + both2,832,510 (about 1.2%) (2010)
- PhilippinesBoth
Chinese community: the latest census or survey for each place; the bar compares community sizes. Map data: Natural Earth (public domain), via world-atlas.
Sources: China's language law; Taiwan Ministry of Education; Hong Kong Education Bureau; Singapore lists, West 2009; Malaysia, Chinese Wikipedia citing 《语文建设》1982; censuses: Singapore, Malaysia, US, Canada, England and Wales, Australia; Indonesia and the Philippines via Wikipedia. All checked 6 October 2026.
Malaysia is the clearest mixed case. It published a list of simplified characters identical to China's in 1981, yet several Chinese-language dailies print Traditional headlines over Simplified text, according to Chinese Wikipedia, citing a 1982 journal report.
In the United States, weekend Chinese schools belong to two national networks, both founded in 1994. One was started by immigrants from mainland China, and the other mainly by immigrants from Taiwan, says the Center for Applied Linguistics. If you're choosing a school for a child, ask which script it teaches.
Traditional isn't a single standard, either. Taiwan's Ministry of Education fixed its forms in 1982, and Hong Kong compiled its own school reference in 1986.
That reference, Hong Kong's list of standard forms for common characters, 常用字字形表, was re-typeset in 2007, says the Education Bureau.
Among the 1,000 most frequent characters, 9 have different default forms in the two places, by our count. "To speak", shuō, is written 說 or 説, and "inside", lǐ, is 裡 or 裏.
A short history of simplification, 1935 to 2026
China made simplification law in 1956, but the idea is older. The Republic of China published, then halted, its own short list in the 1930s:
| Date | What happened |
|---|---|
| January 1934 | Linguist Qian Xuantong (钱玄同) submits a draft list of simplified characters to the Republic of China's national-language committee (West 2009) |
| August 1935 to February 1936 | The Republic of China publishes 324 simplified characters, then halts the programme (West 2009) |
| January 1955 | The People's Republic publishes a draft scheme; about 200,000 people take part in the discussion (State Council, 1956) |
| 28 January 1956 | The State Council adopts the scheme: 515 characters and 54 components, with the first 230 characters standard from 1 February (State Council, 1956) |
| May 1964 | The General List of Simplified Characters: 2,238 characters, with components simplified by analogy (General List, 1964) |
| 1969 to 1976 | Singapore issues 502 simplified characters (1969), about 2,250 (1974), and a list identical to the mainland's (May 1976) (West 2009) |
| 20 December 1977 | A second round of 853 simplified characters is published; the People's Daily uses its first table until July 1978 (draft text; West 2009) |
| 1981 to 1983 | Malaysia publishes a list identical to China's on 28 February 1981, and its Chinese primary schools switch from 1983 (Chinese Wikipedia) |
| September and December 1982 | Taiwan's Ministry of Education issues standard forms for 4,808 common and 6,341 less common characters (Taiwan MOE) |
| 1986 | Hong Kong compiles its standard character forms, 4,721 characters (Education Bureau) |
| 24 June 1986 | China formally withdraws the second round and calls for caution about any further simplification (State Council notice) |
| 10 October 1986 | The General List is republished with small changes: 2,235 characters (1986 note) |
| 5 June 2013 | The Table of General Standard Chinese Characters, 8,105 characters, replaces the older lists; no Traditional character is restored (State Council; Ministry of Education) |
| 1 January 2026 | China's revised language law takes effect; Article 19 lists the six cases where Traditional characters may be kept or used (law text) |
Why simplify? The aim was mass literacy in a country where most adults couldn't read, as Wikipedia's history summarises it.
How much the new shapes helped, as opposed to schooling, is still argued. Taiwan and Hong Kong reached high literacy with Traditional characters, as the debate over the two scripts points out.
The 1977 second round shows the limits. Its forms were mostly new inventions rather than shapes people already wrote, and the People's Daily used them for about seven months before dropping them.
Since 1986, the official line has been stability. Traditional characters are allowed in China for set purposes, such as historic sites, calligraphy and hand-written shop signs. Simplified is the standard everywhere else, under Article 19 of the language law.
Taiwan, Hong Kong and Macau never switched
None of them was governed from Beijing when the reforms happened.
The Republic of China had tried, then halted, its own short list in the 1930s. After its government moved to Taiwan in 1949, it kept the characters it already used.
Hong Kong and Macau were under British and Portuguese administration until 1997 and 1999, as Hong Kong's Basic Law and the Macau handover record.
In Taiwan, the script became a political line. A 1993 article in Taiwan Panorama, a government-funded magazine, quoted a Taipei professor saying "simplified characters are like the Red Guards".
The same article reported that a cross-strait agreement signed that spring was printed in both scripts. Taiwan standardised its own character forms in 1982.
Hong Kong's Basic Law makes Chinese and English official languages without naming a script. China's national laws apply in Hong Kong only if they're listed in its Annex III, and the national language law isn't on that list. Hong Kong's teachers, parents and editors work from the Education Bureau's own list of standard Traditional forms.
Which should you learn?
Learn the script of the place and the people you'll read with. If there's no such place yet, start with Simplified. It's the script of the HSK and of mainland China, and adding Traditional reading later is a smaller job than it looks.
Live, work or study in mainland China; take the HSK
- Learn first
- Simplified
- Why
- The HSK syllabus is in Simplified (CTI)
Live or study in Taiwan; take the TOCFL
- Learn first
- Traditional
- Why
- Taiwan's standard; the TOCFL offers a Traditional or a Simplified paper (TOCFL form)
Family in Hong Kong or Macau
- Learn first
- Traditional, and think about Cantonese
- Why
- Both use Traditional; the spoken language at home may be Cantonese (Mandarin vs Cantonese)
Singapore or Malaysia
- Learn first
- Simplified, then learn to read Traditional signs
- Why
- Schools use Simplified; signs and headlines often use Traditional
Older books, calligraphy, temple inscriptions
- Learn first
- Traditional
- Why
- Books and inscriptions from before the 1950s use the older forms
No particular place
- Learn first
- Simplified, then Traditional for reading
- Why
- The HSK and mainland materials use it; reading Traditional comes later
Switching costs less than it looks
About a third of the 1,000 most common characters look different. But many of those changes come from shared components, so they arrive in batches.
The hard part runs one way. Going from Traditional to Simplified, merged characters collapse: "to send out", 發, and "hair", 髮, both become 发, and there's nothing to choose.
Going from Simplified to Traditional, you have to learn which form each word takes: the 96 merged groups and a few Taiwan-only distinctions.
Olle Linge of Hacking Chinese estimates "around five hundred tricky cases". He advises waiting until you know a few thousand characters before adding the second set.
Traditional takes more strokes, not more of everything
A Traditional character takes about a quarter more strokes in everyday text, as the stroke chart above shows. Most of what makes Chinese hard is the same in both scripts: vocabulary, tones, and the first 2,000 to 3,000 characters.
Learning Traditional with SayMei
Since 4 October 2026, you can switch SayMei to Traditional characters. The setting is called Chinese characters in Settings. It's called Characters in a story's reading settings, the gear, and in lesson settings, flashcard settings and the Lightning Links game.
Visitors whose connection or browser suggests Taiwan, Hong Kong or Macau are asked once which script they want, as the SayMei changelog notes.
What changes is the Chinese text on screen, including stories, lesson transcripts and the whiteboard, flashcards and Lightning Links. The Traditional worksheet and the level test have their own Traditional options.
Where SayMei falls short, plainly:
- Characters, not vocabulary. You'll see 出租車, not the Taiwan word 計程車.
- The converter isn't perfect. Our audit found 273 wrong characters across our stories, and we're fixing the patterns it found.
- Mainland Mandarin only. Mei Lin's voice and the lessons follow the mainland standard and the HSK. There's no Zhuyin, also called bopomofo, and no Cantonese.
- Tools and dictionary pages stay as written, so a dictionary page shows both forms rather than switching.
What to do next
- Pick the script of your place. Learn the one used by the place and the people you'll read with.
- Default to Simplified if you're unsure. It's the script of the HSK and of mainland China.
- Add Traditional reading later. Learn the shared components first, and whole groups of characters switch at once.
- Double-check the merged characters. Converters and learners slip on the same few, like 里, 面 and 干.
- Read in your target script. Switch a story to Traditional and see how much you can already read.
Try a story in Traditional. Free, no account: open any book, tap the gear, choose Traditional. Open the stories
FAQ
Is Mandarin Simplified or Traditional?
Neither. Mandarin is a spoken language, and both scripts write it. Mainland China and Singapore write Mandarin in Simplified, and Taiwan writes it in Traditional. The words and grammar are the same; only the shapes of more than a third of the 1,000 most common characters differ, by our count.
Is Cantonese written in Simplified or Traditional?
Mostly Traditional, because Hong Kong and Macau, where most written Cantonese comes from, never adopted the 1956 reform. Cantonese speakers in Guangdong learn Simplified at school, like the rest of mainland China. Script and spoken language are separate choices, and our Mandarin vs Cantonese guide covers the second.
Does the HSK use Simplified characters? Is there a Traditional test?
CTI's HSK syllabus lists its characters in Simplified at every level. China's revised language law, in force since January 2026, says international Chinese teaching should use the national standard script, under Article 22. Taiwan's TOCFL lets you choose a Traditional or a Simplified paper, as its registration form shows.
Do Taiwan and Hong Kong write Traditional the same way?
Almost. Each has its own standard, set by Taiwan's Ministry of Education and Hong Kong's Education Bureau. Among the 1,000 most common characters, only 9 have different default forms, such as "to speak", written 說 or 説, by our count. Everyday vocabulary differs more than the characters do.
Is Traditional Chinese harder to learn?
It takes more strokes: a Traditional character in everyday text averages 8.8 strokes against 7.1 for Simplified, about a quarter more, by our count. Reading is less affected than handwriting, and the rest of the work, vocabulary, tones and the first few thousand characters, is the same in both scripts.
Can people in mainland China read Traditional characters?
We found no official survey. What we can measure: about 60% of the 1,000 most common characters are identical in both scripts, and many of the rest differ in predictable parts, by our count. Writing Traditional is rarer: schools teach Simplified, and China's law allows Traditional only in set cases such as calligraphy, historic sites and hand-written signs, under its language law.
Should I learn both scripts at the same time?
Usually not. Pick one to read and write, and add reading of the other later. Olle Linge of Hacking Chinese advises waiting until you know a few thousand characters, when the switch is quick. Our count shows why it's manageable: 59% to 65% of the most common characters are identical in both scripts.
Can I trust a Simplified-to-Traditional converter?
Mostly. The converter we use got about 1 character in 1,260 wrong across our graded stories, by our audit. The mistakes clustered in merged characters such as 里, 面 and 干, and in 游, which Taiwan splits in two. Check those, and remember that converters swap characters, not regional vocabulary.
How we measured
Which characters differ. We ranked characters by frequency in SUBTLEX-CH, a corpus of 46.8 million characters of film and TV subtitles, from Cai and Brysbaert 2010. Then we compared each character with two standards:
- China's 2013 correspondence table, Attachment 1 of the Table of General Standard Chinese Characters, from the table text.
- The Taiwan standard forms, from OpenCC's character dictionary and Taiwan variant list.
A character "differs" if any of its Traditional forms differs from it, and "always differs" if it never stays the same. We cross-checked the results four ways:
| Check | Result |
|---|---|
| Our parse of China's 2013 table | Reproduces its own totals: 2,546 characters with Traditional forms, 2,574 Traditional characters and 96 one-to-many groups |
| Jun Da's Modern Chinese frequency list, about 193 million characters | 36.8% of its top 1,000 differ by China's table, and 42.0% by Taiwan's forms |
| Unicode's Unihan database, as a second mapping | 38.2% of the top 1,000 differ |
| Where the two standards disagree | On 58 of the top 1,000: 57 differ only under Taiwan's forms, and 1 only under China's table |
The 57 Taiwan-only differences are 1955 variant choices such as 于/於, print-shape differences such as 没/沒, and a few rare alternatives OpenCC lists, such as 它/牠. The one China-only case is 才/纔.
Real text. We converted every line of SayMei's 166 graded stories, 344,086 characters at HSK 1–4, with the converter the app uses: opencc-js 1.4.2, Simplified to Taiwan standard, a sentence block at a time. Then we counted the characters that changed.
Converter accuracy. We checked the converted stories four ways, with these results:
| Check | Scope |
|---|---|
| CC-CEDICT dictionary entries for words containing a merged character | 32,834 words |
| The stories' own pinyin, for characters whose reading decides the form | 3,386 checks |
| Targeted scans | Each error pattern we found |
| A seeded random sample of tokens neither check had verified | 200 tokens: 9 errors, all in patterns already counted |
| Known errors | 273, or 0.08% |
| Upper bound, allowing for patterns we may have missed | 365, or 0.11% |
The labels are ours and haven't been independently checked yet.
Strokes come from Unicode's Unihan database, kTotalStrokes. HSK lists come from the new HSK syllabus.
Limits. Frequency lists reflect their corpora: subtitles over-represent conversation. "Taiwan forms" follows OpenCC's variant list, which may differ from a given Taiwan textbook in a handful of characters. Population figures come from different years and count different things, like ancestry, ethnicity or citizenship, so compare them with care.
Sources
- China: State Council resolution on the 1956 scheme; General List, 1964 and 1986 editions; second-round draft, 1977; State Council notice on the 2013 table; Ministry of Education explainers (1, 2); Law on the Standard Spoken and Written Chinese Language, revised 2025.
- Taiwan: Ministry of Education, standard character forms; TOCFL registration form; Taiwan Panorama, May 1993.
- Hong Kong: Education Bureau, 常用字字形表; Basic Law.
- Singapore, the 1930s and 1977: Andrew West, Unicode L2/09-260 (2009).
- Malaysia: Chinese Wikipedia, 馬來西亞華語.
- Populations: Singapore Census 2020; Malaysia DOSM 2025; US Census 2020; Statistics Canada 2026; GOV.UK, Census 2021; ABS 2021; Center for Applied Linguistics.
- Data: SUBTLEX-CH; Jun Da frequency list; OpenCC and opencc-js; Unihan; CC-CEDICT; HSK syllabus.
- Background: Wikipedia, Simplified Chinese characters and Debate on traditional and simplified Chinese characters; Hacking Chinese.
Methods
The research was done on 6 October 2026: every date comes from an official text or a named source, linked where it's used, and every number marked "our count" comes from our own scripts run on public data and on SayMei's stories. All Chinese examples were checked against the CC-CEDICT dictionary (6 October 2026 release). No learner data was used. If you spot a mistake, tell us: we fix errors.
