How to learn Chinese tones as an adult: what research shows
Train your ear first, practise two-syllable words, keep the third tone low and get feedback on your voice. The evidence for each step, plus a tone lab.

Adults who have never studied Mandarin can already tell its tones apart better than chance. The hard part comes later: storing each tone as part of the word.
I wrote this guide for adult English speakers, roughly HSK 1 to 3, who learn mostly on their own. You know the four tone marks, but tones 2 and 3 still blur. If you grew up hearing Mandarin, your starting point is different. If one stubborn error needs diagnosing, book a human tutor.
I link the studies behind every section. I can't hear your voice through this page, so wherever a human ear matters, I'll say so.
Key takeaways
- Your ear isn't the obstacle. Adults with no Mandarin tell tones apart better than chance, and they improve within a month of classes.
- Listening practice works fast. Two weeks of training lifted learners' tone identification by 21%, and the gain lasted.
- Using tones in real words takes years. Advanced learners accepted most words said with the wrong tone; native listeners rejected nearly all of them.
- The third tone is usually low. It dips and rises mainly when said alone or at the end of a phrase.
- Practise pairs, not just syllables. Tones change shape next to each other, so two-syllable words train them the way you'll use them.
- True tone deafness is rare. It affects about 1.5% of people, and even they learned new tones in training.
Tones are hard for your memory, not your ear
English uses pitch across whole sentences, to ask a question or stress a word. So an adult English ear has never had to file a syllable's pitch as part of the word itself.
That makes the hard part listening and word memory, not your voice. The best evidence I found comes from people who got very far.
In one study, 16 English speakers with about 11 years of Mandarin identified tones on single syllables almost as well as native listeners. Then they judged whether syllables were real words. When only the tone was wrong, they usually missed it:
Source: Pelzl et al. 2019, Tables 1, 2 and 4. Naming tone 2: identifying the tone of a single tone 2 syllable. Spotting fakes: rejecting a made-up word that differs from a real one only in tone. On single syllables, the learners matched native listeners on every tone except tone 2.
They could hear tones. What they hadn't built was the habit of storing the tone as part of each word.
Nor are beginners' ears the obstacle. Adults with no Mandarin at all tell tones apart better than chance before any study, and they improve after a month of classes.
Tones aren't decoration you can skip, though. Identifying a syllable's tone carries at least as much information as identifying its vowel. And here's what happened when researchers flattened the pitch of Mandarin sentences:
Source: Patel, Xu & Wang 2010. In quiet, flattened sentences were understood about as well as natural ones.
Context rescues you at a quiet table, not in a busy restaurant.
You're probably not tone deaf. Fewer than 2 in 100 people have true tone deafness, and even they can learn word tones. See the tone-deaf section.
Four tones and a light one
Mandarin has four tones, plus a light, unstressed one called the neutral tone. The shapes in our lab describe a syllable said on its own.
In ordinary speech, the third tone usually loses its rise, and every tone bends toward its neighbours. Linguists write tone shapes with Yuen Ren Chao's pitch numbers, from 1 at the bottom of your voice to 5 at the top, as PolyU's introduction to phonetics explains.
The four tones, and the light one
Each tone on one syllable, shi. In connected speech the third tone stays low unless it ends a phrase, and the light tone takes its pitch from the syllable before it.
The four tones and the light tone, on their own
Tone 1
- Pitch numbers
- 55
- What it does
- High and level
- Example
- teacher, 师 shī
Tone 2
- Pitch numbers
- 35
- What it does
- Rises from the middle
- Example
- time, 时 shí
Tone 3
- Pitch numbers
- 214 alone; 21 before tones 1, 2, 4
- What it does
- Low; dips and rises only when alone or at the end
- Example
- to begin, 始 shǐ
Tone 4
- Pitch numbers
- 51
- What it does
- Falls from the top
- Example
- to look at, 视 shì
Tone Light
- Pitch numbers
- No shape of its own
- What it does
- Short; pitch set by the syllable before it
- Example
- mom, 妈妈 māma
Tone 3 changes with its neighbour
Alone
- Example
- 始 to begin
- Written
- shǐ
- Said
- shǐ
- Pitch
- 214
Before tones 1, 2, 4
- Example
- 好吃 tasty
- Written
- hǎochī
- Said
- hǎochī
- Pitch
- 21
Before another 3
- Example
- 你好 hello
- Written
- nǐ hǎo
- Said
- ní hǎo
- Pitch
- 35
| Word | After | Height of the light syllable (1–5) |
|---|---|---|
| 妈妈 mom | tone 1 | 2–3 (sources differ) |
| 朋友 friend | tone 2 | 3 |
| 我们 we | tone 3 | 4 |
| 谢谢 thank you | tone 4 | 1 |
Pitch numbers use Chao's 1–5 scale, where 5 is the top of your voice. The recordings are SayMei's text-to-speech (computer voices). Hear every syllable in the pinyin chart, or check your own voice with the Tone Checker.
For a printable chart and audio for every tone pattern, use the tone-pair trainer. To hear any syllable in all its tones, open the pinyin chart.
Train your ear first, but speak from week one
My advice: start with your ears, and say every word out loud from the first day. Listening training improves your pronunciation only partly, because each skill improves mostly through its own practice.
The listening half is well supported. In a 1999 study, eight American learners trained on tone identification for eight sessions over two weeks. Here's how their accuracy rose:
Source: Wang, Spence, Jongman & Sereno 1999.
The training reached the learners' voices too. After the same kind of listening-only practice, native listeners identified the trainees' own spoken tones 18% more accurately.
The "partly" matters. Across 25 years of studies, perception training had a large effect on listening, and one about half as big on speaking.
In a study of 38 English speakers learning tone words, people who practised only listening or only speaking did far worse on the skill they hadn't practised. So listen first, but say every word out loud too.
Hearing many different voices sounds like an obvious upgrade, and it can help you recognise new speakers later. But the evidence is mixed:
- A 2021 meta-analysis of speech-sound training found that the immediate advantage of many voices shrank to nothing, once outliers were removed and publication bias corrected. The advantages for handling new voices and for long-term retention stayed large.
- In another study, many voices helped learners who already perceived pitch well, and slowed down those who didn't.
If tones feel impossible, I'd start with one clear voice.
Try our ear test below. The default round pits tone 2 against tone 3, the pair that causes the most trouble. Guessing would get you 5 out of 10, so anything clearly above that is your ear working.
Test your ear
Rounds of ten one-syllable words: play each one and pick the tone you heard. On a pair, guessing scores about 5 of 10, so a clearly higher score is your ear working.
The five rounds and what to listen for
2 vs 3
- Words
- mí 迷 / mǐ 米 · tú 图 / tǔ 土 · yán 颜 / yǎn 眼 · shí 时 / shǐ 始 · tóng 同 / tǒng 统 · fán 繁 / fǎn 反 · yí 姨 / yǐ 以
- What to listen for
- Tone 2 starts in the middle and climbs. Tone 3 drops low first and turns up late, if at all. Listen for how low it goes, and when it turns.
1 vs 4
- Words
- mā 妈 / mà 骂 · shī 师 / shì 视 · yī 衣 / yì 意 · tōng 通 / tòng 痛 · fān 翻 / fàn 饭 · bāo 包 / bào 报 · tū 突 / tù 兔 · wū 屋 / wù 物
- What to listen for
- Both start high. Tone 1 stays up and level; tone 4 drops fast to the bottom.
3 vs 4
- Words
- mǎi 买 / mài 卖 · shuǐ 水 / shuì 睡 · wěn 稳 / wèn 问 · gǒu 狗 / gòu 够 · mǎ 马 / mà 骂 · yǎn 眼 / yàn 验 · xiǎng 想 / xiàng 向
- What to listen for
- Tone 4 starts at the top and falls. Tone 3 starts low, sinks further and may rise at the end.
1 vs 2
- Words
- chī 吃 / chí 持 · yī 衣 / yí 姨 · shī 师 / shí 时 · tōng 通 / tóng 同 · fān 翻 / fán 繁 · tū 突 / tú 图 · yān 烟 / yán 颜
- What to listen for
- Tone 1 is high from the first moment. Tone 2 starts lower and climbs.
All four
- Words
- shī 师 / shí 时 / shǐ 始 / shì 视 · yī 衣 / yí 姨 / yǐ 以 / yì 意 · tōng 通 / tóng 同 / tǒng 统 / tòng 痛 · fān 翻 / fán 繁 / fǎn 反 / fàn 饭 · tū 突 / tú 图 / tǔ 土 / tù 兔 · yān 烟 / yán 颜 / yǎn 眼 / yàn 验
- What to listen for
- High and level (1), rising (2), low (3), falling (4).
Every word is one computer voice from SayMei's pinyin chart audio, kept only if its measured pitch clearly matched its tone. Eight short listening sessions raised learners' tone identification by 21% in one study (Wang et al., 1999). Next: the tone-pair trainer's 1-minute ear quiz, or Panda Tone Toss.
When single syllables feel easy, take the 1-minute ear quiz on all 20 two-syllable tone patterns, or play Panda Tone Toss.
Practise tones in two-syllable words
Tones change shape next to each other, and learners find them harder in two-syllable words than in single syllables. That's why I'd practise in pairs.
A tone is pulled toward the one before it. The end of one syllable strongly shapes the start of the next, while the following tone has a small opposite effect.
New learners in a training study learned tones less well in two-syllable words than in one-syllable words. Drilling shī shí shǐ shì one by one is a start. Saying tomorrow, 明天, and movie, 电影, trains tones the way you'll actually use them.
Our grid below has one real word for each tone pattern. Blue dots mark pairs where tone 3 stays low; the coral one changes tone outright.
When tones meet
One real word for each of the 20 two-syllable patterns. Only 3 + 3 changes the tone you say: 水果 shuǐguǒ is said shuíguǒ. Before tones 1, 2, 4 and the light tone, tone 3 stays low.
One word for each tone pattern (first tone + second tone)
Tones 1 + 1
- Word
- 今天 today
- Written
- jīntiān
- Said
- jīntiān
- What happens
- –
Tones 1 + 2
- Word
- 欢迎 welcome
- Written
- huānyíng
- Said
- huānyíng
- What happens
- –
Tones 1 + 3
- Word
- 开始 to begin
- Written
- kāishǐ
- Said
- kāishǐ
- What happens
- –
Tones 1 + 4
- Word
- 工作 work
- Written
- gōngzuò
- Said
- gōngzuò
- What happens
- –
Tones 1 + light
- Word
- 妈妈 mom
- Written
- māma
- Said
- māma
- What happens
- Light second syllable
Tones 2 + 1
- Word
- 明天 tomorrow
- Written
- míngtiān
- Said
- míngtiān
- What happens
- –
Tones 2 + 2
- Word
- 学习 to study
- Written
- xuéxí
- Said
- xuéxí
- What happens
- –
Tones 2 + 3
- Word
- 牛奶 milk
- Written
- niúnǎi
- Said
- niúnǎi
- What happens
- –
Tones 2 + 4
- Word
- 文化 culture
- Written
- wénhuà
- Said
- wénhuà
- What happens
- –
Tones 2 + light
- Word
- 朋友 friend
- Written
- péngyou
- Said
- péngyou
- What happens
- Light second syllable
Tones 3 + 1
- Word
- 北京 Beijing
- Written
- Běijīng
- Said
- Běijīng
- What happens
- Tone 3 stays low (no rise)
Tones 3 + 2
- Word
- 旅行 to travel
- Written
- lǚxíng
- Said
- lǚxíng
- What happens
- Tone 3 stays low (no rise)
Tones 3 + 3
- Word
- 水果 fruit
- Written
- shuǐguǒ
- Said
- shuíguǒ
- What happens
- The first syllable rises: third tone before third tone
Tones 3 + 4
- Word
- 努力 to work hard
- Written
- nǔlì
- Said
- nǔlì
- What happens
- Tone 3 stays low (no rise)
Tones 3 + light
- Word
- 我们 we
- Written
- wǒmen
- Said
- wǒmen
- What happens
- Tone 3 stays low (no rise)
Tones 4 + 1
- Word
- 放心 to relax
- Written
- fàngxīn
- Said
- fàngxīn
- What happens
- –
Tones 4 + 2
- Word
- 问题 question
- Written
- wèntí
- Said
- wèntí
- What happens
- –
Tones 4 + 3
- Word
- 电影 movie
- Written
- diànyǐng
- Said
- diànyǐng
- What happens
- –
Tones 4 + 4
- Word
- 再见 goodbye
- Written
- zàijiàn
- Said
- zàijiàn
- What happens
- –
Tones 4 + light
- Word
- 谢谢 thank you
- Written
- xièxie
- Said
- xièxie
- What happens
- Light second syllable
| Word | Where | Dictionary | Said |
|---|---|---|---|
| 一天 one day | before tone 1 | yī tiān | yì tiān |
| 一年 one year | before tone 2 | yī nián | yì nián |
| 一起 together | before tone 3 | yī qǐ | yì qǐ |
| 一样 the same | before tone 4 | yī yàng | yí yàng |
| 第一 first | counting, ordinals | dì yī | dì yī |
| 不吃 not eat | before tone 1 | bù chī | bù chī |
| 不忙 not busy | before tone 2 | bù máng | bù máng |
| 不好 not good | before tone 3 | bù hǎo | bù hǎo |
| 不是 is not | before tone 4 | bù shì | bú shì |
| 是不是 is it or not? | between repeated words | shì bù shì | shì bu shì |
Spoken forms follow SayMei's tone sandhi checker; dictionaries write the original tones. 一 is yí before a fourth tone and yì before tones 1–3, but stays yī in counting, dates and ordinals. 不 is bú before a fourth tone and light in A-not-A questions. Practise all 20 patterns with 400 words in the tone-pair trainer.
To drill every pattern with 20 words each, use the tone-pair trainer.
On paper, the free Tone-Pair Listening and Speaking Lab covers all 16 tone combinations plus light endings, in four 15-minute sessions.
Is the third tone low or dipping?
Usually low. It dips and rises only on its own or at the end of a phrase. Before tones 1, 2 and 4 it stays low, and before another third tone it rises like tone 2.
Phoneticians call the low version the "half third". Its pitch is about 21 instead of 214, as PolyU and Chen and colleagues describe it. We measured it in the lab's recordings:
tasty, 好吃
- What the first syllable does
- Falls about 5 semitones and never rises
- Pitch
- 196 to 143 Hz
hello, 你好
- What the first syllable does
- Climbs about 6 semitones, because the next syllable is also tone 3
- Pitch
- 211 to 303 Hz
Source: SayMei's computer voice, measured with the Praat pitch tracker. Both words play in the lab above.
Native speakers treat the low version as the normal case. They apply the low change to made-up words more reliably than the rising one. Learners lag behind: second-language speakers in the US used the low third tone before other syllables less often than heritage speakers did.
Part of the problem may be how the tone is taught. A thesis survey found that 14 of 15 beginner textbooks described tone 3 as falling and rising.
Researchers have argued that teaching the dip first may feed the tone 2 and tone 3 confusion, because tone 2 can dip slightly too.
What I'd do: aim for the bottom of your normal speaking range and stay there. Add the rise only at the end of a phrase. To see which third tones rise in a phrase you're learning, paste it into the tone sandhi checker.
Tone 2 and tone 3 are the pair to watch
Both can start low and rise. What separates them is when the pitch turns upward and how far it drops first, and even advanced learners find tone 2 the hardest to identify.
Listeners tell the two apart mainly by that turning point and that first drop. Tone 2 turns up early, after a small dip. Tone 3 sinks further and turns late, if at all. The advanced learners in the study above were near-native on every isolated tone except tone 2.
Even computer voices blur this line. We screened 16 one-syllable tone 2 recordings from SayMei's own audio library for the lab. Seven dipped too deep or turned up too late to count as clear examples, so the ear test leaves them out, as Methods explains.
Which tone is "hardest" depends on the task, so don't trust a single ranking:
- Listening: in a 2024 training study, tone 1 was easiest to identify and tone 3 hardest. Exaggerating the pitch range helped tones 2 and 3 but hurt tones 1 and 4.
- Speaking: American learners after four months of study made errors on 55.6% of tone 4 syllables but only 8.9% of tone 2 syllables, in one older study. Another, from the same year, found errors spread evenly.
The consistent trouble spot is the tone 2 and tone 3 pair. Two habits I'd build:
- Listen for the turn. In the lab's 2 vs 3 round, ask where the pitch bottoms out: near the start for tone 2, or past the middle for tone 3.
- Use more of your range when you speak. Native Chinese speakers use a pitch range about 1.5 times wider than English speakers do in English.
What is the neutral tone, and how do you say it?
A neutral-tone syllable is short and light, with no shape of its own. Its pitch follows the syllable before it: a little higher after a third tone, lowest after a fourth.
The neutral tone has a weak middle target that the previous tone overpowers. It's also about half as long as a full syllable, as Lin's review reports from earlier measurements.
Length is where learners go wrong. In one study, learners' neutral tones differed from native speakers' mainly because they were too long, not because of their pitch, and they got the lowest goodness ratings.
Sources agree that the light syllable sits low after tone 4, as in thank you, 谢谢, and relatively high after tone 3, as in we, 我们. They disagree about its exact height after tone 1, so our lab draws that one as a range.
The light tone can change meaning. Thing is 东西, but 东西 means east and west, as CC-CEDICT lists them.
To practise, make the second syllable quick. Say mom, 妈妈, then shorten the second "ma" until it nearly disappears.
Three tone changes to learn on day one
I'd learn three changes from the start:
- Two third tones in a row: hello, 你好, is said ní hǎo.
- One, 一, changes before a fourth tone.
- Not, 不, changes before a fourth tone.
Dictionaries keep writing the original tones, so learn the spoken forms by ear. Here's how they sound:
fruit, 水果
- Said
- shuíguǒ
- Why
- Third tone before a third tone rises
the same, 一样
- Said
- yíyàng
- Why
- 一 before a fourth tone becomes yí
together, 一起
- Said
- yìqǐ
- Why
- 一 before tones 1, 2 or 3 becomes yì
first, 第一
- Said
- dì yī
- Why
- Counting and ordinals keep yī
is not, 不是
- Said
- bú shì
- Why
- 不 before a fourth tone becomes bú
is it or not?, 是不是
- Said
- shì bu shì
- Why
- 不 between repeated words goes light
Source: these match the rules in SayMei's tone sandhi checker.
With three or more third tones in a row, word grouping decides which ones rise. I'm fine, 我很好, can be said wǒ hén hǎo or wó hén hǎo. I'd leave that to the checker, not to memory.
Knowing the rule is the easy part. American learners with about three years of study applied the third-tone rules even to made-up words, but with less accurate pitch than native speakers.
For printable drills with dictionary and spoken tones side by side, get the free workbook Tone Changes in Everyday Mandarin.
Pitch graphs help; a teacher still hears more
Pitch graphs and teachers both help, in different ways. Seeing your own pitch curve improved tone accuracy in several studies. Automatic checkers judge single syllables imperfectly, and a teacher still catches errors in context.
We found four visual-feedback studies. Each is small, but they point the same way:
- What learners got
- Their own pitch curves next to native ones, 20–25 minutes a week
- Who and how long
- 35 learners, 9 weeks
- Result
- Tones improved; two-thirds found the curves helpful
- What learners got
- Pitch displays plus corrected audio in their own voice
- Who and how long
- 44 beginners, 4 weeks
- Result
- Improved more than a group given feedback by the researcher
- What learners got
- Pitch contours with pinyin
- Who and how long
- First-year US learners
- Result
- Beat numbers with pinyin, and contours alone
- What learners got
- Contours, numbers or colours
- Who and how long
- 303 English-speaking novices, 3 weeks online
- Result
- Contours and numbers slightly beat colours; combining cues added nothing
Source: the four studies, linked.
How you're corrected matters too. In a 14-week one-to-one course, beginners whose teacher simply repeated the word back correctly, a "recast", improved their spoken tones more than those given explicit correction. The study, by Bryfonski and Ma, had 41 learners.
What an automatic checker can hear is narrower than it looks. SayMei's free Tone Checker compares the pitch shape of one word with the model. In our own test, it named the intended tone 84% of the time:
Source: SayMei's internal evaluation file, 6 October 2026: 4,344 labelled clips from seven computer voices, from ElevenLabs and Gemini text-to-speech, plus pitch-shifted, flattened, slowed, sped-up and noisy copies.
A real third tone can dip or stay low, and that confuses the checker. One caveat matters more: every test voice was computer-generated, including the deeper, flatter and noisier altered copies. We haven't measured it on human learners yet, and we'll add those results when we have them. It also doesn't judge sentences, consonants or vowels.
My rule: apps for repetition, a human for diagnosis. If you have one session with a teacher, ask them to listen to the same ten words you practise every day.
Shadowing and gestures can help
One small study says shadowing works for tones. Shadowing means speaking along with a recording, a split second behind it.
In a study of 14 beginners, four weeks of shadowing improved tone accuracy in free speech, whether the recordings were authentic videos or textbook audio. That's a small sample, so treat it as promising rather than proven.
Hand movements help some people. In a study of 106 participants, watching and making pitch gestures, like a hand that rises for tone 2, improved tone identification and word learning in a short session.
To build up to sentences, our Tone-Pair Lab PDF has "ladders" that grow one tone pair into a phrase and then a full sentence, with free audio. For a timed speaking drill you can do alone, see How to speak Chinese.
Your ear improves in weeks; your words take years
Your ear can improve within days or weeks. Using tones automatically in real words takes years. Here's how fast each kind of progress came in the studies we read:
Four days, an hour a day
- What changed
- Adults new to tonal languages improved
Two weeks
- What changed
- Training produced lasting gains
- Study
- Wang et al., 1999
One month of classes
- What changed
- Classroom learners improved; advanced ones matched native speakers on a discrimination task
About 11 years
- What changed
- Learners rejected only 35% of tone-only fake words, against 91% for native speakers; 1 of 16 scored in the native range
- Study
- Pelzl et al., 2019
Source: the four studies, linked.
The slow part is recognising words. The gap between advanced learners and native listeners persisted even in near-ideal listening conditions.
What no study shows. "Most learners master tones in two to three months" appears in blog posts and chatbot answers, with no source. I couldn't find a study that supports it, and the advanced-learner results point the other way. Expect clear progress in weeks, and a long tail of practice in real conversations.
You're almost certainly not tone deaf
True congenital amusia, the clinical form of tone deafness, is rare, and even people who have it learned new word tones in a training study.
The old figure of 4% came from a single 1980 estimate.
A test of 20,000 people put it at about 1.5%. Amusia does make tones harder, though:
- Among 22 Mandarin speakers with amusia, nearly half had trouble telling tones apart. The six with marked difficulty still produced tones normally.
- In a training study, 21 people with amusia were less accurate than 23 typical listeners, but they still learned new word tones.
Musical ability helps, but you don't need it. Pitch perception and musical experience predicted who learned pitch-based words best, with large individual differences.
Musical ability also went with tone accuracy in the listening-or-speaking study. And pitch acuity predicted tone identification in Portuguese-speaking learners.
If you struggle, it usually means more listening practice, not a missing ability.
A six-week plan for adult beginners
Ten to fifteen minutes a day is enough to start. That's our suggestion, not a study finding. Each week adds one skill, and every step uses one of our free tools:
Week 1
- Focus
- Hear single tones
- Do this
- Two rounds of the lab's 2 vs 3 ear test a day; tap through a few pinyin syllables. For consonants and vowels, use our pronunciation guide.
- Free tool
- Tone lab, pinyin chart, pronunciation guide
- Why
- Listening training works fast (Wang et al., 1999)
Week 2
- Focus
- Hear tone pairs
- Do this
- The 1-minute ear quiz on all 20 patterns; a few rounds of the game
- Free tool
- Tone-pair trainer, Panda Tone Toss
- Why
- Pairs are harder and closer to real speech (Chang & Bowles, 2015)
Week 3
- Focus
- Say words and check them
- Do this
- Say each word from the written prompt before you play the audio, then compare
- Free tool
- Tone Checker, Tone-Pair Lab PDF
- Why
- Speaking improves mainly through speaking (Li & DeKeyser, 2017); seeing your pitch helps (Chun et al., 2015)
Week 4
- Focus
- Low third tone and the three changes
- Do this
- Keep tone 3 low; drill 你好, 一 and 不
- Free tool
- Tone sandhi checker, Tone Changes PDF
- Why
- Natives use the low third most (PolyU)
Week 5
- Focus
- Light tone and shadowing
- Do this
- Shorten light syllables; shadow two short sentences a day
- Free tool
- Tone-Pair Lab PDF ladders
- Why
- Learners' light tones run long (Chang & Yao, 2019); shadowing helped in a small study (Lu & Su, 2024)
Week 6
- Focus
- Speak with feedback
- Do this
- Hold a short conversation; ask for corrections
- Free tool
- Free lesson with Mei Lin, or a human teacher
- Why
- Tones fail in real word use (Pelzl et al., 2019); recasts helped (Bryfonski & Ma, 2019)
Keep a fixed list of ten words, and re-test yourself once a week, for listening and speaking separately. A fixed list makes the scores comparable.
Six tone myths the evidence doesn't support
Each of these shows up in guides and chatbot answers. Here's what the studies actually found:
"Most learners master tones in two to three months."
- What the evidence says
- No study we found supports it. Learners with about 11 years still accepted most tone-only fake words.
- Source
- Pelzl et al., 2019
"The third tone always dips and rises."
- What the evidence says
- In connected speech it is usually low; the dip appears mainly alone or phrase-final.
- Source
- PolyU; Zhang & Lai, 2010
"The third tone has to be creaky."
- What the evidence says
- Creaky voice goes with low pitch in general, including other tones, and fades when overall pitch rises.
- Source
- Kuang, 2017
"Hearing many voices is always better."
- What the evidence says
- It helps with new voices and retention, but the immediate edge vanished after bias correction, and many voices held back weaker perceivers in one study.
"Colour-coding tones works best."
- What the evidence says
- Pitch contours and tone numbers did slightly better than colours with 303 novices.
"Context will save you."
- What the evidence says
- Flat-pitch Mandarin fell to 60% intelligibility in noise, against 80% for natural speech.
- Source
- Patel, Xu & Wang, 2010
The pattern is the same in each row: the confident version is simpler than the evidence.
Where SayMei fits, and where it doesn't
Our free tools cover listening and single-word checking, and Mei Lin gives you spoken practice. None of it replaces a teacher's ear for persistent errors.
Here's our free tool for each step, with no sign-up:
| Step | Free SayMei tool |
|---|---|
| Hear every tone of every syllable | Pinyin chart with audio: tap any of 403 syllables |
| Train your ear on tone pairs | Tone-pair trainer: all 20 two-syllable patterns, 400 words and a 1-minute ear quiz |
| Check your own voice | Tone Checker: say one word and see your pitch line next to the model's |
| Learn the tone changes | Tone sandhi checker: type a phrase and see which tones change |
| Play | Panda Tone Toss, a tone listening game |
| Mandarin Pronunciation Map for English Speakers (30 pages), Tone-Pair Listening and Speaking Lab (33 pages), Tone Changes in Everyday Mandarin (24 pages), Pinyin chart and spelling guide (22 pages); all on the downloads page | |
| Listen for fun | SayMei's HSK 1 songs, free with no account |
| Speak | A free 10-minute lesson with Mei Lin |
Where it fits
- The free first lesson puts your tones under conversational load, which is where the research says they break. It runs up to 10 minutes, with no account or card.
- After that, SayMei Premium is $5.99 a month for 200 minutes with Mei Lin. It starts with a 7-day trial with 100 minutes that needs a card; see our pricing page.
- The HSK 1 songs help with words and listening. They aren't tone models; see the FAQ.
Where it falls short
- No human ear. A trained teacher hears register errors, too high or too low, and mistakes in context that SayMei's tools miss.
- The Tone Checker is narrow. It checks one word at a time, tones 1 to 4 only. Its accuracy comes from computer voices, and it's weakest on tone 3. Its earlier version, replaced in October 2026, recognised our own reference recordings on only 50 of 96 cards.
- Mei Lin's ability to catch tone errors hasn't been measured. In the first lesson she's instructed to correct by repeating your sentence back the right way, one fix at a time. That's the style that worked in the recast study above, but that study used human teachers. Treat her as practice, not diagnosis.
- The audio is synthetic and uses few voices. The lab, chart and trainer use SayMei's text-to-speech. Real speakers vary more.
Skip SayMei for tones if you need a teacher to diagnose one persistent error, or detailed feedback on whole sentences. Book a human tutor for that.
What to do next
- Test your ear today. Run the 2 vs 3 round in the lab, then the tone-pair quiz.
- Say every word out loud. Say it from the written prompt first, then check it against the audio.
- Keep the third tone low. Add the rise only at the end of a phrase.
- Learn the three day-one changes. Two third tones in a row, one, 一, and not, 不.
- Get a human ear on your list. Ask a teacher to check the same words you practise every day.
FAQ
Do tones matter if people can guess from context?
Context rescues some errors, not all. In a lab test, native listeners understood flattened-pitch Mandarin about 94% of the time in quiet, but only 60% in noisy babble, against 80% for natural speech. Tones also carry at least as much information as vowels, and a wrong tone can point to a different word.
Should I learn tones before vocabulary, or with each word?
With each word, from the first week. Learners who name tones well can still miss them inside words. In one study, English speakers with about 11 years of Mandarin rejected wrong-tone words only 35% of the time, against 91% for native listeners. Store the tone as part of the word, and say it aloud whenever you review.
Is there an app that tells me if my tones are correct?
Yes, for single words, with limits. Pitch checkers compare the shape of your voice with the target tone. Our free Tone Checker named the intended tone in 84% of 4,344 test clips from seven computer-generated voices. On third tones it managed 66%, because a third tone can dip or stay low. It doesn't judge sentences, consonants or vowels.
Why can I copy a tone right after hearing it, but not say it on my own?
Because imitation and recall are different skills. Echoing uses the sound still in your ear; speaking on your own means pulling the tone from memory. In a training study with 38 English speakers, people who practised only listening did much worse when tested on speaking, and the reverse. Say words from the written prompt first, then check against the audio.
Is it too late to learn tones as an adult?
No. Adults with no Mandarin told tones apart better than chance and improved after one month of classes, and advanced classroom learners matched native speakers on the same task. In another study, adults new to tonal languages improved within four one-hour training days. What takes years is using tones automatically while you talk.
How do I know my tones are improving?
Measure listening and speaking separately, once a week. For listening, score yourself on words you haven't memorised. Guessing gets 25% in a four-way choice and 50% in the lab's two-way rounds, so a score clearly above chance that keeps rising is real progress.
For speaking, record the same ten words each week and compare them with the model, or ask a teacher to mark them. A fixed list keeps the scores comparable.
Is the third tone supposed to be creaky?
Often, but creak isn't what makes it a third tone; low pitch is. Acoustic research shows creaky voice can come with any low pitch in Mandarin, including other tones. The third tone also gets less creaky when a speaker raises their overall pitch. Aim for the bottom of your normal range; if creak appears there, fine, but don't force it.
Do Chinese songs keep the tones?
Often not. A survey of Mandarin songs found that the melody often fails to preserve each word's tone. Some genres follow tones more closely: in two excerpts from Chinese musicals, the melody followed the tones over 65% of the time. Use songs, like our free HSK 1 songs, for words and listening, and learn each word's tones from speech.
Sources
Checked 6–7 October 2026. "Abstract" means we read the abstract or a summary, not the full paper.
How tones work
- PolyU Department of Chinese and Bilingual Studies, "Basic tones" (Chao tone values, half third). https://www.polyu.edu.hk/bepth/introduction-to-phonetics/tones/basic-tones/
- Chen, He, Wayland, Yang, Li & Yuen (2019). Speech Communication 115:67–77 (full text). https://doi.org/10.1016/j.specom.2019.10.008
- Xu (1997). Journal of Phonetics 25:61–83 (author copy). https://www.homepages.ucl.ac.uk/~uclyyix/yispapers/Xu_JP97.pdf
- Moore & Jongman (1997). JASA 102(3):1864–77 (abstract). https://pubmed.ncbi.nlm.nih.gov/9301064/
- Zhang & Lai (2010). Phonology 27(1):153–201 (abstract): the 214-to-21 change applied more reliably than 214-to-35. https://doi.org/10.1017/s0952675710000060
- Chen & Xu (2006). Phonetica 63(1):47–75 (abstract). https://pubmed.ncbi.nlm.nih.gov/16514275/
- Lin, "The register tonal feature and the neutral tone in Mandarin", UVic Working Papers in Linguistics (full text; cites Chao 1968, Yip 1980, Cheng 1973). https://journals.uvic.ca/index.php/WPLC/article/view/5090/1999
- Kuang (2017). JASA 142(3):1693–1706 (abstract). https://hmongstudies.library.wisc.edu/catalog/HmongStudies1589
- Surendran & Levow (2004). Speech Prosody (abstract). https://doi.org/10.21437/speechprosody.2004-23
- Patel, Xu & Wang (2010). Speech Prosody (abstract). https://doi.org/10.21437/speechprosody.2010-238
- CC-CEDICT dictionary (MDBG release of 29 November 2025), CC BY-SA 4.0. https://www.mdbg.net/chinese/dictionary
How adults learn tones
- Wang, Spence, Jongman & Sereno (1999). JASA 106(6):3649–58 (abstract). https://pubmed.ncbi.nlm.nih.gov/10615703/
- Wang, Jongman & Sereno (2003). JASA 113:1033–43 (abstract). https://kuscholarworks.ku.edu/signposting/describedby/e4e66ea1-40f1-42b1-8027-807cbfcca7f8
- Wang, Jongman & Sereno, "L2 acquisition and processing of Mandarin tone" (chapter manuscript; summarises Chen 1974, Shen 1989, Miracle 1989). https://kuppl.ku.edu/sites/kuppl/files/documents/publications/wang_etal_revision.pdf
- Pelzl, Lau, Guo & DeKeyser (2019). Studies in Second Language Acquisition 41:59–86 (full text). https://doi.org/10.1017/S0272263117000444
- Pelzl, Lau, Guo & DeKeyser (2021). Studies in Second Language Acquisition (abstract). https://doi.org/10.1017/s027226312000039x
- Chang & Bowles (2015). JASA 138(6):3703–16 (abstract). https://pubmed.ncbi.nlm.nih.gov/26723326/
- Sakai & Moorman (2018). Applied Psycholinguistics 39(1):187–224 (abstract): d = 0.92 for listening, 0.54 for speaking. https://doi.org/10.1017/s0142716417000418
- Li & DeKeyser (2017). Studies in Second Language Acquisition 39(4):593–620 (abstract). https://doi.org/10.1017/s0272263116000358
- Zhang, Cheng & Zhang (2021). JSLHR 64(12):4802–25 (abstract; a meta-analysis of non-native speech-sound training: 18 studies, 549 people; g = −0.04 for the immediate advantage of many voices after correction, 0.72 for new voices, 1.09 for retention). https://doi.org/10.1044/2021_JSLHR-21-00181
- Perrachione, Lee, Ha & Wong (2011). JASA 130(1):461–72 (abstract). https://pubmed.ncbi.nlm.nih.gov/21786912/
- Cao, Pavlik & Bidelman (2024). Frontiers in Psychology 15:1403816 (abstract). https://pubmed.ncbi.nlm.nih.gov/39233888/
- Chang & Yao (2016). Heritage Language Journal 13(2):134–60 (abstract). https://doi.org/10.46538/hlj.13.2.4
- Chang & Yao (2019). ICPhS proceedings (abstract). https://ira.lib.polyu.edu.hk/handle/10397/93049
- Linge (2011). Lund University thesis on teaching the third tone (abstract). https://www.hackingchinese.com/media/teaching_the_third_tone_in_standard_chinese.pdf
- Li, Jeong & Astikainen (2026). Neuropsychologia 231:109543 (abstract). https://pubmed.ncbi.nlm.nih.gov/42425404/
- Wang, Potter & Saffran (2020). Language Learning and Development 16(3):231–43 (abstract). https://pubmed.ncbi.nlm.nih.gov/33716583/
Feedback, practice and correction
- Chun, Jiang, Meyr & Yang (2015). Journal of Second Language Pronunciation 1(1):86–114 (abstract). https://doi.org/10.1075/jslp.1.1.04chu
- Chen (2024). Computer Assisted Language Learning 37(3):363–88 (abstract). https://doi.org/10.1080/09588221.2022.2037652
- Liu, Wang, Perfetti, Brubaker, Wu & MacWhinney (2011). Language Learning 61(4):1119–41 (abstract). https://doi.org/10.1111/j.1467-9922.2011.00673.x
- Godfroid, Lin & Ryu (2017). Language Learning 67:819–57 (abstract). https://doi.org/10.1111/lang.12246
- Bryfonski & Ma (2019). Studies in Second Language Acquisition (abstract): d = .75. https://doi.org/10.1017/s0272263119000317
- Lu & Su (2024). Journal of Second Language Pronunciation 10(1):59–84 (abstract). https://doi.org/10.1075/jslp.22033.lu
- Baills, Suárez-González, González-Fuente & Prieto (2019). Studies in Second Language Acquisition 41(1):33–58 (abstract). https://doi.org/10.1017/s0272263118000074
Tone deafness and musical ability
- Peretz & Vuvan (2017). European Journal of Human Genetics 25(5):625–30 (abstract). https://pubmed.ncbi.nlm.nih.gov/28224991/
- Nan, Sun & Peretz (2010). Brain 133(9):2635–42 (abstract). https://pubmed.ncbi.nlm.nih.gov/20685803/
- Zhu, Chen, Chen, Zhang, Shao & Wiener (2023). JSLHR 66(7):2461–77 (abstract). https://pubmed.ncbi.nlm.nih.gov/37267445/
- Wong & Perrachione (2007). Applied Psycholinguistics 28(4):565–85 (abstract). https://doi.org/10.1017/s0142716407070312
- Zhou & Veríssimo (2025). Bilingualism: Language and Cognition (abstract). https://doi.org/10.1017/s1366728925100114
Songs
- Wee (2007). "Unraveling the relation between Mandarin tones and musical melody." Journal of Chinese Linguistics 35(1):128–44 (abstract). https://scholars.hkbu.edu.hk/en/publications/unraveling-the-relation-between-mandarin-tones-and-musical-melody-10
- Zhang & Kong (2026). "Tone-tune correspondence in Chinese musicals." Speech Prosody (abstract). https://doi.org/10.21437/speechprosody.2026-172
SayMei (first-party)
- Tone Checker page and SayMei's internal evaluation file (6 October 2026). https://www.saymei.app/tools/chinese-pronunciation-diagnosis
- Tone sandhi checker (rules engine). https://www.saymei.app/tools/tone-sandhi-checker
- Free first lesson and pricing, checked 6 October 2026. https://www.saymei.app/live/first-lesson · https://www.saymei.app/pricing
Methods
This guide was checked against the linked sources and the CC-CEDICT dictionary. If you spot a mistake, tell us: we fix errors.
- Sources. Every number links to its source. We read the full text of the Pelzl 2019, Chen 2019 and Xu 1997 papers, and the abstracts or summaries of the rest; the list above says which.
- Chinese examples. Every example was checked against CC-CEDICT, and every spoken form against SayMei's tone-change rules, the same engine behind the tone sandhi checker.
- Audio and the tone lab. Every sound is a recording from SayMei's own text-to-speech library, in ElevenLabs computer voices: single syllables from the pinyin chart's audio, words from the tone-pair workbook and the tone sandhi checker. We measured the lab's dashed pitch lines from those recordings with Praat. For the ear test, we kept only single syllables whose measured pitch clearly matched their tone, using simple rules for each tone. 9 of 16 candidate tone 2 recordings passed, and 44 of 46 for tones 1, 3 and 4. This is a rough screen for choosing clear examples, not a scientific classifier.
- Tone Checker figures. These come from SayMei's internal evaluation of 6 October 2026, on computer voices only. No learner data was used in this guide.
- What we didn't do. We did not test other apps hands-on, and this guide doesn't rank them.
