SayMei

How to speak Chinese: what the research says, and how to practise

Speaking is a separate skill from understanding. What research says works for speaking Mandarin, a 4/3/2 drill to try now and a routine for solo learners.

You can follow a slow podcast in Chinese, but when someone asks you a question, nothing comes out. That gap is normal, and it has a clear explanation.

If you're an English-speaking adult, from beginner to about HSK 4, and you understand more Mandarin than you can say, this guide is for you. We read the research on learning to speak a second language, counted what happens in SayMei's lessons, and turned both into a drill and a daily routine.

Key takeaways

  • Speaking is its own skill. Speaking practice builds speaking, and listening practice builds listening, so you have to talk to get better at talking.
  • Feedback works. Across 33 studies, correcting learners' mistakes had a medium effect that lasted.
  • Speaking turns are scarce in group classes. One-to-one formats give you many more, and solo drills give you unlimited turns, with nobody to answer.
  • Shadowing and timed retelling build fluency alone. Neither needs a partner, and neither can tell you whether your Chinese is right.
  • Tones live in sentences. They change shape next to each other, so practise them in phrases, not just as single syllables.
  • Lower the stakes first. Anxiety goes with slower progress, so start alone and work up to real people.

Why can you understand Chinese but not speak it?

Because speaking is a different skill from understanding, and it's the one that has to run in real time.

Psycholinguists describe speaking as a production line. Willem Levelt's model, the standard account since 1989, breaks it into five stages:

  1. Decide what you want to say.
  2. Find the words and put them into grammar.
  3. Build the sounds, including Mandarin's tones.
  4. Move your mouth.
  5. Listen to yourself and fix slips.

Fluent speakers run this line at two or three words a second. In your first language, finding the words and building the sounds take no effort. In Chinese they're slow and conscious, so they eat the attention you need for the message.

Listening runs on different processes. The words and grammar arrive from the speaker, and your job is to match them to what you know.

Your vocabulary shows the same split. You can recognise more words than you can produce, and the gap widens for less common words. So a word you "know" from flashcards may still be out of reach mid-sentence.

The finding I'd put first is that practice transfers narrowly. In an experiment with 82 learners, listening practice helped people understand a grammar point, and speaking practice helped them produce it.

A later meta-analysis of 35 comparisons found that both kinds of teaching worked. For producing language, though, production practice did better on delayed tests.

Linguist Merrill Swain gave this idea a name in 1985: the output hypothesis. Speaking lets you notice gaps, test your guesses about how the language works and think about it. In a study with Sharon Lapkin, producing language made learners notice problems and rework their sentences.

None of this means listening is wasted. It fills the store of words and sounds that speaking draws on. It just doesn't train you to pull them out.

If you grew up hearing Chinese at home and understand far more than you can say, this gap is your whole story. Our guide for heritage speakers picks it up from there.

To speak better, you have to speak.

Four ingredients make speaking improve

In the research we read, four things have the strongest evidence: producing language, getting feedback, repeating the same task, and doing it often over weeks.

Computers can supply some of this. A 2025 meta-analysis of 16 studies of dialogue-based systems found a moderate improvement in learners' speaking.

One caveat: it pooled different kinds of systems and learners. Treat the size of the effect as a rough guide, not a promise.

A routine worth keeping has all four: you speak, someone or something answers, you say it again, and you keep going for weeks.

Group classes leave little time to talk

Speaking turns are scarce in group classes, plentiful one-to-one and unlimited alone, though alone nobody answers.

Here's the arithmetic we'd start from. Take a 50-minute class of 30 students, in which the teacher talks for half the time.

That leaves 25 minutes. Split evenly, each student gets about 50 seconds, before any time goes to reading, silence or admin.

Michael Long and Patricia Porter made this argument in 1985, to push for pair and group work. It still holds for big lecture-style classes.

One-to-one formats change the arithmetic. We counted how often learners spoke in our free first lesson with Mei Lin:

Most learners took 20 to 29 speaking turns in 10 minutes

Learner speaking turns per 10 minutes in SayMei's free first lesson: 1,790 learners, median about 26

Speaking turns per 10 minutesLearners
Under 10219
10–19274
20–29711
30–39406
40 or more180

Source: SayMei product analytics, read-only, 15 July to 7 October 2026. One row per person: their first lesson, if it lasted at least 60 seconds. The middle half of learners took 18.8 to 32.2 turns per 10 minutes. A turn is one finished learner utterance, spoken or typed, and can be a single word. Internal and test traffic is excluded, no transcripts or recordings were used, and every published number covers at least 50 learners.

We also timed the turns. They came quickly, and almost all of them were spoken:

MilestoneMedian timeSessions
Fifth speaking turn100 seconds1,392
Tenth speaking turn190 seconds1,077

Source: SayMei lesson records, every first lesson in the same window. Each time counts only the sessions that got that far. About 91% of these turns were spoken rather than typed.

One limit: turns aren't minutes of speech, so they don't compare directly with the classroom estimate. These learners also chose to try SayMei, and this is one product's design, not a benchmark for all tutors.

The point isn't that an app beats a class. It's that most learners need to add speaking somewhere. One-to-one sessions, with a person or an AI tutor, give you many turns with a listener. Solo methods give you unlimited turns without one.

Does shadowing work for Mandarin?

Yes, for rhythm, fluency and tones in phrases. It won't teach you to build your own sentences.

Shadowing means speaking along with a recording, a split second behind the speaker. A systematic review of 44 studies found it can make learners easier to understand and improve their fluency and intonation.

Evidence for individual sounds was inconclusive, though. Many of the studies also tested learners only on material they'd rehearsed.

We found two studies that give a sense of a realistic dose:

  • Foote & McDonough, 2017

    Who
    16 learners
    How much
    8 weeks, 4+ times a week, 10+ minutes, recording themselves
    What changed
    Easier to understand and more fluent; accent unchanged
  • Lu & Su, 2024

    Who
    14 beginners learning Mandarin
    How much
    4 weeks
    What changed
    More accurate tones in spontaneous sentences

Source: the two studies, linked. In the first, learners shadowed short dialogues on iPods. In the second, textbook audio worked as well as authentic videos.

The Mandarin study is small, so it's encouraging rather than conclusive. Still, it's the closest evidence we have for tones.

Here's how I'd shadow, in three passes:

  1. Pick 20 to 40 seconds of audio you mostly understand, with a transcript. Our listening dialogues are free and come with pinyin.
  2. Listen twice, then read the transcript aloud along with the audio.
  3. Shadow without the text, record one take, and compare it with the original.

Try it with one sentence, "First I have coffee, then I go to work": 我先喝咖啡,然后去上班。 wǒ xiān hē kāfēi, ránhòu qù shàngbān

Try it: a 4/3/2 speaking drill for Mandarin

Tell the same short story three times, in four, three and then two minutes. It's built for fluency: fewer pauses and faster recall.

The activity comes from linguist Paul Nation. It suits solo learners well, because you need only a timer and a topic.

It also has a known weak spot. When the time shrinks, people mostly repeat their first version word for word, so they get faster but not more accurate. Here's what we found in the studies:

  • de Jong & Perfetti, 2011

    Learners
    24 ESL students, 4-, 3- and 2-minute talks in three sessions
    What they found
    Everyone got more fluent; only those who repeated the same topic kept the gain on later tests
  • Thai & Boers, 2016 and Boers, 2014

    Learners
    Learners of English
    What they found
    Shrinking time gave the biggest fluency gains but none in accuracy or complexity; constant time gave smaller fluency gains and some accuracy gains
  • Huang & Liu, 2022

    Learners
    28 learners talking to themselves for four weeks
    What they found
    More fluent, especially with rising time pressure

Source: the studies, linked. All of them tested learners of English.

The researchers' advice was to build in an early chance to fix your language. So the drill below adds a one-minute look-up after the first round, and offers a mode that keeps the time constant.

The 4/3/2 drill

Tell the same short story three times with less time each round, with a one-minute look-up after round 1. It measures time and long pauses, not whether your Chinese is right.

  1. Pick a topic and a level.
  2. Talk about it for the first round without stopping to fix mistakes.
  3. Take one minute to look up a word or phrase you reached for and didn't have.
  4. Tell the same story again in less time, using what you looked up.
  5. Tell it once more, fastest. If you can, record each round and listen back.
  6. Repeat the same topic tomorrow, then move to a new one.

Level

Mode

Your prompt

Say your name, where you're from, where you live, what you do and one thing you like.

Chunks you can use

  • 我叫…… wǒ jiào · My name is…
  • 现在住在…… xiàn zài zhù zài · now live in…
  • 我喜欢……,也喜欢…… wǒ xǐ huan yě xǐ huan · I like…, and I also like…

Three ways to time the rounds

  • Short 2 / 1.5 / 1 min

    Rounds
    2:00, 1:30, 1:00
    Why
    Short mode keeps the 4:3:2 ratio in less time for beginners. It's our adaptation and hasn't been tested in research.
  • Fluency 4 / 3 / 2 min

    Rounds
    4:00, 3:00, 2:00
    Why
    The research format: 4, 3, then 2 minutes. Best for speed and smoothness.
  • Accuracy 3 / 3 / 3 min

    Rounds
    3:00, 3:00, 3:00
    Why
    Same time every round. In one study this gave smaller fluency gains but some accuracy and complexity gains.

The drill shows model answers for each topic at two levels, Starter (about HSK 1–2) and Building (about HSK 3–4), and the prompt and chunks for each of its 5 topics. Pinyin shows dictionary tones, except 一 and 不, which show the tone you say; check any sentence in the tone sandhi checker. Recording is optional and stays on your device: nothing is uploaded.

One caveat: every study of this drill we found was with learners of English. We found no published study with Mandarin learners. The drill's Short mode, with rounds of two minutes, a minute and a half and one minute, is our own adaptation for beginners, not a tested format.

The drill also can't tell you whether your Chinese was right. It measures only time and, if you allow the microphone, long pauses, on your device.

Practise tones in whole sentences

Practise tones inside phrases and sentences. That's where they change, and where single-syllable drills transfer least.

Mandarin tones shift in context. You'll meet three changes early:

  • Two third tones in a row: the first one rises. Hello, 你好(nǐ hǎo), is said ní hǎo, and very good, 很好(hěn hǎo), is said hén hǎo.
  • "Not" before a fourth tone: 不(bù) rises, as in is not, 不是(bú shì).
  • "One" changes too: 一(yī) is said yì in together, 一起(yìqǐ), and yí in the same, 一样(yíyàng).

Learners pick these rules up. But in one study, English-speaking learners said the changed tones with less native-like pitch shapes. Our free tone sandhi checker shows the spoken tones for any sentence you paste in.

Pronunciation teaching works on average. A meta-analysis of 86 reports found a large effect, bigger when learners got feedback.

A later review of 77 studies adds a catch. The gains were clearest in careful, monitored speech, and less clear when judged in spontaneous speech.

Drilling the four tones on one syllable is a start: mom, 妈(mā); hemp, 麻(má); horse, 马(mǎ); scold, 骂(mà). Saying whole sentences at speed is the goal.

Listening practice helps your speaking here too. American learners who trained only on hearing tones afterwards produced tones that native listeners identified 18% more accurately.

Will people understand you if your tones slip? Often, in a quiet room. When researchers flattened all the pitch out of Mandarin sentences, listeners still understood about 94% of them in quiet.

In background noise, understanding fell to 60%, against 80% for natural speech. And some slips change the word: buy is 买(mǎi), and sell is 卖(mài).

For practice, our free Tone Checker records you and draws your pitch against the target shape. The tone-pair trainer drills the two-syllable combinations, and our guide to learning tones goes further.

Feedback helps, if your listener can hear tones

For tones, hearing the right version straight after your attempt works well. Software that can't really hear tones will let mistakes slide.

Researchers distinguish two kinds of correction. In a recast, the listener says your sentence back correctly. In a prompt, they signal a mistake and make you fix it.

In classrooms, prompts beat recasts. For Mandarin tones, one study points the other way.

In a 14-week one-to-one online course with 41 beginners, the group corrected with recasts improved its tone production more than the group corrected explicitly. It was a medium effect, and both students and teacher preferred recasts.

Our reading: let a listener model your tones, and ask them to make you retry your grammar.

Machines are uneven listeners. Speech-recognition practice improved pronunciation in a meta-analysis of 15 studies, most of all when the software gave explicit corrections.

Dictation-style feedback, where you simply see what the software heard, helped less. The effect was small for people practising alone.

General voice assistants are the weakest judges of pronunciation. OpenAI's own help page says ChatGPT Voice transcripts aren't verbatim, and Hacking Chinese put ChatGPT to the test in September 2026.

In separate chats, ChatGPT scored the same native teacher's recording anywhere from 6.5 to 8.5 out of ten. When the identical file was uploaded again, it invented an "improvement".

A transcript can catch a slip that turns buy, 买(mǎi), into sell, 卖(mài). But a correct transcript doesn't prove your tones were right. That's also why SayMei's scripted dialogue practice, which matches your syllables against the text, deliberately gives no tone score.

Here's how we do it at SayMei. In first lessons, Mei Lin is told to recast rather than say "wrong", and to fix at most one error per turn, so the conversation keeps moving.

The upside is flow, and a model to copy. The downside is that a recast is easy to miss, and the classroom evidence favours prompts for grammar. If you want to be made to try again, say so, whoever your teacher is.

Nervous? Lower the stakes first

Anxiety and progress are linked, so start where mistakes cost nothing and raise the stakes as you go.

Researchers have measured foreign language anxiety since a 1986 paper by Horwitz, Horwitz and Cope, now cited more than 4,500 times. Two meta-analyses have since put a number on the link:

Meta-analysisLearnersCorrelation
Teimouri, Goetze & Plonsky, 201919,933−.36
Zhang, 2019more than 10,000−.34

Source: the two meta-analyses, linked. The first measured anxiety against achievement, the second against performance. Both links are moderate and negative, and the second held steady across proficiency levels.

That's correlation, not proof that anxiety slows you down. But it fits what learners describe: if speaking feels awful, you do less of it.

Low-stakes practice can help. In a study of 232 secondary-school students, practising with speech-recognition websites lowered their speaking anxiety, compared with regular classes. That was English, not Mandarin, but the logic carries.

Here's the ladder I'd suggest, one that works with your nerves rather than against them:

  1. Alone: shadowing and the drill on this page. Nobody hears you.
  2. Recorded: listen back. Hearing yourself is uncomfortable for a week, then useful.
  3. A patient listener: an AI tutor or a tutor you choose, where a mistake costs nothing.
  4. Real people: a language exchange, a café order, a call with family.

Match your practice to your budget and nerves

No single method covers everything. Match the mix to your budget, your time and your nerves.

We compared seven ways to practise on cost and nerves. Prices change often, so treat them as a snapshot:

Solo methods are free and calm; nerves rise once real people join in

Each line is one method; a bar is a price range. Details for every method are in the cards below.

Nerves ↑

  • Medium to high

    • Language exchange $0
  • Medium

    • Online tutor ~28–38¢
  • Low

    • Speech recognition $0
    • Voice assistant Go ~2¢; Plus ~4¢
    • AI tutor ~2–6¢
  • Very low

    • Shadowing $0
    • Timed retelling $0

Cost per practice minute →

  • Shadowing

    What feedback you get
    None automatic; compare your recording
    Price
    $0
    Cost per practice minute*
    $0
    Nerves
    Very low
    Best for
    Rhythm, tones in phrases
  • Timed retelling (4/3/2), self-talk

    What feedback you get
    Self-review only
    Price
    $0
    Cost per practice minute*
    $0
    Nerves
    Very low
    Best for
    Fluency, recall
  • Reading aloud to speech recognition (phone dictation, SayMei dialogue practice)

    What feedback you get
    Did the transcript match? Not a tone verdict
    Price
    $0 (SayMei: 1 free session a day)
    Cost per practice minute*
    $0
    Nerves
    Low
    Best for
    Checking you're understood
  • General voice assistant (ChatGPT Voice, Gemini Live)

    What feedback you get
    Conversation; unreliable on pronunciation
    Price
    ChatGPT Free (limited) $0, Go $8, Plus $19.99; Gemini Live $0
    Cost per practice minute*
    Go ~2¢; Plus ~4¢
    Nerves
    Low
    Best for
    Lots of low-stakes turns
  • Dedicated AI tutor (SayMei, Speak, Praktika, Pingo AI, Talkpal, SuperChinese CHAO, Gliglish)

    What feedback you get
    Built-in corrections (vendor claims; untested here)
    Price
    SayMei $5.99 for 200 min (+$2.99 per 100); Praktika $9.99; Pingo $14.99; Speak $17.99; Talkpal $19.99; CHAO $24.99; Gliglish 10 min/day free
    Cost per practice minute*
    ~2–6¢
    Nerves
    Low
    Best for
    Structured daily speaking
  • Online tutor (italki, Preply)

    What feedback you get
    A trained ear plus targeted correction
    Price
    italki $20/hr (median lowest hourly price of 763 Chinese teachers); Preply $14 average class, $23/hr native average
    Cost per practice minute*
    ~28–38¢
    Nerves
    Medium
    Best for
    Diagnosing tone and grammar habits
  • Language exchange (HelloTalk, Tandem, meetups)

    What feedback you get
    Friendly but untrained
    Price
    $0 (HelloTalk VIP $12.99 optional)
    Cost per practice minute*
    $0
    Nerves
    Medium to high
    Best for
    Real conversation, culture

Sources: OpenAI, Google, SayMei pricing, Speak, Preply and italki; other app prices are from their US App Store listings. Prices checked 6 and 7 October 2026. *At 15 minutes a day for 30 days, or 450 minutes. Costs are per minute of session time, not of your own speech; we didn't measure talk share.

Here's how we'd choose, by what you need most:

  • Pick a general voice assistant if you want the most conversation for the least money. ChatGPT Go gives you up to three hours of voice a day, and Gemini Live is free. Expect to supply your own structure, and don't trust it on tones.
  • Pick a dedicated AI tutor if you want someone else to plan the speaking: a lesson path, questions at your level, corrections. SayMei is the cheapest paid plan in this group, but caps you at 200 minutes a month. Gliglish's free daily minutes need no account, and Speak's Mandarin course is new, from June 2026, running from beginner to B1.
  • Pick a tutor if you've practised for a while and want your habits diagnosed. Even one session a month adds the trained ear no app reliably has.
  • Pick a language exchange once you can keep a few minutes going and want real people. Partners can be unreliable, and conversations drift into English.
  • Skip SayMei if you want unlimited free conversation, a human diagnosis of your tones or offline practice in a native app. Use Gemini Live or ChatGPT for the first, book a tutor for the second; SayMei runs in the browser only.

For app-by-app detail, see our comparison of AI apps for learning Chinese.

A 15-minute daily routine for solo learners

Fifteen minutes a day, six days a week, built from the evidence above. No study tested this exact routine. It's our recommendation, and we built each part from the studies in this guide.

Here's what I'd do each day, at two stages:

  • 5 min

    Beginner (first months)
    Shadow one short dialogue line by line
    HSK 3–4
    Shadow 30 seconds of a podcast or story
  • 5–7 min

    Beginner (first months)
    4/3/2 drill, Short mode, with a Starter model answer
    HSK 3–4
    4/3/2 drill, full mode; same topic two days running, then a new one
  • 3–5 min

    Beginner (first months)
    Tone pairs in the tone-pair trainer, or a few turns with a conversation partner
    HSK 3–4
    One conversation with feedback (tutor, AI tutor or exchange)
  • Weekly

    Beginner (first months)
    One session with someone who answers back; re-record Monday's talk on Saturday and compare
    HSK 3–4
    Two conversations with feedback; keep a list of the words you looked up

The weekly line matters most. Solo practice builds fluency, but only a listener tells you what to fix. If you're preparing for the HSKK speaking exam, the same habits apply. Our HSKK guide times each paper and has a practice timer, and our HSK exam guide covers the HSK itself.

What to do next

  • Speak a little every day. Shadow one short recording, then run the drill on this page.
  • Repeat before you move on. Tell the same story twice, fix one gap between rounds, then change topic.
  • Practise tones in phrases. Check the spoken tones of any sentence before you shadow it.
  • Add a listener every week. A tutor, an AI tutor or an exchange partner, so someone can tell you what to fix.
  • Ask to be corrected. Tell your listener to make you retry, not just to say it back.

FAQ

How long does it take to learn to speak Chinese?

The US Foreign Service Institute estimates about 88 weeks, or 2,200 class hours, for US diplomats in full-time training to reach professional-level speaking and listening. They add 17 hours of self-study a week.

A basic conversation comes much sooner. SayMei doesn't publish time-to-level data yet; our guide to how long Mandarin takes breaks the estimate down by hours per week.

Can I learn to speak Chinese without a teacher or partner?

You can build fluency alone. Learners who recorded the same talk three times became more fluent, and so did learners who practised talking to themselves for four weeks. What solo practice lacks is feedback: speech-recognition practice had only a small effect for people practising alone. Add a listener, a person or an AI tutor, at least weekly.

Is shadowing a good way to learn Chinese?

Yes, as one part of practice. A review of 44 studies found shadowing makes you easier to understand and improves fluency and intonation, though not reliably individual sounds. In a small Mandarin study, 14 beginners improved their tone accuracy in spontaneous sentences after four weeks. It doesn't teach you to build your own sentences.

Is ChatGPT voice good for practising Chinese speaking?

It's good for lots of low-pressure conversation: ChatGPT Go allows up to three hours of voice a day for $8 a month. But it's a poor judge of pronunciation. OpenAI says voice transcripts aren't verbatim. In a September 2026 test, ChatGPT gave the same native speaker different scores in separate chats, then invented progress on an identical recording.

How many minutes a day should I practise speaking Chinese?

No study sets a minimum. In one shadowing study that worked, learners practised at least 10 minutes, four times a week, for eight weeks, while short bursts of speech-recognition practice did no better than none. We recommend 15 minutes a day, six days a week, with one session a week where someone answers back.

What's the cheapest way to practise speaking Chinese every day?

Shadowing, self-talk and timed retelling cost nothing. For conversation, Gemini Live is free with a Google account, and ChatGPT's free tier gives limited voice.

Among paid options at 15 minutes a day, ChatGPT Go works out at about 2¢ a minute and SayMei about 3¢. An average Preply class, at $14, is about 28¢ a minute if it runs 50 minutes. Prices checked 6 October 2026.

Should I wait until I know more before I start speaking?

No. Trying to speak is what shows you the gaps: Swain and Lapkin found that producing language made learners notice problems and rework their sentences. Across 35 comparisons, production-based teaching also did better for productive knowledge on delayed tests. Start with sentences you can almost say, like "My name is…", 我叫…… wǒ jiào.

Will people understand me if my tones are wrong?

Often, in a quiet conversation with context. When researchers removed all pitch from Mandarin sentences, listeners understood about 94% in quiet, but only 60% in noise, against 80% for natural speech. Some tone errors change the word entirely, like buy, 买(mǎi), and sell, 卖(mài), so tones are worth the work.

Sources

  1. Levelt, W. J. M. (1989). Speaking: From intention to articulation. MIT Press. https://doi.org/10.5860/choice.27-1947
  2. Levelt, Roelofs & Meyer (1999). Behavioral and Brain Sciences 22(1). https://doi.org/10.1017/S0140525X99001776
  3. Webb (2008). Studies in Second Language Acquisition 30(1). https://doi.org/10.1017/S0272263108080042
  4. DeKeyser & Sokalski (1996). Language Learning 46(4). https://doi.org/10.1111/j.1467-1770.1996.tb01354.x
  5. Shintani, Li & Ellis (2013). Language Learning 63(2). https://doi.org/10.1111/lang.12001
  6. Swain & Lapkin (1995). Applied Linguistics 16(3). https://doi.org/10.1093/applin/16.3.371
  7. Sales (2020). Interview with Merrill Swain. https://doi.org/10.15210/interfaces.v20i0.18775
  8. Li (2010). Language Learning 60(2). https://doi.org/10.1111/j.1467-9922.2010.00561.x
  9. Lyster & Saito (2010). SSLA 32(2): 15 classroom studies, 827 learners. https://doi.org/10.1017/S0272263109990520
  10. Bryfonski & Ma (2019). SSLA: d = .75. https://doi.org/10.1017/S0272263119000317
  11. de Jong & Perfetti (2011). Language Learning 61(2). https://doi.org/10.1111/j.1467-9922.2010.00620.x
  12. Lee, Jang & Plonsky (2015). Applied Linguistics 36(3). https://doi.org/10.1093/applin/amu040
  13. Ngo, Chen & Lai (2024). ReCALL 36(1). https://doi.org/10.1017/S0958344023000113
  14. Hou & Min (2026). ReCALL 38(1). https://doi.org/10.1017/S0958344025100268
  15. Long & Porter (1985). TESOL Quarterly 19(2). https://doi.org/10.2307/3586827
  16. Whitworth & Rose (2025). Research Synthesis in Applied Linguistics 1(2). https://doi.org/10.1080/29984475.2025.2546827
  17. Foote & McDonough (2017). Journal of Second Language Pronunciation 3(1). https://doi.org/10.1075/jslp.3.1.02foo
  18. Lu & Su (2024). Journal of Second Language Pronunciation 10(1). https://doi.org/10.1075/jslp.22033.lu
  19. Nation (1989). System 17(3). https://doi.org/10.1016/0346-251X(89)90010-9
  20. Thai & Boers (2016). TESOL Quarterly 50(2). https://doi.org/10.1002/tesq.232
  21. Boers (2014). RELC Journal 45(3). https://doi.org/10.1177/0033688214546964
  22. Huang & Liu (2022). English Teaching & Learning 47(2). https://europepmc.org/article/PMC/PMC8934908
  23. Chen et al. (2019). Speech Communication 115. https://doi.org/10.1016/j.specom.2019.10.008
  24. Saito & Plonsky (2019). Language Learning 69(3). https://doi.org/10.1111/lang.12345
  25. Wang, Jongman & Sereno (2003). Journal of the Acoustical Society of America 113(2). https://doi.org/10.1121/1.1531176
  26. Patel, Xu & Wang (2010). Speech Prosody 2010. https://doi.org/10.21437/speechprosody.2010-238
  27. Horwitz, Horwitz & Cope (1986). Modern Language Journal 70(2). https://doi.org/10.1111/j.1540-4781.1986.tb05256.x
  28. Teimouri, Goetze & Plonsky (2019). SSLA 41(2). https://doi.org/10.1017/S0272263118000311
  29. Zhang (2019). Modern Language Journal 103(4). https://doi.org/10.1111/modl.12590
  30. Bashori, van Hout, Strik & Cucchiarini (2021). System 99. https://doi.org/10.1016/j.system.2021.102496
  31. OpenAI Help, ChatGPT Voice (checked 6 Oct 2026). https://help.openai.com/en/articles/20001274-chatgpt-voice
  32. Google, Gemini Live help (checked 6 Oct 2026). https://support.google.com/gemini/answer/15274899
  33. Linge, O. (2026). No, ChatGPT can't analyse your Mandarin pronunciation. Hacking Chinese. https://www.hackingchinese.com/no-chatgpt-cant-analyse-your-mandarin-pronunciation/
  34. US Foreign Service Institute, Foreign Language Training. https://www.state.gov/foreign-language-training/
  35. SayMei data: first-lesson speaking turns, 15 Jul to 7 Oct 2026 (aggregates only).
  36. Vendor pages checked 6 Oct 2026: SayMei pricing and first lesson; Speak; Preply; US App Store listings for ChatGPT, Speak, Praktika, Pingo AI, Talkpal, SuperChinese, HelloTalk.
  37. italki Chinese teacher directory, all public pages read 7 Oct 2026 (763 teachers; median lowest hourly price $20, community tutors $18, professional teachers $22). https://www.italki.com/en/teachers/chinese

Methods

This guide draws on a research dossier assembled on 6 October 2026: 30 research papers and books, each claim checked against at least the published abstract; public price and help pages; and aggregate SayMei lesson data, with no transcripts and no personal data. Every published SayMei number covers at least 50 learners. Every Mandarin example, including the drill's model answers, was checked against the CC-CEDICT dictionary (CC BY-SA 4.0). No app was tested hands-on. If you spot a mistake, tell us: we fix errors.

Mei Lin, SayMei's AI teacher, smiling and waving

Ready to get speaking?

Talk with Mei Lin, your AI teacher, free for 10 minutes.

Start speaking

Mei Lin is an AI tutor. No account or card needed.