Adaptive AI feedback and the acquisition of the Mandarin Tone 2–Tone 3 contrast: A mixed-methods study with beginner learners in a non-target-language environment
Keywords:
Chinese as a foreign language; lexical tone; speech perception training; adaptive feedback; generative and speech AI; pronunciation instructionAbstract
The Mandarin Tone 2–Tone 3 contrast is a persistent perceptual obstacle for adult learners, least reliably resolved by non-tonal first-language listeners. This study investigates whether adaptive feedback organisation, rather than algorithmic tutor presence, alters beginner learners' perception of this contrast in non-Chinese-speaking environments.Sixty-eight beginner CFL learners (dominant L1 Urdu or English, recruited from four Pakistan and Nigeria universities) were split into three 10-session groups: contingent AI feedback re-targeting stimuli per learner error profile (n = 23), same AI engine with fixed pre-sequenced items (n = 22), no-feedback control with identical perceptual exposure (n = 23).Outcomes were measured pre-intervention, post-intervention and four weeks later, covering two-alternative forced-choice identification accuracy, AXB discrimination sensitivity (d′), and generalisation to untrained talkers. Twelve learners were interviewed on system engagement.Results show a Time × Group interaction for identification (F(4, 122) = 37.19, p < .001, partial η² = .549) and discrimination sensitivity (F(4, 122) = 37.37, p < .001, partial η² = .551): the adaptive condition outperformed the fixed-sequence condition by 10.5 percentage points (d = 1.19) at post-test, retaining the 10.5-point advantage (d = 1.12) at delayed post-test; the fixed condition outperformed the no-feedback control by 10.7 points; the adaptive advantage reached 10.4 points (d = 0.92) for untrained talkers.
Interviews show learner benefit depends on the ability to calibrate trust in the system, which is damaged by connected-speech Tone 3 classification errors. The study concludes the active factor is feedback loop contingency rather than artificial intelligence itself, noting implications for tone instruction design in low-input settings.