Dataset composition as a curricular decision: A systematic review and equity audit of datasets, corpora and analytics in Chinese education
Keywords:
Chinese language education; learner corpora; datasets; knowledge graphs; learning analytics; equity audit; systematic review; research transparencyAbstract
The infrastructure on which AI-mediated Chinese language education depends — learner corpora, speech and tone datasets, character and handwriting resources, lexical knowledge graphs, and learning-analytics pipelines — is not a neutral substrate. What a dataset contains determines what a system can learn to see, and what it omits determines whose learning will be misread. This article reports a registered systematic review of that infrastructure, combined with a five-dimension equity audit of the resources themselves. Ninety-six studies published between 2015 and 2025 met the inclusion criteria, selected from 2,847 records across four databases. The review characterises the resource landscape, appraises reporting practice against the transparency standard adopted here, and audits each resource for representational, definitional, procedural, distributional and epistemic equity. Resources were dominated by university learners in target-language environments: 68 of 96 were collected where Chinese is the surrounding language, only 24 in non-target-language settings; heritage learners were explicitly identified in 7 resources (7.3 per cent), learners under 18 were the primary population in 19, of which only 8 reported parental consent procedures, non-Mandarin varieties appeared in 6, and learners writing in traditional characters in 9. Reporting practice was uneven: model versions in 61 resources (64 per cent), full prompt or instruction specifications in 9 of the 38 generative-AI resources (24 per cent), obtainable data in 11 of the 44 carrying a data-availability statement, and ethics approval in 57 (59 per cent). The audit identifies a systematic mismatch between the populations Chinese-language AI research draws on and those the field's own literature identifies as under-served. We argue that dataset composition is a curricular decision with measurable consequences, propose a six-item reporting standard for resources intended for Chinese language education, and set out an agenda around the four absences the audit reveals.