Beyond agreement statistics: subgroup fairness in automated Chinese writing assessment
Keywords:
Chinese language assessment; automated writing evaluation; fairness and validity; heritage language learners; differential item functioning; many-facets Rasch measurementAbstract
Automated scoring of second-language writing is now used for placement, progress monitoring and increasingly consequential decisions. Its fairness relies on an untested (outside English) assumption: the construct scored by the engine is identical for all learner subgroups. This study tests the assumption for Chinese, using a corpus of 412 argumentative and narrative essays from 206 Chinese learners: 138 essays from 69 heritage learners (Mandarin/Cantonese family background, no formal literacy instruction) and 274 essays from 137 non-heritage learners. All essays were independently rated by two trained human raters on a six-point holistic scale and five-category analytic rubric, with discrepant cases adjudicated. A transformer-based scoring engine was then fitted and evaluated.
Engine-human agreement was lower for heritage learners than non-heritage learners (quadratic weighted κ = .65 versus .82). A many-facets Rasch analysis found a reliable heritage × engine bias interaction: at equal human-rated proficiency, heritage learners received engine scores approximately 0.38 scale points lower. Feature decomposition showed the gap derived from the engine's weighting of character-form accuracy and treatment of explicit discourse connectives, both of which penalize discourse-competent but orthographically less secure writing — a profile common among heritage learners. A script-variety sub-analysis found an additional penalty for traditional character essays. At the placement cut score, the misclassification rate was 14.5% for heritage learners versus 6.2% for non-heritage learners.
We argue that automated Chinese scoring fairness cannot be confirmed by aggregate agreement statistics, heritage status is a construct-relevant rather than nuisance variable, and automated feedback from such engines risks transmitting systematic bias to learners with the most uneven literacy development.