Beyond agreement statistics: subgroup fairness in automated Chinese writing assessment

Authors

  • Cong Liu Author
  • Zhiqiang Zhao Author

Keywords:

Chinese language assessment; automated writing evaluation; fairness and validity; heritage language learners; differential item functioning; many-facets Rasch measurement

Abstract

Automated scoring of second-language writing is now used for placement, progress monitoring and increasingly consequential decisions. Its fairness relies on an untested (outside English) assumption: the construct scored by the engine is identical for all learner subgroups. This study tests the assumption for Chinese, using a corpus of 412 argumentative and narrative essays from 206 Chinese learners: 138 essays from 69 heritage learners (Mandarin/Cantonese family background, no formal literacy instruction) and 274 essays from 137 non-heritage learners. All essays were independently rated by two trained human raters on a six-point holistic scale and five-category analytic rubric, with discrepant cases adjudicated. A transformer-based scoring engine was then fitted and evaluated.

Engine-human agreement was lower for heritage learners than non-heritage learners (quadratic weighted κ = .65 versus .82). A many-facets Rasch analysis found a reliable heritage × engine bias interaction: at equal human-rated proficiency, heritage learners received engine scores approximately 0.38 scale points lower. Feature decomposition showed the gap derived from the engine's weighting of character-form accuracy and treatment of explicit discourse connectives, both of which penalize discourse-competent but orthographically less secure writing — a profile common among heritage learners. A script-variety sub-analysis found an additional penalty for traditional character essays. At the placement cut score, the misclassification rate was 14.5% for heritage learners versus 6.2% for non-heritage learners.

We argue that automated Chinese scoring fairness cannot be confirmed by aggregate agreement statistics, heritage status is a construct-relevant rather than nuisance variable, and automated feedback from such engines risks transmitting systematic bias to learners with the most uneven literacy development.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-15

Issue

Section

Articles