Image-native automated scoring of handwritten mathematical responses: reliability evidence and teacher–AI collaboration

JiEun Janet Song, Young-seok Oh, Dong Joong Kim

Abstract


This study examines the reliability of an image-native multimodal AI system for automated scoring of handwritten responses to Advanced Placement (AP) Calculus free-response items without requiring optical character recognition (OCR) preprocessing. Using inter-rater agreement indices and test–retest reliability analyses, we found substantial to almost perfect agreement between artificial intelligence (AI)-generated scores and calibrated human ratings, as well as almost perfect stability across repeated scoring sessions. These results suggest that the observed reliability of the AI scoring system warrants further investigation of validity-related evidence and inferences. As a practical implication for assessment practice, we propose a human-in-the-loop teacher–AI collaborative (TAC) framework in which automated scoring operates under teacher oversight. Taken together, these findings provide initial evidence of reliability supporting the responsible use of AI-based scoring as a measurement instrument in high-stakes educational assessment.

Keywords


Automated scoring; Handwritten mathematics; Image-native grading; Multimodal large language model; Reliability; Rubric-based scoring; Teacher–AI collaboration

Full Text:

PDF


DOI: http://doi.org/10.11591/ijere.v15i4.38456

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 JiEun Janet Song, Young-Seok Oh, Dong Joong Kim

International Journal of Evaluation and Research in Education (IJERE)
p-ISSN: 2252-8822e-ISSN: 2620-5440
The journal is published by Institute of Advanced Engineering and Science (IAES).

View IJERE Stats

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.