Speakers
Description
Assessing argumentative writing remains a challenge because it requires not only consistent scoring but also informed judgment of students' reasoning, creativity, and personal perspectives. Recent advances in generative artificial intelligence (AI) have created new opportunities to support writing assessment. However, most existing studies have focused on scoring accuracy, with limited empirical evidence on how AI performs when assessing argumentative writing in educational contexts where evaluation relies heavily on teacher judgment, such as Vietnamese lower secondary education. This study investigates the role of generative AI in argumentative writing assessment by comparing AI-generated evaluations with those of teachers to identify which aspects of assessment can be effectively supported by AI and which continue to require teacher expertise. A quasi-experimental design was conducted using 80 argumentative essays written in Vietnamese by Grade 8 students. Each essay was independently assessed using the same analytic rubric by five experienced lower secondary Vietnamese language and literature teachers and ChatGPT 5.5. Data were analyzed using descriptive and comparative statistics within Messick's (1989) Unified Validity Framework. The findings suggest that AI provides more consistent and transparent scoring for criteria explicitly defined in the rubric, whereas teachers remain better at evaluating students' reasoning, creativity, contextual interpretation, and personal expression. The study proposes an AI–teacher hybrid assessment model and offers practical implications for AI-assisted writing assessment, teacher AI literacy, and the responsible integration of AI into language education.
Keywords: artificial intelligence, writing assessment, AI-assisted assessment, teacher AI literacy