Speaker
Description
Functioning as a powerful Large Language Model (LLM), generative artificial intelligence (Gen AI) has helped encourage learners’ learning autonomy in many aspects which used to be considered challenging for self-learning like skill practice. In writing, it is utilized to create sample responses, which assists learners in visualizing the expectation interpreted from the band descriptors. However, concerns have been raised against the authenticity and alignment of those models as they are usually overly polished. This research aimed to investigate Gen AI’s ability to generate level-calibrated writing for VSTEP task 2. The researchers employed mixed-methods comparative design, in which 40 AI-generated written responses at level B1 and B2 were evaluated to check whether they aligned with intended proficiency levels and resembled authentic learner performances. The results show that Gen AI struggles to mimic lower-level discourse organizations and lexical resources whereas it performs better at higher bands, which clearly distort low-level learners’ expectations.
Key words: VSTEP, Gen AI, writing, assessment alignment, score validity