Large language models (LLMs) are increasingly used for automated essay scoring, yet their underlying scoring mechanisms remain insufficiently understood. This study systematically compared the scoring behavior of three LLMs (Qwen, GPT, and Gemini) with human raters on English essays written by non-native learners. Sixteen textual features were analyzed to compare score alignment, feature weighting, subgroup consistency, and feature interactions. Results showed strong overall alignment but distinct feature weighting patterns between the LLMs and human raters. Specifically, the LLMs placed greater emphasis on grammatical accuracy, lexical sophistication, and syntactic complexity, indicating a stronger preference for formal precision and linguistic sophistication, whereas human raters prioritized content completeness and visual presentation, displaying greater tolerance toward minor linguistic errors. Across proficiency levels, human raters exhibited a more stable scoring framework. LLMs, however, showed larger cross-group shifts, placing more weight on language errors for low-proficiency students and increasingly rewarding linguistic sophistication for high-proficiency students. Interaction analysis further revealed that the LLMs integrated multiple features when scoring, with this integration pattern varies by proficiency level. These findings highlight both the potential and limitations of LLM-based scoring and underscore the importance of interpretability and transparency to enhance the validity of automated scoring in educational settings.
- Inicie sesión para enviar comentarios