Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity

Sección IA: Inteligencia artificial Bibliografía

Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity

Manru Wang
Yihan Chen
Xiaoting Huang
Yuxuan Lai
2026
Computers & Education: Artificial Intelligence
10
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
expresión escrita
evaluación
grandes modelos de lenguaje
enseñanza/aprendizaje de lenguas
IA y enseñanza-aprendizaje de lenguas
IA y evaluación
modelos de lenguaje (LLM)
estudio empírico

Texto completo

Large language models (LLMs) are increasingly used for automated essay scoring, yet their underlying scoring mechanisms remain insufficiently understood. This study systematically compared the scoring behavior of three LLMs (Qwen, GPT, and Gemini) with human raters on English essays written by non-native learners. Sixteen textual features were analyzed to compare score alignment, feature weighting, subgroup consistency, and feature interactions. Results showed strong overall alignment but distinct feature weighting patterns between the LLMs and human raters. Specifically, the LLMs placed greater emphasis on grammatical accuracy, lexical sophistication, and syntactic complexity, indicating a stronger preference for formal precision and linguistic sophistication, whereas human raters prioritized content completeness and visual presentation, displaying greater tolerance toward minor linguistic errors. Across proficiency levels, human raters exhibited a more stable scoring framework. LLMs, however, showed larger cross-group shifts, placing more weight on language errors for low-proficiency students and increasingly rewarding linguistic sophistication for high-proficiency students. Interaction analysis further revealed that the LLMs integrated multiple features when scoring, with this integration pattern varies by proficiency level. These findings highlight both the potential and limitations of LLM-based scoring and underscore the importance of interpretability and transparency to enhance the validity of automated scoring in educational settings.

Texto completo en abierto (CC BY-NC-ND 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability
  • A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning
  • Level-specific feedback generation for scene descriptions via fine-tuning multimodal large language models
  • How well can LLMs grade essays in Arabic?
  • Enhancing language learning through generative AI feedback on picture-cued writing tasks
  • LLMs do not grade essays like humans
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos