Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • LLMs do not grade essays like humans

Sección IA: Inteligencia artificial Bibliografía

LLMs do not grade essays like humans

Jerin George Mathew
Sumayya Taher
Anindita Kundu
Denilson Barbosa
2026
Computers & Education: Artificial Intelligence
11
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
expresión escrita
evaluación
grandes modelos de lenguaje
IA y enseñanza-aprendizaje de lenguas
IA y evaluación
modelos de lenguaje (LLM)
análisis de producción de IA
estudio empírico

Texto completo

Large language models have recently been proposed as tools for automated essay scoring, but their agreement with human grading remains unclear. In this work, we evaluate how LLM-generated scores compare with human grades and analyze the grading behavior of several models from the GPT and Llama families in an out-of-the-box setting, without task-specific training. Our results show that agreement between LLM and human scores remains relatively weak and varies with essay characteristics. In particular, compared to human raters, LLMs tend to assign higher scores to short or underdeveloped essays, while assigning lower scores to longer essays that contain minor grammatical or spelling errors. We also find that the scores generated by LLMs are generally consistent with the feedback they generate: essays receiving more praise tend to receive higher scores, while essays receiving more criticism tend to receive lower scores. These results suggest that LLM-generated scores and feedback follow coherent patterns but rely on signals that differ from those used by human raters, resulting in limited alignment with human grading practices. Nevertheless, our work shows that LLMs produce feedback that is consistent with their grading and that they can be reliably used in supporting essay scoring.

Texto completo en abierto (CC BY-NC 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • How well can LLMs grade essays in Arabic?
  • Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability
  • Automated reading passage generation with OpenAI's large language model
  • Evaluating large language models as raters in large-scale writing assessments: A psychometric framework for reliability and validity
  • A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning
  • Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos