Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning

Sección IA: Inteligencia artificial Bibliografía

A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning

Tongxi Liu
Luyao Ye
Wei Yan
2026
Computers & Education: Artificial Intelligence
10
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
expresión escrita
evaluación
grandes modelos de lenguaje
enseñanza/aprendizaje de lenguas
educación superior
IA y enseñanza-aprendizaje de lenguas
IA y evaluación
modelos de lenguaje (LLM)
estudio empírico

Texto completo

Recent advances in large language models have revitalized research on automated essay evaluation, yet critical concerns remain regarding their reliability, validity, and interpretability. This study presents a comparative analysis of five LLMs (GPT-4.1, Llama 4 Maverick, Gemini 2.5 Flash, Claude Sonnet 4, and DeepSeek R1) in the assessment of long English essays authored by non-native speakers in higher education. The analysis draws on LLM-generated scores for 60 essays to examine (a) intra-model reliability across repeated scoring runs, (b) the degree of alignment between model outputs and expert human ratings, and (c) causal feature dependencies that clarify how linguistic characteristics influence model scoring behavior. Findings reveal substantial variation: some models achieved near-perfect reproducibility and strong alignment with human raters, whereas others displayed inconsistency, score compression, or systematic underestimation. Causal discovery analysis further uncovered distinct evaluative heuristics, with most models prioritizing lexical precision and fluency, while others emphasized syntactic complexity or cross-domain integration. Collectively, these results establish model-specific reliability profiles and application contexts, providing empirical benchmarks and practical guidance for the responsible use of LLMs in educational writing assessment.

Texto completo en abierto (CC BY 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability
  • Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity
  • Level-specific feedback generation for scene descriptions via fine-tuning multimodal large language models
  • How well can LLMs grade essays in Arabic?
  • Enhancing language learning through generative AI feedback on picture-cued writing tasks
  • LLMs do not grade essays like humans
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos