Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability

Sección IA: Inteligencia artificial Bibliografía

Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability

Austin Pack
Alex Barrett
Juan Escalante
2024
Computers & Education: Artificial Intelligence
6
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
expresión escrita
evaluación
grandes modelos de lenguaje
enseñanza/aprendizaje de lenguas
IA y enseñanza-aprendizaje de lenguas
IA y evaluación
modelos de lenguaje (LLM)
estudio empírico

Texto completo

Advancements in generative AI, such as large language models (LLMs), may serve as a potential solution to the burdensome task of essay grading often faced by language education teachers. Yet, the validity and reliability of leveraging LLMs for automatic essay scoring (AES) in language education is not well understood. To address this, we evaluated the cross-sectional and longitudinal validity and reliability of four prominent LLMs, Google's PaLM 2, Anthropic's Claude 2, and OpenAI's GPT-3.5 and GPT-4, for the AES of English language learners' writing. 119 essays taken from an English language placement test were assessed twice by each LLM, on two separate occasions, as well as by a pair of human raters. GPT-4 performed the best, demonstrating excellent intrarater reliability and good validity. All models, with the exception of GPT-3.5, improved over time in their intrarater reliability. The interrater reliability of GPT-3.5 and GPT-4, however, decreased slightly over time. These findings indicate that some models perform better than others in AES and that all models are subject to fluctuations in their performance. We discuss potential reasons for such variability, and offer suggestions for prospective avenues of research.

Texto completo en abierto (CC BY-NC 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning
  • Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity
  • Level-specific feedback generation for scene descriptions via fine-tuning multimodal large language models
  • How well can LLMs grade essays in Arabic?
  • Enhancing language learning through generative AI feedback on picture-cued writing tasks
  • LLMs do not grade essays like humans
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos