Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • How well can LLMs grade essays in Arabic?

Sección IA: Inteligencia artificial Bibliografía

How well can LLMs grade essays in Arabic?

Rayed Ghazawi
Edwin Simpson
2025
Computers & Education: Artificial Intelligence
9
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
expresión escrita
evaluación
grandes modelos de lenguaje
IA y enseñanza-aprendizaje de lenguas
IA y evaluación
modelos de lenguaje (LLM)
prompts
estudio empírico

Texto completo

This research assesses the effectiveness of state-of-the-art large language models (LLMs), including ChatGPT, Llama, Aya, Jais, and ACEGPT, in the task of Arabic automated essay scoring (AES) using the AR-AES dataset. It explores various evaluation methodologies, including zero-shot, few-shot in context learning, and fine-tuning, and examines the influence of instruction-following capabilities through the inclusion of marking guidelines within the prompts. A mixed-language prompting strategy, integrating English prompts with Arabic content, was implemented to improve model comprehension and performance. Among the models tested, ACEGPT demonstrated the strongest performance across the dataset, achieving a Quadratic Weighted Kappa (QWK) of 0.67, but was outperformed by a smaller BERT-based model with a QWK of 0.88. The study identifies challenges faced by LLMs in processing Arabic, including tokenization complexities and higher computational demands. Performance variation across different courses underscores the need for adaptive models capable of handling diverse assessment formats and highlights the positive impact of effective prompt engineering on improving LLM outputs. To the best of our knowledge, this study is the first to empirically evaluate the performance of multiple generative Large Language Models (LLMs) on Arabic essays using authentic student data.

Texto completo en abierto (CC BY 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • LLMs do not grade essays like humans
  • Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability
  • Evaluating large language models as raters in large-scale writing assessments: A psychometric framework for reliability and validity
  • A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning
  • Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity
  • Level-specific feedback generation for scene descriptions via fine-tuning multimodal large language models
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos