Inicio
todoELE
  • Inicio
  • Materiales
    • ๐Ÿ“‹ Actividades
    • ๐Ÿ“ Conjugación
    • ๐Ÿ“Š Corpus
    • ๐Ÿ“” Diccionarios
    • โœ… Evaluación
    • โš™๏ธ Gramática
    • ๐Ÿ“— Manuales
    • โœ๏ธ Ortografía
    • ๐Ÿ“… Programación
    • ๐Ÿ—ฃ๏ธ Pronunciación
    • ๐Ÿ“ Recursos
    • ๐Ÿ”ค Vocabulario
    • ๐Ÿ’ป Herramientas digitales
  • Formación
    • ๐Ÿ“š Bibliografía
    • ๐Ÿ‘ฅ Congresos
    • ๐ŸŽ“ Cursos
    • ๐Ÿซ Centros
    • ๐Ÿข Organizaciones
    • ๐Ÿ“ฐ Revistas
    • ๐ŸŒ Atlas de ELE
  • Trabajo
    • ๐Ÿ’ผ Ofertas de trabajo
    • โ„น๏ธ Trabajo - Recursos
  • En la red
    • ๐ŸŒ Sitios ELE
    • ๐Ÿ“ฐ Agregador
    • ๐Ÿ“ง Formespa
  • IA
    • โœจ Nuevos contenidos
    • ๐Ÿ“š Bibliografía IA
    • ๐Ÿงฐ Herramientas IA
    • ๐Ÿ’ฌ Prompts
    • ๐Ÿงช Experiencias IA
    • ๐ŸŒ Sitios web IA
    • ๐Ÿ“ฐ Actualidad IA
  • Comunidad
    • ๐Ÿ“ฐ Actualidad ELE
    • ๐Ÿ˜Š Anécdotas ELE
    • ๐Ÿ“ Blog
    • ๐Ÿ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Evaluating the capability of large language models in characterising relational feedback: A comparative analysis of prompting strategies

Sección IA: Inteligencia artificial Bibliografía

Evaluating the capability of large language models in characterising relational feedback: A comparative analysis of prompting strategies

Wei Dai
Yixin Cheng
Ahmad Ari Aldino
Yi-Shan Tsai
Dragan Gaševiฤ‡
Guanliang Chen
2025
Computers & Education: Artificial Intelligence
8
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
feedback/retroalimentación
grandes modelos de lenguaje
modelos de lenguaje (LLM)
prompts
IA y evaluación
estudio empírico

Texto completo

Relational feedback is increasingly recognised for its crucial role in enhancing student-instructor relationships and promoting the assimilation of feedback. Despite its significance, no studies have tried to develop automated methods to analyse written feedback for properties of relational feedback to promote its use at scale and assist feedback providers with their relational feedback practices. This automated analysis of relational feedback can be performed as a classification task. However, traditional machine and deep learning methods for text classification typically require extensive human labelling and pose a significant challenge for educators and researchers lacking machine learning and data science expertise. Large language models offer a promising solution due to their advancements in text classification tasks and their capacity to interact using natural-language prompts. Prompting strategies can significantly influence model performance; however, it remains unclear how the prompt should be designed to enable the accurate characterisation of relational feedback. Therefore, this study aims to investigate the capability of GPT-4o, the versatile and flagship model by OpenAI, in characterising relational feedback and evaluate how its effectiveness varies across zero-shot, one-shot and few-shot prompting strategies. Results from extensive experiments conducted on a real-world dataset comprising 793 feedback sentences revealed that: i) GPT-4o achieved an average Accuracy exceeding 0.8 in identifying nine out of ten relational characteristics, and an average F1 score exceeding 0.7 in identifying six out of ten relational feedback characteristics; ii) GPT-4o's classification performance demonstrated no significant differences across various prompting strategies for eight out of ten relational characteristics; iii) GPT-4o's classification performance was enhanced by explicitly distinguishing between related relational characteristics within the prompt. Our findings underscore the potential of large language models in identifying relational feedback and indicate that providing a clear definition of relational characteristics enhances classification performance more effectively than incorporating exemplars in the prompt.

Texto completo en abierto (CC BY 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Assessing the proficiency of large language models in automatic feedback generation: An evaluation study
  • Large language models meet user interfaces: The case of provisioning feedback
  • Applying large language models and chain-of-thought for automatic scoring
  • Can large language models meet the challenge of generating school-level questions?
  • Automatic item generation in various STEM subjects using large language model prompting
  • Assessing student errors in experimentation using artificial intelligence and large language models: A comparative study with human raters
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos