Inicio
todoELE
  • Inicio
  • Materiales
    • ๐Ÿ“‹ Actividades
    • ๐Ÿ“ Conjugación
    • ๐Ÿ“Š Corpus
    • ๐Ÿ“” Diccionarios
    • โœ… Evaluación
    • โš™๏ธ Gramática
    • ๐Ÿ“— Manuales
    • โœ๏ธ Ortografía
    • ๐Ÿ“… Programación
    • ๐Ÿ—ฃ๏ธ Pronunciación
    • ๐Ÿ“ Recursos
    • ๐Ÿ”ค Vocabulario
    • ๐Ÿ’ป Herramientas digitales
  • Formación
    • ๐Ÿ“š Bibliografía
    • ๐Ÿ‘ฅ Congresos
    • ๐ŸŽ“ Cursos
    • ๐Ÿซ Centros
    • ๐Ÿข Organizaciones
    • ๐Ÿ“ฐ Revistas
    • ๐ŸŒ Atlas de ELE
  • Trabajo
    • ๐Ÿ’ผ Ofertas de trabajo
    • โ„น๏ธ Trabajo - Recursos
  • En la red
    • ๐ŸŒ Sitios ELE
    • ๐Ÿ“ฐ Agregador
    • ๐Ÿ“ง Formespa
  • IA
    • โœจ Nuevos contenidos
    • ๐Ÿ“š Bibliografía IA
    • ๐Ÿงฐ Herramientas IA
    • ๐Ÿ’ฌ Prompts
    • ๐Ÿงช Experiencias IA
    • ๐ŸŒ Sitios web IA
    • ๐Ÿ“ฐ Actualidad IA
  • Comunidad
    • ๐Ÿ“ฐ Actualidad ELE
    • ๐Ÿ˜Š Anécdotas ELE
    • ๐Ÿ“ Blog
    • ๐Ÿ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Assessing the quality of automatic-generated short answers using GPT-4

Sección IA: Inteligencia artificial Bibliografía

Assessing the quality of automatic-generated short answers using GPT-4

Luiz Rodrigues
Filipe Dwan Pereira
Luciano Cabral
Dragan Gaševiฤ‡
Geber Ramalho
Rafael Ferreira Mello
2024
Computers & Education: Artificial Intelligence
7
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
grandes modelos de lenguaje
evaluación
educación superior
análisis de producción de IA
modelos de lenguaje (LLM)
IA y evaluación
estudio empírico

Texto completo

Open-ended assessments play a pivotal role in enabling instructors to evaluate student knowledge acquisition and provide constructive feedback. Integrating large language models (LLMs) such as GPT-4 in educational settings presents a transformative opportunity for assessment methodologies. However, existing literature on LLMs addressing open-ended questions lacks breadth, relying on limited data or overlooking question difficulty levels. This study evaluates GPT-4's proficiency in responding to open-ended questions spanning diverse topics and cognitive complexities in comparison to human responses. To facilitate this assessment, we generated a dataset of 738 open-ended questions across Biology, Earth Sciences, and Physics and systematically categorized it based on Bloom's Taxonomy. Each question included eight human-generated responses and two from GPT-4. The outcomes indicate GPT-4's superior performance over humans, encompassing both native and non-native speakers, irrespective of gender. Nevertheless, this advantage was not sustained in ’remembering’ or ’creating’ questions aligned with Bloom's Taxonomy. These results highlight GPT-4's potential for underpinning advanced question-answering systems, its promising role in supporting non-native speakers, and its capacity to augment teacher assistance in assessments. However, limitations in nuanced argumentation and creativity underscore areas necessitating refinement in these models, guiding future research toward bolstering pedagogical support.

Texto completo en abierto (CC BY-NC-ND 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Assessing the proficiency of large language models in automatic feedback generation: An evaluation study
  • Automatic question-answer pairs generation using pre-trained large language models in higher education
  • How do LLMs perform in the context of MCQs across different levels of thinking skills in a business education course at higher education? A comparison of ChatGPT, Gemini, and Copilot
  • A hybrid reasoning framework for artificial intelligence assessment rubric generation in human and automated contexts: Evidence from an undergraduate programming course
  • Is GPT-4 fair? An empirical analysis in automatic short answer grading
  • LLMs do not grade essays like humans
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos