Inicio
todoELE
  • Inicio
  • Materiales
    • ๐Ÿ“‹ Actividades
    • ๐Ÿ“ Conjugación
    • ๐Ÿ“Š Corpus
    • ๐Ÿ“” Diccionarios
    • โœ… Evaluación
    • โš™๏ธ Gramática
    • ๐Ÿ“— Manuales
    • โœ๏ธ Ortografía
    • ๐Ÿ“… Programación
    • ๐Ÿ—ฃ๏ธ Pronunciación
    • ๐Ÿ“ Recursos
    • ๐Ÿ”ค Vocabulario
    • ๐Ÿ’ป Herramientas digitales
  • Formación
    • ๐Ÿ“š Bibliografía
    • ๐Ÿ‘ฅ Congresos
    • ๐ŸŽ“ Cursos
    • ๐Ÿซ Centros
    • ๐Ÿข Organizaciones
    • ๐Ÿ“ฐ Revistas
    • ๐ŸŒ Atlas de ELE
  • Trabajo
    • ๐Ÿ’ผ Ofertas de trabajo
    • โ„น๏ธ Trabajo - Recursos
  • En la red
    • ๐ŸŒ Sitios ELE
    • ๐Ÿ“ฐ Agregador
    • ๐Ÿ“ง Formespa
  • IA
    • โœจ Nuevos contenidos
    • ๐Ÿ“š Bibliografía IA
    • ๐Ÿงฐ Herramientas IA
    • ๐Ÿ’ฌ Prompts
    • ๐Ÿงช Experiencias IA
    • ๐ŸŒ Sitios web IA
    • ๐Ÿ“ฐ Actualidad IA
  • Comunidad
    • ๐Ÿ“ฐ Actualidad ELE
    • ๐Ÿ˜Š Anécdotas ELE
    • ๐Ÿ“ Blog
    • ๐Ÿ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Is GPT-4 fair? An empirical analysis in automatic short answer grading

Sección IA: Inteligencia artificial Bibliografía

Is GPT-4 fair? An empirical analysis in automatic short answer grading

Luiz Rodrigues
Cleon Xavier
Newarney Costa
Dragan Gaševiฤ‡
Rafael Ferreira Mello
2025
Computers & Education: Artificial Intelligence
8
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
evaluación
grandes modelos de lenguaje
ética
inclusión y equidad
IA y evaluación
modelos de lenguaje (LLM)
ética de la IA
estudio empírico

Texto completo

Short open-ended questions represent a central resource in formative and summative assessments both face-to-face and online settings, ranging from elementary to higher education. However, grading these questions remains challenging for instructors, raising attention to the field of Automatic Short Answer Grading (ASAG). While ASAG has yielded valuable contributions to learning analytics, it often faces generalizability issues. Accordingly, the rapid advancement in Large Language Models (LLMs) has motivated their adoption to empower ASAG systems. Despite that, previous research has not investigated whether LLMs are fair graders in the context of ASAG. Therefore, this paper presents an empirical analysis aimed to understand LLMs' fairness in ASAG by using human grades as a baseline, comparing them to GPT-4's answers, and investigating whether the LLM's grades are equivalent in grading answers from varied groups of humans. Our results demonstrated that, while GPT-4 tended to be more lenient in its grading, it maintained consistent evaluation standards when assessing responses from different groups of students. GPT-4 remained consistent for questions of different subjects and levels of Bloom's taxonomy and for people with different demographics. These findings suggest GPT-4 is a fair grader, supporting its potential to empower educators and developers in using and designing ASAG systems. Nevertheless, we recommend further research to investigate these findings and understand how to optimize GPT-4's grades.

Texto completo en abierto (CC BY-NC-ND 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Towards responsible AI in education: A Delphi-AHP-based framework for evaluating educational large language models
  • Assessing the quality of automatic-generated short answers using GPT-4
  • Appraisal of high-stake examinations during SARS-CoV-2 emergency with responsible and transparent AI: Evidence of fair and detrimental assessment
  • Not for people like me: How frontier AI models redirect skeptical rural school staff
  • Beyond binary outcomes: Evaluating and mitigating bias in national standardized test score prediction
  • Assessing the proficiency of large language models in automatic feedback generation: An evaluation study
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos