Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Assisting quality assurance of examination tasks: Using a GPT model and Bayesian testing for formative assessment

Sección IA: Inteligencia artificial Bibliografía

Assisting quality assurance of examination tasks: Using a GPT model and Bayesian testing for formative assessment

Nico Willert
Phi Katharina Würz
2025
Computers & Education: Artificial Intelligence
8
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
evaluación
ChatGPT
educación superior
herramientas
tecnología educativa
IA y evaluación
chatGPT
prompts
estudio empírico

Texto completo

Formative quality assurance in the creation of examination tasks has always been an extremely time-consuming process. Especially due to the changing and short-lived content of computer science, new questions have to be created regularly, which in turn requires quality assurance. With the emergence of artificial intelligence (AI) systems such as ChatGPT and their ability to solve a range of different tasks, the question arises as to what extent this ability can also be utilized as part of a quality assurance process. One aspect of the formative quality assurance of multiple-choice questions involves checking the correct classification of alternative answers into correct and incorrect answers. As AI systems inherently lack transparency and predictability in their output, we present a simplified approach using Bayesian hypothesis testing to estimate the tendencies of an AI towards the classification. To evaluate the approach, the process is implemented and connected to the OpenAI API to handle inconsistent responses and other aspects that contribute to the robustness and reliability. This research is concluded by an evaluation carried out by means of the gpt-3.5-turbo model, using the examination tasks of two programming courses. This provides insights into the response scheme of the AI in relation to the prompt pattern used and the usability of AI for the subsequent quality assurance process.

Texto completo en abierto (CC BY-NC 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Integrating the adapted UTAUT model with moral obligation, trust and perceived risk to predict ChatGPT adoption for assessment support: A survey with students
  • Comparing expert tutor evaluation of reflective essays with marking by generative artificial intelligence (AI) tool
  • The rapid rise of generative AI and its implications for academic integrity: Students’ perceptions and use of chatbots for assistance with assessments
  • Investigating the affordances of OpenAI's large language model in developing listening assessments
  • Reinventing assessments with ChatGPT and other online tools: Opportunities for GenAI-empowered assessment practices
  • How to train your dragon: Evaluating prompting and fine-tuning for GPT-based item generation in L2 listening assessment
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos