Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Using convolutional neural networks to automatically score eight TIMSS 2019 graphical response items

Sección IA: Inteligencia artificial Bibliografía

Using convolutional neural networks to automatically score eight TIMSS 2019 graphical response items

Lillian Tyack
Lale Khorramdel
Matthias von Davier
2024
Computers & Education: Artificial Intelligence
6
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
evaluación
educación primaria y secundaria
IA y evaluación
estudio empírico

Texto completo

International large-scale assessments (ILSAs) have used graphical response-based items to measure student ability for decades, but they have yet to implement automated scoring of these responses and instead rely on human scoring alone. To investigate how scores provided by machine algorithms compare to those provided by human raters, we applied convolutional neural networks (CNNs) to classify image-based responses from eight Timss 2019 items. Our results show that the most accurate CNN models classified over 99% of the image responses into the appropriate scoring category for dichotomous items and almost 98% for one trichotomous item. Additionally, during the modeling process, the CNNs correctly classified numerous image responses that human raters had scored incorrectly. For most items, the number of incorrectly human-scored responses exceeded the average number of responses misclassified by the most accurate models. These results suggest that automated scoring using CNNs is comparable to, and in many cases more accurate, than human raters, even across a wide variety of graphing tasks. This paper argues that the machine learning procedure explored could be implemented in ILSAs as a verification method to improve the accuracy and consistency of graphical response item scores. In lieu of additional human raters, ILSAs could implement CNN-based automated scoring to provide a second set of scores, thus reducing the workload and costs associated with human scoring.

Texto completo en abierto (CC BY-NC-ND 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • An interactive test dashboard with diagnosis and feedback mechanisms to facilitate learning performance
  • Can large language models meet the challenge of generating school-level questions?
  • Identifying active ingredients and uptake patterns in the implementation of an AI-based writing support tool: Insights from a randomized controlled trial
  • Comparing Teacher and Artificial Intelligence Scoring in Writing Assessment: A Generalizability Theory Analysis
  • Exploring AI-generated feedback in peer-discussion contexts: A mixed-methods study of essay writing in secondary classrooms
  • Assessing reading fluency in elementary grades: A machine learning approach
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos