Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Applying psychometric methods to distinguish between human and generative AI responses to multiple-choice assessments

Sección IA: Inteligencia artificial Bibliografía

Applying psychometric methods to distinguish between human and generative AI responses to multiple-choice assessments

Alona Strugatski
Giora Alexandron
2026
Computers & Education: Artificial Intelligence
11
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
integridad académica
evaluación
IA y evaluación
análisis de producción de IA
estudio empírico

Texto completo

The growing use of generative AI (GenAI) tools like ChatGPT raises serious concerns about academic integrity, especially in the context of assessments. While detection efforts have focused on open-ended responses, multiple-choice questions (MCQs), which are common in high-stakes testing, remain largely overlooked, partly due to their perceived detection difficulty. The present work establishes a theoretical and empirical foundation for applications of psychometric theory to separate GenAI and human responses. Specifically, it investigates whether person-fit statistics (PFS), a class of methods within Item Response Theory (IRT) used to evaluate how well an examinee’s response pattern fits the expectations of the IRT model, can distinguish between GenAI and human responses to MCQ assessments. To study this, we use data from two authentic assessment contexts: a high-school level chemistry test and the national university entrance exam, each with approximately 1000 human respondents. Our results demonstrate that PFS reveal significant differences between human responses and those generated by advanced chatbots (ChatGPT, Claude, and Gemini), which appear as ‘aberrant’ respondents. We also demonstrate that different chatbots present significantly different response patterns, suggesting that they should be treated as a heterogeneous group of ‘intelligences’ rather than as a single one. Using the PFS measures, we also demonstrate, however, that newer GenAI versions not only improve in performance but also become more ‘human-like’ in their response patterns. Together, these results position IRT as a robust framework for characterizing and separating human and GenAI response patterns in MCQ assessments, providing a theoretical foundation and empirical evidence.

Texto completo en abierto (CC BY-NC-ND 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • GenAI detection tools, adversarial techniques and implications for inclusivity in higher education
  • Testing of detection tools for AI-generated text
  • Do teachers spot AI? Evaluating the detectability of AI-generated texts among student essays
  • GPT detectors are biased against non-native English writers
  • Accused: How students respond to allegations of using ChatGPT on assessments
  • The AI Assessment Scale (AIAS) in action: A pilot implementation of GenAI-supported assessment
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos