Inicio
todoELE
  • Inicio
  • Materiales
    • ๐Ÿ“‹ Actividades
    • ๐Ÿ“ Conjugación
    • ๐Ÿ“Š Corpus
    • ๐Ÿ“” Diccionarios
    • โœ… Evaluación
    • โš™๏ธ Gramática
    • ๐Ÿ“— Manuales
    • โœ๏ธ Ortografía
    • ๐Ÿ“… Programación
    • ๐Ÿ—ฃ๏ธ Pronunciación
    • ๐Ÿ“ Recursos
    • ๐Ÿ”ค Vocabulario
    • ๐Ÿ’ป Herramientas digitales
  • Formación
    • ๐Ÿ“š Bibliografía
    • ๐Ÿ‘ฅ Congresos
    • ๐ŸŽ“ Cursos
    • ๐Ÿซ Centros
    • ๐Ÿข Organizaciones
    • ๐Ÿ“ฐ Revistas
    • ๐ŸŒ Atlas de ELE
  • Trabajo
    • ๐Ÿ’ผ Ofertas de trabajo
    • โ„น๏ธ Trabajo - Recursos
  • En la red
    • ๐ŸŒ Sitios ELE
    • ๐Ÿ“ฐ Agregador
    • ๐Ÿ“ง Formespa
  • IA
    • โœจ Nuevos contenidos
    • ๐Ÿ“š Bibliografía IA
    • ๐Ÿงฐ Herramientas IA
    • ๐Ÿ’ฌ Prompts
    • ๐Ÿงช Experiencias IA
    • ๐ŸŒ Sitios web IA
    • ๐Ÿ“ฐ Actualidad IA
  • Comunidad
    • ๐Ÿ“ฐ Actualidad ELE
    • ๐Ÿ˜Š Anécdotas ELE
    • ๐Ÿ“ Blog
    • ๐Ÿ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • Comparative analysis of NLP-driven MCQ generators from text sources

Sección IA: Inteligencia artificial Bibliografía

Comparative analysis of NLP-driven MCQ generators from text sources

Asmae Azzi
Ferenc Erdล‘s
Richárd Németh
Vijayakumar Varadarajan
Stephen Afrifa
2025
Computers & Education: Artificial Intelligence
9
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
creación de materiales
evaluación
grandes modelos de lenguaje
procesamiento del lenguaje natural
IA y creación de materiales
IA y evaluación
modelos de lenguaje (LLM)
estudio empírico

Texto completo

The application of learning sciences with technology has been shown to boost learner interactions, yet the potential of advanced tool, particularly those that leverage Natural Language Processing (NLP), still very much untapped in learning contexts. This paper speaks to this age-old problem of generating quality Multiple-Choice questions (MCQs) – a prevalent but time-consuming mode of assessment – via the suggested comprehensive comparison study of template-based AI solutions. The study contrasts general-purpose Large Language Models (LLMs) with specialized MCQ-focused AI programs. The scientific approach employed was quite stringent, where each of the software applications was benchmarked using a common dataset of text across varying levels of complexity and topic. Results indicate that general-purpose LLMs, especially DeepSeek and ChatGPT, consistently present higher performance and reliability, especially when processing complex textual content. Whereas specialized tools offer distinctive formatting options, they exhibit decreasing performance as texts become more complex and signify strong operation constriction at the free versions. Developing solid and effective distractors turned out to be a complicated task for all the tested tools. We conclude the paper by presenting a standardized assessment model, making evidence-based recommendations for developers and teachers, and suggesting ways to incorporate various AI capabilities into modern educational assessment effectively.

Texto completo en abierto (CC BY 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Can large language models meet the challenge of generating school-level questions?
  • Automatic item generation in various STEM subjects using large language model prompting
  • Automated reading passage generation with OpenAI's large language model
  • Automatic question-answer pairs generation using pre-trained large language models in higher education
  • Analysis of LLMs for educational question classification and generation
  • Investigating the affordances of OpenAI's large language model in developing listening assessments
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos