Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Bibliografia
  • How do LLMs perform in the context of MCQs across different levels of thinking skills in a business education course at higher education? A comparison of ChatGPT, Gemini, and Copilot

Sección IA: Inteligencia artificial Bibliografía

How do LLMs perform in the context of MCQs across different levels of thinking skills in a business education course at higher education? A comparison of ChatGPT, Gemini, and Copilot

Laurens Goorts
Ryan Hollevoet
Vanessa Xia
Felix Cammaerts
Almer Güngör
2025
Computers & Education: Artificial Intelligence
9
https://www.sciencedirect.com/science/a…
artículo
estudio empírico
inteligencia artificial
grandes modelos de lenguaje
ChatGPT
evaluación
educación superior
herramientas
tecnología educativa
análisis de producción de IA
modelos de lenguaje (LLM)
chatGPT
estudio empírico

Texto completo

This exploratory study investigates the performance of three widely-used and freely available large language models (LLMs)— ChatGPT (GPT-3.5 Turbo), Gemini, and Copilot—in answering multiple-choice questions (MCQs) categorized by cognitive complexity based on the revised Bloom's Taxonomy. Although MCQs offer a structured, efficient, and scalable method for evaluation, a gap exists in the literature on how LLMs handle varying cognitive levels, particularly in a business course at higher education. Understanding LLM performance in this context is crucial, as students increasingly use LLMs for searching answers but also to receive tailored and scaffolded responses for an interactive and personalized learning experience. Using 100 MCQs on the Business Intelligence & Data Analytics chapter from this course—classified into Lower-Order Thinking Skills (LOTS) and Higher-Order Thinking Skills (HOTS)—LLMs' accuracy scores were compared across these categories as well as Bloom's Subcategories. Findings from Generalized Linear Mixed Models reveal that three LLMs perform better on LOTS questions than HOTS. While descriptive trends in performance were observed across models, these differences were not statistically significant. However, prompt engineering, specifically FS-CD, can enhance LLM performance on complex tasks, and its effectiveness varies by model and cognitive level, significantly improving performance on HOTS for all LLMs, particularly for Apply-level tasks. These insights have potential implications for higher education, especially in business education. These findings underscore the necessity for a critical approach when using LLMs as learning tools. Educators can leverage these insights to guide students in the effective use of LLMs for mastering complex concepts and skills.

Texto completo en abierto (CC BY-NC-ND 4.0).
  • Inicie sesión para enviar comentarios

Enviar publicación

Contenidos relacionados

  • Performance of ChatGPT on the US fundamentals of engineering exam: Comprehensive assessment of proficiency and potential implications for professional environmental engineering practice
  • ¿Tienen GPT-3.5 y GPT-4 un estilo de escritura diferente del estilo humano?: un estudio exploratorio para el español
  • Fine-tuning ChatGPT for automatic scoring
  • Students’ use of large language models in engineering education: A case study on technology acceptance, perceptions, efficacy, and detection chances
  • Comparing expert tutor evaluation of reflective essays with marking by generative artificial intelligence (AI) tool
  • Analysis of LLMs for educational question classification and generation
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos