Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Inteligencia artificial
  • Bibliografía
  • Etiquetas
  • IA y evaluación

IA: Bibliografía: IA y evaluación

Nº de publicaciones: 140
Volver a lista completa

Paginación

  • Primera página «
  • Página anterior ‹
  • …
  • Página 6
  • Página 7
  • Página 8
  • Página 9
  • Página 10
  • Página 11
  • Página actual 12
  • Página 13
  • Página 14
  • Siguiente página ›
  • Última página »
Assessing the quality of automatic-generated short answers using GPT-4
Luiz Rodrigues, Filipe Dwan Pereira, Luciano Cabral (2024)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial grandes modelos de lenguaje evaluación

Temas IA: análisis de producción de IA modelos de lenguaje (LLM) IA y evaluación

Resumen:

Texto completo

Open-ended assessments play a pivotal role in enabling instructors to evaluate student knowledge acquisition and provide constructive feedback. Integrating large language models (LLMs) such as GPT-4 in educational settings presents a transformative opportunity for assessment methodologies. However, existing literature on LLMs addressing open-ended questions lacks breadth, relying on limited data or overlooking question difficulty levels. This study evaluates GPT-4's proficiency in responding to open-ended questions spanning diverse topics and cognitive complexities in comparison to human responses. To facilitate this assessment, we generated a dataset of 738 open-ended questions across...

Ver ficha completa

A gamified web based system for computer programming learning
Giuseppina Polito, Marco Temperini (2021)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial gamificación pensamiento computacional

Temas IA: IA y evaluación estudio empírico

Resumen:

Texto completo

The availability of Automated Assessment tools for computer programming tasks can be a significant asset in Computer Science education. Systems providing such kind of service are built around an interface, allowing to administer the tasks (exercises to train programming skills), and show the results, accompanied by meaningful feedback. To produce such results, they apply techniques ranging from static analysis of program correctness, to testing-based evaluation. These systems can also support Competitive Programming, which is known to have educational meaning too. We developed the 2TSW system, supporting the automated correction of computer programming tasks, in a gamified web-based...

Ver ficha completa

Evaluating the potential of ChatGPT-reformulated essays as written feedback in L2 writing
Yingzhao Chen (2025)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial ChatGPT expresión escrita

Temas IA: IA y enseñanza-aprendizaje de lenguas chatGPT análisis de producción de IA

Resumen:

Texto completo

Reformulation is a form of written corrective feedback to help second language (L2) learners improve their writing. This study examined whether ChatGPT could produce reformulations that (1) retain the meanings of the original essays and (2) are linguistically more developed than learners’ original essays. In addition, three types of ChatGPT prompts were compared to see which type yielded better reformulations. One thousand two hundred argumentative essays written for the TOEFL iBT® independent writing task were submitted to ChatGPT. ROUGE-L scores, used as a proxy for meaning retention, showed that ChatGPT reformulations largely retained the meaning of the original essays. A qualitative...

Ver ficha completa

Evaluating the capability of large language models in characterising relational feedback: A comparative analysis of prompting strategies
Wei Dai, Yixin Cheng, Ahmad Ari Aldino (2025)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial feedback/retroalimentación grandes modelos de lenguaje

Temas IA: modelos de lenguaje (LLM) prompts IA y evaluación

Resumen:

Texto completo

Relational feedback is increasingly recognised for its crucial role in enhancing student-instructor relationships and promoting the assimilation of feedback. Despite its significance, no studies have tried to develop automated methods to analyse written feedback for properties of relational feedback to promote its use at scale and assist feedback providers with their relational feedback practices. This automated analysis of relational feedback can be performed as a classification task. However, traditional machine and deep learning methods for text classification typically require extensive human labelling and pose a significant challenge for educators and researchers lacking machine...

Ver ficha completa

Augmenting assessment with AI coding of online student discourse: A question of reliability
Kamila Misiejuk, Rogers Kaliisa, Jennifer Scianna (2024)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial evaluación grandes modelos de lenguaje

Temas IA: IA y evaluación modelos de lenguaje (LLM) estudio empírico

Resumen:

Texto completo

Currently, many generative Artificial Intelligence (AI) tools are being integrated into the educational technology landscape for instructors. Our paper examines the potential and challenges of using Large Language Models (LLMs) to code student-generated content in online discussions based on intended learning outcomes and how instructors could use this to assess the intended and enacted learning design. If instructors were to rely on LLMs as a means of assessment, the reliability of these models to code the data accurately is crucial. Employing a diverse set of LLMs from the GPT family and prompting techniques on an asynchronous online discussion dataset from a blended-learning bachelor-...

Ver ficha completa

Assessing student errors in experimentation using artificial intelligence and large language models: A comparative study with human raters
Arne Bewersdorff, Kathrin Seßler, Armin Baur (2023)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial evaluación grandes modelos de lenguaje

Temas IA: IA y evaluación modelos de lenguaje (LLM) estudio empírico

Resumen:

Texto completo

Identifying logical errors in complex, incomplete or even contradictory and overall heterogeneous data like students’ experimentation protocols is challenging. Recognizing the limitations of current evaluation methods, we investigate the potential of Large Language Models (LLMs) for automatically identifying student errors and streamlining teacher assessments. Our aim is to provide a foundation for productive, personalized feedback. Using a dataset of 65 student protocols, an Artificial Intelligence (AI) system based on the GPT-3.5 and GPT-4 series was developed and tested against human raters. Our results indicate varying levels of accuracy in error detection between the AI system and...

Ver ficha completa

Coauthorship integrity: Reconceptualising assessment validity for the age of generative artificial intelligence
Mohsen Ebrahimzadeh, Antonette Shibani, Simon Buckingham Shum (2026)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: trabajo teórico

Temas: inteligencia artificial integridad académica evaluación

Temas IA: IA y evaluación ética de la IA

Resumen:

Texto completo

We consider how a future of pervasive human/AI coauthorship challenges current notions of validity and academic integrity. Specifically, we address widespread concerns that in non-proctored contexts, students are using generative artificial intelligence (GenAI) to submit texts they do not understand. Adopting an assessment validity lens, we show how GenAI undermines the integrity of multiple forms of validity evidence, leading us to propose Coauthorship Integrity as a new conceptual source of validity evidence for addressing these threats. Coauthorship Integrity is violated when students submit AI-generated content that they do not understand. To hold students accountable in this regard, a...

Ver ficha completa

Adaptive serious games assessment: The case of the blood transfusion game in nursing education
Dirk Ifenthaler, Muhittin ŞahiΜ‡n, Ivan Boo (2025)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial juegos evaluación

Temas IA: IA y evaluación estudio empírico

Resumen:

Texto completo

Highlights:

  • The analysis of gameplay data from the Blood Transfusion Serious Game (BTSG) revealed crucial insights into the engagement, performance, and competency levels of participating nurses.
  • The study highlighted the efficiency of the adaptive assessment algorithm in determining competency indicators through gameplay activities.
  • Notable discrepancies existed between the game-embedded result and the adaptive assessment result, suggesting potential biases in the game result, which tended to overestimate the demonstrated competence.

Ver ficha completa

Can students judge like experts? A large-scale study on the pedagogical quality of AI and human personalized formative feedback
Tanya Nazaretsky, Hagit Gabbay, Tanja Käser (2026)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial feedback/retroalimentación educación superior

Temas IA: IA y evaluación estudio empírico

Resumen:

Texto completo

While feedback is essential for guiding student learning, providing timely and personalized guidance in large-scale educational settings remains a significant challenge. Generative AI offers a scalable solution, yet little is known about students’ perceptions of AI-generated feedback. In this paper, we aim to investigate how the identity of the feedback provider (human vs. AI) affects students’ ability to assess feedback quality and whether their judgments are biased. We propose a comprehensive rubric for assessing the pedagogical quality of formative feedback. We use it to compare the objective quality of AI-generated and human-crafted feedback (N = 979). Next, using data collected from...

Ver ficha completa

Evaluating the psychometric properties of ChatGPT-generated questions
Shreya Bhandari, Yunting Liu, Yerin Kwak (2024)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial evaluación ChatGPT

Temas IA: IA y evaluación IA y creación de materiales análisis de producción de IA

Resumen:

Texto completo

Not much is known about how LLM-generated questions compare to gold-standard, traditional formative assessments concerning their difficulty and discrimination parameters, which are valued properties in the psychometric measurement field. We follow a rigorous measurement methodology to compare a set of ChatGPT-generated questions, produced from one lesson summary in a textbook, to existing questions from a published Creative Commons textbook. To do this, we collected and analyzed responses from 207 test respondents who answered questions from both item pools and used a linking methodology to compare IRT properties between the two pools. We find that neither the difficulty nor discrimination...

Ver ficha completa

Paginación

  • Primera página «
  • Página anterior ‹
  • …
  • Página 6
  • Página 7
  • Página 8
  • Página 9
  • Página 10
  • Página 11
  • Página actual 12
  • Página 13
  • Página 14
  • Siguiente página ›
  • Última página »

Etiquetas

  • análisis de producción de IA (54)
  • catálogo de herramientas (15)
  • chatbots (68)
  • chatGPT (155)
  • conceptos básicos (38)
  • estudio empírico (487)
  • ética de la IA (119)
  • guía (26)
  • IA y aprendizaje (159)
  • IA y creación de materiales (63)
  • IA y educación (360)
  • IA y ELE (31)
  • IA y enseñanza-aprendizaje de lenguas (208)
  • IA y evaluación (140)
  • IA y variedades del español (1)
  • modelos de lenguaje (LLM) (87)
  • prompts (65)
  • revisión de bibliografía (136)
Sobre Todoele Índice Publica Contacto: [email protected]
Política de privacidad Créditos