IA: Bibliografía: IA y evaluación
Nº de publicaciones: 140Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial evaluación grandes modelos de lenguaje
Temas IA: IA y evaluación modelos de lenguaje (LLM) prompts
Resumen:
Automated scoring (AS) has become increasingly prevalent in educational measurement. However, applying it to international reading assessments remains challenging, particularly due to the length and complexity of the required prompting, driven by the need to include lengthy reading passages and detailed scoring guides. Processing these lengthy inputs results in high computational costs and may impede the performance of large language models (LLMs). This study explored the potential of optimizing AS with prompt compression using OpenAI's LLM, GPT-4o. Our results show that prompt compression significantly reduces the length of reading passages and scoring guides while maintaining their...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial feedback/retroalimentación educación superior
Temas IA: IA y evaluación estudio empírico
Resumen:
In higher education, delivering effective feedback is pivotal for enhancing student learning but remains challenging due to the scale and diversity of student populations. Learner-centered feedback, a robust approach to effective feedback that tailors to individual student needs, encompasses three key dimensions—Future Impact, Sensemaking, and Agency, which collectively include eight specific components, thereby enhancing its relevance and impact in the learning process. However, providing consistent and effective learner-centered feedback at scale is challenging for educators. This study addresses this challenge by automating the analysis of feedback content to promote effective learner-...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial grandes modelos de lenguaje evaluación
Temas IA: análisis de producción de IA modelos de lenguaje (LLM) IA y evaluación
Resumen:
Open-ended assessments play a pivotal role in enabling instructors to evaluate student knowledge acquisition and provide constructive feedback. Integrating large language models (LLMs) such as GPT-4 in educational settings presents a transformative opportunity for assessment methodologies. However, existing literature on LLMs addressing open-ended questions lacks breadth, relying on limited data or overlooking question difficulty levels. This study evaluates GPT-4's proficiency in responding to open-ended questions spanning diverse topics and cognitive complexities in comparison to human responses. To facilitate this assessment, we generated a dataset of 738 open-ended questions across...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial gamificación pensamiento computacional
Temas IA: IA y evaluación estudio empírico
Resumen:
The availability of Automated Assessment tools for computer programming tasks can be a significant asset in Computer Science education. Systems providing such kind of service are built around an interface, allowing to administer the tasks (exercises to train programming skills), and show the results, accompanied by meaningful feedback. To produce such results, they apply techniques ranging from static analysis of program correctness, to testing-based evaluation. These systems can also support Competitive Programming, which is known to have educational meaning too. We developed the 2TSW system, supporting the automated correction of computer programming tasks, in a gamified web-based...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial ChatGPT expresión escrita
Temas IA: IA y enseñanza-aprendizaje de lenguas chatGPT análisis de producción de IA
Resumen:
Reformulation is a form of written corrective feedback to help second language (L2) learners improve their writing. This study examined whether ChatGPT could produce reformulations that (1) retain the meanings of the original essays and (2) are linguistically more developed than learners’ original essays. In addition, three types of ChatGPT prompts were compared to see which type yielded better reformulations. One thousand two hundred argumentative essays written for the TOEFL iBT® independent writing task were submitted to ChatGPT. ROUGE-L scores, used as a proxy for meaning retention, showed that ChatGPT reformulations largely retained the meaning of the original essays. A qualitative...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial feedback/retroalimentación grandes modelos de lenguaje
Temas IA: modelos de lenguaje (LLM) prompts IA y evaluación
Resumen:
Relational feedback is increasingly recognised for its crucial role in enhancing student-instructor relationships and promoting the assimilation of feedback. Despite its significance, no studies have tried to develop automated methods to analyse written feedback for properties of relational feedback to promote its use at scale and assist feedback providers with their relational feedback practices. This automated analysis of relational feedback can be performed as a classification task. However, traditional machine and deep learning methods for text classification typically require extensive human labelling and pose a significant challenge for educators and researchers lacking machine...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial evaluación grandes modelos de lenguaje
Temas IA: IA y evaluación modelos de lenguaje (LLM) estudio empírico
Resumen:
Currently, many generative Artificial Intelligence (AI) tools are being integrated into the educational technology landscape for instructors. Our paper examines the potential and challenges of using Large Language Models (LLMs) to code student-generated content in online discussions based on intended learning outcomes and how instructors could use this to assess the intended and enacted learning design. If instructors were to rely on LLMs as a means of assessment, the reliability of these models to code the data accurately is crucial. Employing a diverse set of LLMs from the GPT family and prompting techniques on an asynchronous online discussion dataset from a blended-learning bachelor-...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial evaluación grandes modelos de lenguaje
Temas IA: IA y evaluación modelos de lenguaje (LLM) estudio empírico
Resumen:
Identifying logical errors in complex, incomplete or even contradictory and overall heterogeneous data like students’ experimentation protocols is challenging. Recognizing the limitations of current evaluation methods, we investigate the potential of Large Language Models (LLMs) for automatically identifying student errors and streamlining teacher assessments. Our aim is to provide a foundation for productive, personalized feedback. Using a dataset of 65 student protocols, an Artificial Intelligence (AI) system based on the GPT-3.5 and GPT-4 series was developed and tested against human raters. Our results indicate varying levels of accuracy in error detection between the AI system and...
Detalles Cerrar β
Tipo: artículo
Metodología: trabajo teórico
Temas: inteligencia artificial integridad académica evaluación
Temas IA: IA y evaluación ética de la IA
Resumen:
We consider how a future of pervasive human/AI coauthorship challenges current notions of validity and academic integrity. Specifically, we address widespread concerns that in non-proctored contexts, students are using generative artificial intelligence (GenAI) to submit texts they do not understand. Adopting an assessment validity lens, we show how GenAI undermines the integrity of multiple forms of validity evidence, leading us to propose Coauthorship Integrity as a new conceptual source of validity evidence for addressing these threats. Coauthorship Integrity is violated when students submit AI-generated content that they do not understand. To hold students accountable in this regard, a...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial juegos evaluación
Temas IA: IA y evaluación estudio empírico
Resumen:
Highlights:
- The analysis of gameplay data from the Blood Transfusion Serious Game (BTSG) revealed crucial insights into the engagement, performance, and competency levels of participating nurses.
- The study highlighted the efficiency of the adaptive assessment algorithm in determining competency indicators through gameplay activities.
- Notable discrepancies existed between the game-embedded result and the adaptive assessment result, suggesting potential biases in the game result, which tended to overestimate the demonstrated competence.
Etiquetas
- análisis de producción de IA (54)
- catálogo de herramientas (15)
- chatbots (68)
- chatGPT (155)
- conceptos básicos (38)
- estudio empírico (487)
- ética de la IA (119)
- guía (26)
- IA y aprendizaje (159)
- IA y creación de materiales (63)
- IA y educación (360)
- IA y ELE (31)
- IA y enseñanza-aprendizaje de lenguas (208)
- IA y evaluación (140)
- IA y variedades del español (1)
- modelos de lenguaje (LLM) (87)
- prompts (65)
- revisión de bibliografía (136)