IA: Bibliografía: IA y evaluación
Nº de publicaciones: 140Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial juegos evaluación
Temas IA: IA y evaluación estudio empírico
Resumen:
This study explores the use of serious games combined with machine learning techniques to evaluate student efficiency and determine appropriate degree programs. The main research question addressed is: How can predictive machine learning algorithms, applied to soft skills data collected through serious games, identify the most suitable study path for each student, thereby improving academic orientation and enhancing the likelihood of educational success? The research involved 211 university students from a single university in Italy, focusing on the application of predictive algorithms based on participants' soft skills and academic performance (GPA). The study employs logistic regression...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial integridad académica evaluación
Temas IA: IA y evaluación análisis de producción de IA estudio empírico
Resumen:
The growing use of generative AI (GenAI) tools like ChatGPT raises serious concerns about academic integrity, especially in the context of assessments. While detection efforts have focused on open-ended responses, multiple-choice questions (MCQs), which are common in high-stakes testing, remain largely overlooked, partly due to their perceived detection difficulty. The present work establishes a theoretical and empirical foundation for applications of psychometric theory to separate GenAI and human responses. Specifically, it investigates whether person-fit statistics (PFS), a class of methods within Item Response Theory (IRT) used to evaluate how well an examinee’s response pattern fits...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial evaluación grandes modelos de lenguaje
Temas IA: IA y evaluación modelos de lenguaje (LLM) estudio empírico
Resumen:
Integrating artificial intelligence (AI) into educational assessments represents a paradigm shift, especially in STEM subjects such as chemistry, which require complex problem-solving and written feedback. This study focuses on the effects of fine-tuning on each of the four AI models—Gemini 1.5, Gemini 1.0, BERT, and XLNet—and their performance on chemistry tasks. The models were implemented using three evaluation methods: Seq2Seq for Sodium Reaction Grading, a Regression Task for Overall Grading, and Attention-Based Grading for Key Steps Grading. The fine-tuning process significantly improved the models' accuracy, precision, and stability. Gemini 1.5 outperformed the other models across...
Detalles Cerrar β
Tipo: artículo
Metodología: revisión bibliográfica
Temas: inteligencia artificial evaluación procesamiento del lenguaje natural
Temas IA: IA y evaluación revisión de bibliografía
Resumen:
Text-based open-ended questions in academic formative and summative assessments help students become deep learners and prepare them to understand concepts for a subsequent conceptual assessment. However, grading text-based questions, especially in large (>50 enrolled students) courses, is tedious and time-consuming for instructors. Text processing models continue progressing with the rapid development of Artificial Intelligence (AI) tools and Natural Language Processing (NLP) algorithms. Especially after breakthroughs in Large Language Models (LLM), there is immense potential to automate rapid assessment and feedback of text-based responses in education. This systematic review adopts a...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial evaluación grandes modelos de lenguaje
Temas IA: IA y evaluación modelos de lenguaje (LLM) estudio empírico
Resumen:
Semi open-ended multipart questions consist of multiple sub questions within a single question, requiring students to provide certain factual information while allowing them to express their opinion within a defined context. Human grading of such questions can be tedious, constrained by the marking scheme and susceptible to the subjective judgement of instructors. The emergence of large language models (LLMs) such as ChatGPT has significantly advanced the prospect of automatic grading in educational settings. This paper introduces a topic-based grading approach that harnesses LLM capabilities alongside a refined marking scheme to ensure fair and explainable assessment processes. The...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial evaluación tecnología educativa
Temas IA: IA y evaluación IA y aprendizaje estudio empírico
Resumen:
Computerized adaptive testing (CAT) can effectively facilitate student assessment by dynamically selecting questions on the basis of learner knowledge and item difficulty. However, most CAT models are designed for one-time evaluation rather than improving learning through formative assessment. Since students cannot remember everything, encouraging them to repeatedly evaluate their knowledge state and identify their weaknesses is critical when developing an adaptive formative assessment system in real educational contexts. This study aims to achieve this goal by proposing an adaptive formative assessment system based on CAT and the learning memory cycle to enable the repeated evaluation of...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial expresión escrita evaluación
Temas IA: IA y enseñanza-aprendizaje de lenguas IA y evaluación modelos de lenguaje (LLM)
Resumen:
Recent advances in large language models have revitalized research on automated essay evaluation, yet critical concerns remain regarding their reliability, validity, and interpretability. This study presents a comparative analysis of five LLMs (GPT-4.1, Llama 4 Maverick, Gemini 2.5 Flash, Claude Sonnet 4, and DeepSeek R1) in the assessment of long English essays authored by non-native speakers in higher education. The analysis draws on LLM-generated scores for 60 essays to examine (a) intra-model reliability across repeated scoring runs, (b) the degree of alignment between model outputs and expert human ratings, and (c) causal feature dependencies that clarify how linguistic characteristics...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial evaluación expresión escrita
Temas IA: IA y evaluación estudio empírico
Resumen:
The use of standardized test formats in the assessment of historical competencies has recently come under severe criticism, especially in the United States, where standardized tests are particularly common. History researchers have argued that open-ended items are more appropriate for assessment. However, providing largescale evaluations of open-ended answers is time consuming and poses challenges regarding the objectivity, validity, and replicability of ratings. To address this issue, we investigated the extent to which computer-based evaluation methods are suitable for evaluating student answers by combining qualitative methods from history education research with quantitative, computer-...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial expresión escrita feedback/retroalimentación
Temas IA: IA y enseñanza-aprendizaje de lenguas IA y evaluación estudio empírico
Resumen:
This mixed-methods study investigates how an AI-powered writing tool providing automated feedback compares with peer-generated feedback in supporting secondary students' essay writing. Two research questions guided the study: (1) Did using an AI-powered tool with AI-generated feedback yield greater gains in writing quality from the first to the final draft compared with a standard editor with peer feedback? (2) How did students’ engagement in the writing process differ in the target (used AI-generated feedback) and comparison (used peer feedback) groups, as evidenced by teacher–student and peer–peer interactions and how were these patterns associated with their conceptual understanding of...
Detalles Cerrar β
Tipo: artículo
Metodología: estudio empírico
Temas: inteligencia artificial expresión escrita evaluación
Temas IA: IA y enseñanza-aprendizaje de lenguas IA y evaluación chatGPT
Resumen:
Researchers have sought for decades to automate holistic essay scoring. Over the years, these programs have improved significantly. However, accuracy requires significant amounts of training on human-scored texts—reducing the expediency and usefulness of such programs for routine uses by teachers across the nation on non-standardized prompts. This study analyzes the output of multiple versions of ChatGPT scoring of secondary student essays from three extant corpora and compares it to quality human ratings. We find that the current iteration of ChatGPT scoring is not statistically significantly different from human scoring; substantial agreement with humans is achievable and may be...
Etiquetas
- análisis de producción de IA (54)
- catálogo de herramientas (15)
- chatbots (68)
- chatGPT (155)
- conceptos básicos (38)
- estudio empírico (487)
- ética de la IA (119)
- guía (26)
- IA y aprendizaje (159)
- IA y creación de materiales (63)
- IA y educación (360)
- IA y ELE (31)
- IA y enseñanza-aprendizaje de lenguas (208)
- IA y evaluación (140)
- IA y variedades del español (1)
- modelos de lenguaje (LLM) (87)
- prompts (65)
- revisión de bibliografía (136)