Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Inteligencia artificial
  • Bibliografía
  • Etiquetas
  • IA y evaluación

IA: Bibliografía: IA y evaluación

Nº de publicaciones: 140
Volver a lista completa

Paginación

  • Primera página «
  • Página anterior ‹
  • …
  • Página 6
  • Página 7
  • Página 8
  • Página 9
  • Página 10
  • Página actual 11
  • Página 12
  • Página 13
  • Página 14
  • Siguiente página ›
  • Última página »
A machine learning framework for soft skills assessment: Leveraging serious games in higher education
Agostino Marengo, Alessandro Pagano, Vito Santamato (2025)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial juegos evaluación

Temas IA: IA y evaluación estudio empírico

Resumen:

Texto completo

This study explores the use of serious games combined with machine learning techniques to evaluate student efficiency and determine appropriate degree programs. The main research question addressed is: How can predictive machine learning algorithms, applied to soft skills data collected through serious games, identify the most suitable study path for each student, thereby improving academic orientation and enhancing the likelihood of educational success? The research involved 211 university students from a single university in Italy, focusing on the application of predictive algorithms based on participants' soft skills and academic performance (GPA). The study employs logistic regression...

Ver ficha completa

Applying psychometric methods to distinguish between human and generative AI responses to multiple-choice assessments
Alona Strugatski, Giora Alexandron (2026)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial integridad académica evaluación

Temas IA: IA y evaluación análisis de producción de IA estudio empírico

Resumen:

Texto completo

The growing use of generative AI (GenAI) tools like ChatGPT raises serious concerns about academic integrity, especially in the context of assessments. While detection efforts have focused on open-ended responses, multiple-choice questions (MCQs), which are common in high-stakes testing, remain largely overlooked, partly due to their perceived detection difficulty. The present work establishes a theoretical and empirical foundation for applications of psychometric theory to separate GenAI and human responses. Specifically, it investigates whether person-fit statistics (PFS), a class of methods within Item Response Theory (IRT) used to evaluate how well an examinee’s response pattern fits...

Ver ficha completa

Fine-tuning AI models for enhanced consistency and precision in chemistry educational assessments
Sri Yamtinah, Antuni Wiyarsi, Hayuni Retno Widarti (2025)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial evaluación grandes modelos de lenguaje

Temas IA: IA y evaluación modelos de lenguaje (LLM) estudio empírico

Resumen:

Texto completo

Integrating artificial intelligence (AI) into educational assessments represents a paradigm shift, especially in STEM subjects such as chemistry, which require complex problem-solving and written feedback. This study focuses on the effects of fine-tuning on each of the four AI models—Gemini 1.5, Gemini 1.0, BERT, and XLNet—and their performance on chemistry tasks. The models were implemented using three evaluation methods: Seq2Seq for Sodium Reaction Grading, a Regression Task for Overall Grading, and Attention-Based Grading for Key Steps Grading. The fine-tuning process significantly improved the models' accuracy, precision, and stability. Gemini 1.5 outperformed the other models across...

Ver ficha completa

Automatic assessment of text-based responses in post-secondary education: A systematic review
Rujun Gao, Hillary E. Merzdorf, Saira Anwar (2024)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: revisión bibliográfica

Temas: inteligencia artificial evaluación procesamiento del lenguaje natural

Temas IA: IA y evaluación revisión de bibliografía

Resumen:

Texto completo

Text-based open-ended questions in academic formative and summative assessments help students become deep learners and prepare them to understand concepts for a subsequent conceptual assessment. However, grading text-based questions, especially in large (>50 enrolled students) courses, is tedious and time-consuming for instructors. Text processing models continue progressing with the rapid development of Artificial Intelligence (AI) tools and Natural Language Processing (NLP) algorithms. Especially after breakthroughs in Large Language Models (LLM), there is immense potential to automate rapid assessment and feedback of text-based responses in education. This systematic review adopts a...

Ver ficha completa

Harnessing the power of AI-instructor collaborative grading approach: Topic-based effective grading for semi open-ended multipart questions
Phyo Yi Win Myint, Siaw Ling Lo, Yuhao Zhang (2024)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial evaluación grandes modelos de lenguaje

Temas IA: IA y evaluación modelos de lenguaje (LLM) estudio empírico

Resumen:

Texto completo

Semi open-ended multipart questions consist of multiple sub questions within a single question, requiring students to provide certain factual information while allowing them to express their opinion within a defined context. Human grading of such questions can be tedious, constrained by the marking scheme and susceptible to the subjective judgement of instructors. The emergence of large language models (LLMs) such as ChatGPT has significantly advanced the prospect of automatic grading in educational settings. This paper introduces a topic-based grading approach that harnesses LLM capabilities alongside a refined marking scheme to ensure fair and explainable assessment processes. The...

Ver ficha completa

Adaptive formative assessment system based on computerized adaptive testing and the learning memory cycle for personalized learning
Albert C.M. Yang, Brendan Flanagan, Hiroaki Ogata (2022)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial evaluación tecnología educativa

Temas IA: IA y evaluación IA y aprendizaje estudio empírico

Resumen:

Texto completo

Computerized adaptive testing (CAT) can effectively facilitate student assessment by dynamically selecting questions on the basis of learner knowledge and item difficulty. However, most CAT models are designed for one-time evaluation rather than improving learning through formative assessment. Since students cannot remember everything, encouraging them to repeatedly evaluate their knowledge state and identify their weaknesses is critical when developing an adaptive formative assessment system in real educational contexts. This study aims to achieve this goal by proposing an adaptive formative assessment system based on CAT and the learning memory cycle to enable the repeated evaluation of...

Ver ficha completa

A framework for evaluation of large language models in essay assessment: Reliability, alignment, and causal reasoning
Tongxi Liu, Luyao Ye, Wei Yan (2026)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial expresión escrita evaluación

Temas IA: IA y enseñanza-aprendizaje de lenguas IA y evaluación modelos de lenguaje (LLM)

Resumen:

Texto completo

Recent advances in large language models have revitalized research on automated essay evaluation, yet critical concerns remain regarding their reliability, validity, and interpretability. This study presents a comparative analysis of five LLMs (GPT-4.1, Llama 4 Maverick, Gemini 2.5 Flash, Claude Sonnet 4, and DeepSeek R1) in the assessment of long English essays authored by non-native speakers in higher education. The analysis draws on LLM-generated scores for 60 essays to examine (a) intra-model reliability across repeated scoring runs, (b) the degree of alignment between model outputs and expert human ratings, and (c) causal feature dependencies that clarify how linguistic characteristics...

Ver ficha completa

Artificial intelligence in history education. Linguistic content and complexity analyses of student writings in the CAHisT project (Computational assessment of historical thinking)
Christiane Bertram, Zarah Weiss, Lisa Zachrich (2021)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial evaluación expresión escrita

Temas IA: IA y evaluación estudio empírico

Resumen:

Texto completo

The use of standardized test formats in the assessment of historical competencies has recently come under severe criticism, especially in the United States, where standardized tests are particularly common. History researchers have argued that open-ended items are more appropriate for assessment. However, providing largescale evaluations of open-ended answers is time consuming and poses challenges regarding the objectivity, validity, and replicability of ratings. To address this issue, we investigated the extent to which computer-based evaluation methods are suitable for evaluating student answers by combining qualitative methods from history education research with quantitative, computer-...

Ver ficha completa

Exploring AI-generated feedback in peer-discussion contexts: A mixed-methods study of essay writing in secondary classrooms
Irina Engeness (2025)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial expresión escrita feedback/retroalimentación

Temas IA: IA y enseñanza-aprendizaje de lenguas IA y evaluación estudio empírico

Resumen:

Texto completo

This mixed-methods study investigates how an AI-powered writing tool providing automated feedback compares with peer-generated feedback in supporting secondary students' essay writing. Two research questions guided the study: (1) Did using an AI-powered tool with AI-generated feedback yield greater gains in writing quality from the first to the final draft compared with a standard editor with peer feedback? (2) How did students’ engagement in the writing process differ in the target (used AI-generated feedback) and comparison (used peer feedback) groups, as evidenced by teacher–student and peer–peer interactions and how were these patterns associated with their conceptual understanding of...

Ver ficha completa

Can AI provide useful holistic essay scoring?
Tamara P. Tate, Jacob Steiss, Drew Bailey (2024)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial expresión escrita evaluación

Temas IA: IA y enseñanza-aprendizaje de lenguas IA y evaluación chatGPT

Resumen:

Texto completo

Researchers have sought for decades to automate holistic essay scoring. Over the years, these programs have improved significantly. However, accuracy requires significant amounts of training on human-scored texts—reducing the expediency and usefulness of such programs for routine uses by teachers across the nation on non-standardized prompts. This study analyzes the output of multiple versions of ChatGPT scoring of secondary student essays from three extant corpora and compares it to quality human ratings. We find that the current iteration of ChatGPT scoring is not statistically significantly different from human scoring; substantial agreement with humans is achievable and may be...

Ver ficha completa

Paginación

  • Primera página «
  • Página anterior ‹
  • …
  • Página 6
  • Página 7
  • Página 8
  • Página 9
  • Página 10
  • Página actual 11
  • Página 12
  • Página 13
  • Página 14
  • Siguiente página ›
  • Última página »

Etiquetas

  • análisis de producción de IA (54)
  • catálogo de herramientas (15)
  • chatbots (68)
  • chatGPT (155)
  • conceptos básicos (38)
  • estudio empírico (487)
  • ética de la IA (119)
  • guía (26)
  • IA y aprendizaje (159)
  • IA y creación de materiales (63)
  • IA y educación (360)
  • IA y ELE (31)
  • IA y enseñanza-aprendizaje de lenguas (208)
  • IA y evaluación (140)
  • IA y variedades del español (1)
  • modelos de lenguaje (LLM) (87)
  • prompts (65)
  • revisión de bibliografía (136)
Sobre Todoele Índice Publica Contacto: [email protected]
Política de privacidad Créditos