Inicio
todoELE
  • Inicio
  • Materiales
    • πŸ“‹ Actividades
    • πŸ“ Conjugación
    • πŸ“Š Corpus
    • πŸ“” Diccionarios
    • βœ… Evaluación
    • βš™οΈ Gramática
    • πŸ“— Manuales
    • ✍️ Ortografía
    • πŸ“… Programación
    • πŸ—£οΈ Pronunciación
    • πŸ“ Recursos
    • πŸ”€ Vocabulario
    • πŸ’» Herramientas digitales
  • Formación
    • πŸ“š Bibliografía
    • πŸ‘₯ Congresos
    • πŸŽ“ Cursos
    • 🏫 Centros
    • 🏒 Organizaciones
    • πŸ“° Revistas
    • 🌍 Atlas de ELE
  • Trabajo
    • πŸ’Ό Ofertas de trabajo
    • ℹ️ Trabajo - Recursos
  • En la red
    • 🌐 Sitios ELE
    • πŸ“° Agregador
    • πŸ“§ Formespa
  • IA
    • ✨ Nuevos contenidos
    • πŸ“š Bibliografía IA
    • 🧰 Herramientas IA
    • πŸ’¬ Prompts
    • πŸ§ͺ Experiencias IA
    • 🌐 Sitios web IA
    • πŸ“° Actualidad IA
  • Comunidad
    • πŸ“° Actualidad ELE
    • 😊 Anécdotas ELE
    • πŸ“ Blog
    • πŸ“ŒTablón de anuncios
  • Buscar

Ruta de navegación

  • Inicio
  • Inteligencia artificial
  • Bibliografía
  • Etiquetas
  • análisis de producción de IA

IA: Bibliografía: análisis de producción de IA

Nº de publicaciones: 54
Volver a lista completa

Paginación

  • Primera página «
  • Página anterior ‹
  • Página 1
  • Página 2
  • Página 3
  • Página actual 4
  • Página 5
  • Página 6
  • Siguiente página ›
  • Última página »
LLMs do not grade essays like humans
Jerin George Mathew, Sumayya Taher, Anindita Kundu (2026)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial expresión escrita evaluación

Temas IA: IA y enseñanza-aprendizaje de lenguas IA y evaluación modelos de lenguaje (LLM)

Resumen:

Texto completo

Large language models have recently been proposed as tools for automated essay scoring, but their agreement with human grading remains unclear. In this work, we evaluate how LLM-generated scores compare with human grades and analyze the grading behavior of several models from the GPT and Llama families in an out-of-the-box setting, without task-specific training. Our results show that agreement between LLM and human scores remains relatively weak and varies with essay characteristics. In particular, compared to human raters, LLMs tend to assign higher scores to short or underdeveloped essays, while assigning lower scores to longer essays that contain minor grammatical or spelling errors. We...

Ver ficha completa

Style, sentiment, and quality of undergraduate writing in the AI era: A cross-sectional and longitudinal analysis of 4,820 authentic empirical reports
Matthew H.C. Mak, Lukasz Walasek (2025)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial escritura académica educación superior

Temas IA: IA y enseñanza-aprendizaje de lenguas análisis de producción de IA estudio empírico

Resumen:

Texto completo

As generative artificial intelligence (GenAI) becomes widespread in education, its influence on students' academic practice raises concern. We conducted pre-registered analyses of 4820 empirical reports authored by 2000 psychology undergraduates (2016–2025) to examine how ChatGPT's launch in November 2022 may influence the style, sentiment, and quality of students' writing. Following ChatGPT's release, prevalence of ChatGPT-associated lexical markers (e.g., delve, intricate) surged until 2024, then declined in 2025, possibly due to some students actively masking GenAI traces. Writing style (indexed by lexical diversity/density/sophistication, nominalisation, readability) became increasingly...

Ver ficha completa

Evaluating the performance of ChatGPT and GPT-4o in coding classroom discourse data: A study of synchronous online mathematics instruction
Simin Xu, Xiaowei Huang, Chung Kwan Lo (2024)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial ChatGPT grandes modelos de lenguaje

Temas IA: análisis de producción de IA chatGPT modelos de lenguaje (LLM)

Resumen:

Texto completo

High-quality instruction is essential to facilitating student learning, prompting many professional development (PD) programmes for teachers to focus on improving classroom dialogue. However, during PD programmes, analysing discourse data is time-consuming, delaying feedback on teachers' performance and potentially impairing the programmes' effectiveness. We therefore explored the use of ChatGPT (a fine-tuned GPT-3.5 series model) and GPT-4o to automate the coding of classroom discourse data. We equipped these AI tools with a codebook designed for mathematics discourse and academically productive talk. Our dataset consisted of over 400 authentic talk turns in Chinese from synchronous online...

Ver ficha completa

Do teachers spot AI? Evaluating the detectability of AI-generated texts among student essays
Johanna Fleckenstein, Jennifer Meyer, Thorben Jansen (2024)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial integridad académica evaluación

Temas IA: IA y enseñanza-aprendizaje de lenguas análisis de producción de IA IA y evaluación

Resumen:

Texto completo

The potential application of generative artificial intelligence (AI) in schools and universities poses great challenges, especially for the assessment of students’ texts. Previous research has shown that people generally have difficulty distinguishing AI-generated from human-written texts; however, the ability of teachers to identify an AI-generated text among student essays has not yet been investigated. Here we show in two experimental studies that novice (N = 89) and experienced teachers (N = 200) could not identify texts generated by ChatGPT among student-written texts. However, there are some indications that more experienced teachers made more differentiated and more accurate...

Ver ficha completa

Validating AI-generated classroom observations: Reliability, accuracy, and limits of LLM-based pedagogical judgment
Carolina Melo, Javiera de la Maza, Matías Recabarren (2026)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial grandes modelos de lenguaje práctica docente

Temas IA: modelos de lenguaje (LLM) análisis de producción de IA estudio empírico

Resumen:

Texto completo

This study examines the reliability and accuracy of large language models (LLMs) for automated classroom observation using the World Bank's TEACH Primary framework. As education systems increasingly explore AI-based tools to scale teacher feedback and professional development, empirical validation of these systems is critical. Using a corpus of 12 primary classroom videos, we compared 8618 AI-generated evaluations from eight LLM endpoints against consensus-based ratings from certified TEACH experts. To account for model stochasticity, each model produced 10 independent evaluations per video–element pair. Reliability was assessed using variability and inter-rater consistency indicators,...

Ver ficha completa

Leveraging generative AI for course learning outcome categorization using Bloom's taxonomy
Omaima Almatrafi, Aditya Johri (2025)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial grandes modelos de lenguaje diseño curricular

Temas IA: modelos de lenguaje (LLM) análisis de producción de IA estudio empírico

Resumen:

Texto completo

Learning outcomes are clear and concise statements that describe what students should be able to do or know at the end of a particular course. These statements are crucial in instructional planning, curriculum development, and assessment of student progress and learning. Although there is no universal guidance on how to develop learning outcomes, Bloom's taxonomy is one widely used framework that helps instructors develop outcomes that reflect different levels of thinking, from basic remembering to creative problem-solving. This study investigates the potential of generative AI, specifically GPT-4, in classifying course learning outcomes according to their respective cognitive levels within...

Ver ficha completa

Evaluating the psychometric properties of ChatGPT-generated questions
Shreya Bhandari, Yunting Liu, Yerin Kwak (2024)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial evaluación ChatGPT

Temas IA: IA y evaluación IA y creación de materiales análisis de producción de IA

Resumen:

Texto completo

Not much is known about how LLM-generated questions compare to gold-standard, traditional formative assessments concerning their difficulty and discrimination parameters, which are valued properties in the psychometric measurement field. We follow a rigorous measurement methodology to compare a set of ChatGPT-generated questions, produced from one lesson summary in a textbook, to existing questions from a published Creative Commons textbook. To do this, we collected and analyzed responses from 207 test respondents who answered questions from both item pools and used a linking methodology to compare IRT properties between the two pools. We find that neither the difficulty nor discrimination...

Ver ficha completa

Applying psychometric methods to distinguish between human and generative AI responses to multiple-choice assessments
Alona Strugatski, Giora Alexandron (2026)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial integridad académica evaluación

Temas IA: IA y evaluación análisis de producción de IA estudio empírico

Resumen:

Texto completo

The growing use of generative AI (GenAI) tools like ChatGPT raises serious concerns about academic integrity, especially in the context of assessments. While detection efforts have focused on open-ended responses, multiple-choice questions (MCQs), which are common in high-stakes testing, remain largely overlooked, partly due to their perceived detection difficulty. The present work establishes a theoretical and empirical foundation for applications of psychometric theory to separate GenAI and human responses. Specifically, it investigates whether person-fit statistics (PFS), a class of methods within Item Response Theory (IRT) used to evaluate how well an examinee’s response pattern fits...

Ver ficha completa

Can large language models write reflectively
Yuheng Li, Lele Sha, Lixiang Yan (2023)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial grandes modelos de lenguaje expresión escrita

Temas IA: análisis de producción de IA modelos de lenguaje (LLM) chatGPT

Resumen:

Texto completo en HTML

Generative Large Language Models (LLMs) demonstrate impressive results in different writing tasks and have already attracted much attention from researchers and practitioners. However, there is limited research to investigate the capability of generative LLMs for reflective writing. To this end, in the present study, we have extensively reviewed the existing literature and selected 9 representative prompting strategies for ChatGPT – the chatbot based on state-of-art generative LLMs to generate a diverse set of reflective responses, which are combined with student-written reflections. Next, those responses were evaluated by experienced teaching staff following a theory-aligned...

Ver ficha completa

How reliable are large language models in analyzing the quality of written lesson plans? A mixed-methods study from a teacher internship program
Dennis Hauk, Nina Soujon (2026)
Computers & Education: Artificial Intelligence
Detalles Cerrar βœ•

Tipo: artículo

Metodología: estudio empírico

Temas: inteligencia artificial grandes modelos de lenguaje formación de profesores

Temas IA: modelos de lenguaje (LLM) IA y evaluación análisis de producción de IA

Resumen:

Texto completo

This study investigates the reliability of Large Language Models (LLMs) in evaluating the quality of written lesson plans from pre-service teachers. A total of 32 lesson plans, each ranging from 60 to 100 pages, were collected during a teacher internship program for civic education pre-service teachers. Using the ChatGPT-o1 reasoning model, we compared a human expert standard with LLM coding outcomes in a two-phase explanatory sequential mixed-methods design that combined quantitative reliability testing with a qualitative follow-up analysis to interpret inter-dimensional patterns of agreement. Quantitatively, overall reliability across six qualitative components of written lessons plans (...

Ver ficha completa

Paginación

  • Primera página «
  • Página anterior ‹
  • Página 1
  • Página 2
  • Página 3
  • Página actual 4
  • Página 5
  • Página 6
  • Siguiente página ›
  • Última página »

Etiquetas

  • análisis de producción de IA (54)
  • catálogo de herramientas (15)
  • chatbots (68)
  • chatGPT (155)
  • conceptos básicos (38)
  • estudio empírico (487)
  • ética de la IA (119)
  • guía (26)
  • IA y aprendizaje (159)
  • IA y creación de materiales (63)
  • IA y educación (360)
  • IA y ELE (31)
  • IA y enseñanza-aprendizaje de lenguas (208)
  • IA y evaluación (140)
  • IA y variedades del español (1)
  • modelos de lenguaje (LLM) (87)
  • prompts (65)
  • revisión de bibliografía (136)
Sobre Todoele Índice Publica Contacto: todoele@gmail.com
Política de privacidad Créditos